Towards Supervised and Unsupervised Neural Machine Translation Baselines for Nigerian Pidgin
arXiv (Cornell University)2020
Citations Over Time
Abstract
Nigerian Pidgin is arguably the most widely spoken language in Nigeria. Variants of this language are also spoken across West and Central Africa, making it a very important language. This work aims to establish supervised and unsupervised neural machine translation (NMT) baselines between English and Nigerian Pidgin. We implement and compare NMT models with different tokenization methods, creating a solid foundation for future works.
Related Papers
- → Unsupervised tokenization for machine translation(2009)55 cited
- → Improving Statistical Machine Translation with Word Class Models(2013)41 cited
- → Language Model Pre-training Method in Machine Translation Based on Named Entity Recognition(2020)14 cited
- → Towards State-of-the-art English-Vietnamese Neural Machine Translation(2017)8 cited
- → Linguistic Features of Pidgin Arabic in Kuwait(2013)11 cited