NaijaSenti: A Nigerian Twitter Sentiment Corpus for Multilingual Sentiment Analysis
Citations Over Time
Abstract
Sentiment analysis is one of the most widely studied applications in NLP, but most work focuses on languages with large amounts of data. We introduce the first large-scale human-annotated Twitter sentiment dataset for the four most widely spoken languages in Nigeria (Hausa, Igbo, Nigerian-Pidgin, and Yorùbá ) consisting of around 30,000 annotated tweets per language (and 14,000 for Nigerian-Pidgin), including a significant fraction of code-mixed tweets. We propose text collection, filtering, processing and labeling methods that enable us to create datasets for these low-resource languages. We evaluate a rangeof pre-trained models and transfer strategies on the dataset. We find that language-specific models and language-adaptivefine-tuning generally perform best. We release the datasets, trained models, sentiment lexicons, and code to incentivizeresearch on sentiment analysis in under-represented languages.
Related Papers
- → Genre-based Discourse Analysis of Wedding Invitation Cards in Nigeria: A Comparison between Hausa and Igbo(2019)2 cited
- → Kunle Afolayan, director. October 1. 2014. 115 minutes. English, Nigerian Pidgin, Igbo Yoruba, and Hausa, with Yoruba/English subtitles. Nigeria. GMEDIA. $2.00.(2016)1 cited
- Genre-based Discourse Analysis of Wedding Invitation Cards in Nigeria: A Comparison between Hausa and Igbo(2019)
- → An Appraisal of Social and Political Relationship between Hausa and Igbo in Nigeria(2019)
- → A Brief Introduction to Traditional Music in Nigeria: A Case Study of the Three Major Ethnic Groups.(2021)