| International Journal of Computer Applications |
| Foundation of Computer Science (FCS), NY, USA |
| Volume 187 - Number 141 |
| Year of Publication: 2026 |
| Authors: Rajvinder Kaur |
10.5120/ijca7e7e541c0c0b
|
Rajvinder Kaur . Sentiment Analysis of Code-Mixed Punjabi–English Social Media Text: A Survey of Approaches, Datasets, and Open Challenges. International Journal of Computer Applications. 187, 141 ( Sep 2026), 50-58. DOI=10.5120/ijca7e7e541c0c0b
Sentiment analysis of user-generated text has matured rapidly for high-resource languages, but Punjabi—despite a speaker population exceeding one hundred million—remains poorly served by annotated corpora, lexicons, and evaluation benchmarks. The problem is compounded on social media, where users routinely produce Punjabi–English code-mixed text written variously in Gurmukhi, Shahmukhi, and Romanized Punjabi, with unstandardized spelling, transliteration variants, slang, and informal grammar. This paper surveys research on sentiment analysis of Punjabi and Punjabi–English code-mixed text. Following a documented search protocol, this survey identifies and critically examines the primary studies published between 2014 and 2025, covering lexicon- and rule-based methods, classical machine learning classifiers, deep neural architectures, and multilingual transformer models. This survey catalogues the datasets used in this line of work, including their domains, scripts, sizes, and availability; summarizes the preprocessing pipelines, feature representations, and evaluation metrics reported; and tabulates the published results of each study to enable direct comparison. The survey shows that the genuinely code-mixed Punjabi–English literature is still small and fragmented: datasets are largely private and domain-specific, no shared benchmark exists, and transformer-based methods have so far been demonstrated mainly on the related Urdu–Punjabi (Shahmukhi) setting rather than on Romanized Punjabi–English text. This survey closes by identifying concrete research needs—public benchmark corpora, script normalization and transliteration resources, and Punjabi-adapted pretrained models—that would allow this area to progress beyond isolated proof-of-concept studies.