CFP last date
20 October 2026
Reseach Article

Sentiment Analysis of Code-Mixed Punjabi–English Social Media Text: A Survey of Approaches, Datasets, and Open Challenges

by Rajvinder Kaur
International Journal of Computer Applications
Foundation of Computer Science (FCS), NY, USA
Volume 187 - Number 141
Year of Publication: 2026
Authors: Rajvinder Kaur
10.5120/ijca7e7e541c0c0b

Rajvinder Kaur . Sentiment Analysis of Code-Mixed Punjabi–English Social Media Text: A Survey of Approaches, Datasets, and Open Challenges. International Journal of Computer Applications. 187, 141 ( Sep 2026), 50-58. DOI=10.5120/ijca7e7e541c0c0b

@article{ 10.5120/ijca7e7e541c0c0b,
author = { Rajvinder Kaur },
title = { Sentiment Analysis of Code-Mixed Punjabi–English Social Media Text: A Survey of Approaches, Datasets, and Open Challenges },
journal = { International Journal of Computer Applications },
issue_date = { Sep 2026 },
volume = { 187 },
number = { 141 },
month = { Sep },
year = { 2026 },
issn = { 0975-8887 },
pages = { 50-58 },
numpages = {9},
url = { https://ijcaonline.org/archives/volume187/number141/sentiment-analysis-of-code-mixed-punjabienglish-social-media-text-a-survey-of-approaches-datasets-and-open-challenges/ },
doi = { 10.5120/ijca7e7e541c0c0b },
publisher = {Foundation of Computer Science (FCS), NY, USA},
address = {New York, USA}
}
%0 Journal Article
%1 2026-09-19T02:57:22.007942+05:30
%A Rajvinder Kaur
%T Sentiment Analysis of Code-Mixed Punjabi–English Social Media Text: A Survey of Approaches, Datasets, and Open Challenges
%J International Journal of Computer Applications
%@ 0975-8887
%V 187
%N 141
%P 50-58
%D 2026
%I Foundation of Computer Science (FCS), NY, USA
Abstract

Sentiment analysis of user-generated text has matured rapidly for high-resource languages, but Punjabi—despite a speaker population exceeding one hundred million—remains poorly served by annotated corpora, lexicons, and evaluation benchmarks. The problem is compounded on social media, where users routinely produce Punjabi–English code-mixed text written variously in Gurmukhi, Shahmukhi, and Romanized Punjabi, with unstandardized spelling, transliteration variants, slang, and informal grammar. This paper surveys research on sentiment analysis of Punjabi and Punjabi–English code-mixed text. Following a documented search protocol, this survey identifies and critically examines the primary studies published between 2014 and 2025, covering lexicon- and rule-based methods, classical machine learning classifiers, deep neural architectures, and multilingual transformer models. This survey catalogues the datasets used in this line of work, including their domains, scripts, sizes, and availability; summarizes the preprocessing pipelines, feature representations, and evaluation metrics reported; and tabulates the published results of each study to enable direct comparison. The survey shows that the genuinely code-mixed Punjabi–English literature is still small and fragmented: datasets are largely private and domain-specific, no shared benchmark exists, and transformer-based methods have so far been demonstrated mainly on the related Urdu–Punjabi (Shahmukhi) setting rather than on Romanized Punjabi–English text. This survey closes by identifying concrete research needs—public benchmark corpora, script normalization and transliteration resources, and Punjabi-adapted pretrained models—that would allow this area to progress beyond isolated proof-of-concept studies.

References
  1. M. Kumar, L. Khan, and H.-T. Chang, “Evolving techniques in sentiment analysis: A comprehensive review,” PeerJ Computer Science, vol. 11, e2592, 2025, doi: 10.7717/peerj-cs.2592.
  2. D. M. Eberhard, G. F. Simons, and C. D. Fennig, Eds., Ethnologue: Languages of the World, 27th ed. Dallas, TX, USA: SIL International, 2024.
  3. Office of the Registrar General & Census Commissioner, India, Census of India 2011: Language and Mother Tongue Data. New Delhi, India: Ministry of Home Affairs, Government of India, 2011.
  4. S. Hochreiter and J. Schmidhuber, “Long short-term memory,” Neural Computation, vol. 9, no. 8, pp. 1735–1780, 1997.
  5. Y. Kim, “Convolutional neural networks for sentence classification,” in Proc. Conf. Empirical Methods in Natural Language Processing (EMNLP), 2014, pp. 1746–1751.
  6. J. Devlin, M.-W. Chang, K. Lee, and K. Toutanova, “BERT: Pre-training of deep bidirectional transformers for language understanding,” in Proc. NAACL-HLT, 2019, pp. 4171–4186.
  7. A. Conneau et al., “Unsupervised cross-lingual representation learning at scale,” in Proc. 58th Annu. Meeting Assoc. Computational Linguistics (ACL), 2020, pp. 8440–8451.
  8. D. Kakwani et al., “IndicNLPSuite: Monolingual corpora, evaluation benchmarks and pre-trained multilingual language models for Indian languages,” in Findings of EMNLP, 2020, pp. 4948–4961.
  9. S. Khanuja et al., “MuRIL: Multilingual representations for Indian languages,” arXiv preprint arXiv:2103.10730, 2021.
  10. K. Bali, J. Sharma, M. Choudhury, and Y. Vyas, “‘I am borrowing ya mixing?’ An analysis of English-Hindi code mixing in Facebook,” in Proc. 1st Workshop on Computational Approaches to Code Switching, 2014, pp. 116–126.
  11. R. Bhargava, Y. Sharma, and S. Sharma, “Sentiment analysis for mixed script Indic sentences,” in Proc. Int. Conf. Advances in Computing, Communications and Informatics (ICACCI), 2016, pp. 524–529.
  12. S. Swami, A. Khandelwal, V. Singh, S. S. Akhtar, and M. Shrivastava, “A corpus of English-Hindi code-mixed tweets for sarcasm detection,” arXiv preprint arXiv:1805.11869, 2018.
  13. P. Patwa et al., “SemEval-2020 Task 9: Overview of sentiment analysis of code-mixed tweets,” in Proc. 14th Workshop on Semantic Evaluation (SemEval-2020), 2020, pp. 774–790.
  14. S. Khanuja, S. Dandapat, A. Srinivasan, S. Sitaram, and M. Choudhury, “GLUECoS: An evaluation benchmark for code-switched NLP,” in Proc. 58th Annu. Meeting Assoc. Computational Linguistics (ACL), 2020, pp. 3575–3585.
  15. P. Arora and B. Kaur, “Sentiment analysis of political reviews in Punjabi language,” International Journal of Computer Applications, vol. 126, no. 14, pp. 20–23, 2015, doi: 10.5120/ijca2015906297.
  16. A. Kaur and V. Gupta, “A novel approach for sentiment analysis of Punjabi text using SVM,” International Arab Journal of Information Technology, vol. 14, no. 5, 2017.
  17. J. Singh, G. Singh, R. Singh, and P. Singh, “Morphological evaluation and sentiment analysis of Punjabi text using deep learning classification,” Journal of King Saud University – Computer and Information Sciences, vol. 33, no. 5, pp. 508–517, 2021, doi: 10.1016/j.jksuci.2018.04.003.
  18. M. Singh, V. Goyal, and S. Raj, “Sentiment analysis of English-Punjabi code mixed social media content for agriculture domain,” in Proc. 4th Int. Conf. Information Systems and Computer Networks (ISCON), 2019, doi: 10.1109/ISCON47742.2019.9036204.
  19. K. Yadav, A. Lamba, D. Gupta, A. Gupta, P. Karmakar, and S. Saini, “Bilingual sentiment analysis for a code-mixed Punjabi English social media text,” in Proc. 2020 5th Int. Conf. Computing, Communication and Security (ICCCS), 2020, pp. 1–5, doi: 10.1109/ICCCS49678.2020.9277309.
  20. A. Tiwari, J. Sehgal, M. Singh, and A. Mishra, “Sentiment analysis in English-Punjabi mixed social media posts,” in Proc. 2025 IEEE Int. Conf. Interdisciplinary Approaches in Technology and Management for Social Innovation (IATMSI), vol. 3, 2025, pp. 1–6.
  21. S. Sekhri and K. Bansal, “A review on Punjabi language sentiment analysis using machine learning,” Journal of Computer Technology & Applications, vol. 16, no. 2, pp. 116–122, 2025.
  22. S. Sekhri, N. Shrivastav, P. Kaur, K. Bansal, S. Wadhwa, and P. Bhalla, “Punjabi text sentiment analysis using Indic NLP,” OmniScience: A Multi-disciplinary Journal, vol. 15, no. 1, pp. 38–42, 2025.
  23. M. Hussain, S. Ali, H. Sattar, A. Raza, M. H. Akbar, and M. A. Rafiq, “Urdu-Punjabi code switched sentiment analysis empowered by a deep learning framework integrating XLM-R and GPT,” VAWKUM Transactions on Computer Sciences, vol. 13, no. 2, pp. 1–20, 2025, doi: 10.21015/vtcs.v13i2.2144.
  24. M. Shabbir et al., “Advancing NLP for Shahmukhi Punjabi: Word embedding and text classification with a novel dataset,” VAWKUM Transactions on Computer Sciences, vol. 13, no. 1, pp. 22–43, 2025.
  25. T. Mikolov, I. Sutskever, K. Chen, G. S. Corrado, and J. Dean, “Distributed representations of words and phrases and their compositionality,” in Advances in Neural Information Processing Systems (NeurIPS), 2013, pp. 3111–3119.
  26. J. Pennington, R. Socher, and C. D. Manning, “GloVe: Global vectors for word representation,” in Proc. Conf. Empirical Methods in Natural Language Processing (EMNLP), 2014, pp. 1532–1543.
  27. P. Bojanowski, E. Grave, A. Joulin, and T. Mikolov, “Enriching word vectors with subword information,” Transactions of the Association for Computational Linguistics, vol. 5, pp. 135–146, 2017.
Index Terms

Computer Science
Information Sciences

Keywords

Sentiment analysis code-mixed text Punjabi–English low-resource languages natural language processing transformer models social media