$$\rightleftharpoonup{xx}$$
$$\longleftharp{xx}$$,
$$\longrightharp{xx}$$,
When two or more languages are mixed together in a single line or speech, this is called a code-mixed language. It is common in casual dialog like Hinglish. There are multiple ways human emotions can be understood, and to computationally model a series of emotional statements is to annotate them by the people who uttered those sentences. It can be understood in terms of biological, physiological, psychological levels, and so on. According to scientists such as Roger Penrose, many phenomena in our world are non-computational, and scientists such as Wolfram consider that everything (every phenomenon) can be modeled computationally1. Penrose believes that consciousness involves processes (perhaps related to quantum mechanics within the brain) that go beyond what any step-by-step algorithmic procedure can achieve. He often cites Gödel's incompleteness theorems to support the idea that human mathematical insight, for instance, transcends formal systems2. If consciousness is non-computational, then emotions, as a key aspect of conscious experience, might also have non-computational elements. Stephen Wolfram, known for Mathematica and his work on cellular automata, proposes the "Principle of Computational Equivalence." This suggests that even very complex systems, including potentially the universe itself and phenomena within it (like emotions), can ultimately be described and modeled by computational rules, even if those rules are very simple, generating complex behavior. But practically, this is not possible, and we need someone who is referred to as either an expert or simply an annotator who can do emotion analysis3.
In this research, we propagate the idea of building computational models. But that model will be quasi-computational. Our research in this context aims to be computational in form but might not capture all aspects perfectly, perhaps leaving room for complexities that are difficult or impossible to compute fully. Emotions are difficult to model computationally because they depend on subjective experiences, cultural context, and nuanced expressions that cannot be fully captured through fixed algorithms.
Therefore, for modeling human emotions using variable-based computational approaches, it is necessary to annotate human emotional utterances. This annotation should be performed by an expert or an annotator skilled in emotion analysis¹. Understanding the complexities of human emotions is not an easy task, particularly when dealing with mixed languages. Furthermore, problems related to scale mean that relying solely on manual annotation by humans is not a viable option. Recent research indicates a consistent need for a human-in-the-loop approach when building systems for such complex tasks. Consequently, a semi-automatic approach, which involves automating the more straightforward parts while reserving tasks requiring human nuance for annotators, appears most appropriate for developing natural language systems in this domain.
A human annotator will, of course, be doing work manually, and in the age of computation, this is not what is expected from contemporary scientists. If the annotator (manual, semi-automatic, or fully automatic) is able to intelligently guess the type of emotion embodied in the utterances, utterances that consist of multiple types of emotions expressed as symbols, with colloquialism or code-mixed and using multiple modalities, then the task is hard and easy at the same time. The complexity of emotion annotation in Hinglish utterances depends on the nature of the expression. When emotions are conveyed clearly using familiar words or emojis, annotation is relatively straightforward. However, the task becomes challenging when utterances involve multiple emotions, code-mixing, or ambiguous symbolic expressions. Therefore, annotation can be both easy and difficult, depending on how directly the emotion is expressed.
Contemporary approaches in the identification of emotions and sentiments deal with these challenges, including the subjective nature of emotions, the ambiguity in human expressions, the complexity of code-mixed languages like Hinglish, and the time-consuming and inconsistent nature of manual annotation. associated with constructing computational models and managing tedious annotation tasks. Recent research indicates that researchers are employing a diverse range of methods toward this goal, including machine learning, deep learning, and various hybrid approaches. Recent research shows that to overcome these issues, researchers are employing a variety of techniques, such as machine learning, deep learning, and hybrid models.
Recent research shows that researchers are employing all kinds of approaches, including machine learning, deep learning3, and hybrid approaches. The term sentiment analysis refers to a procedure used when the polarity of the emotions is believed to be a marker to understand the raw emotion of humans3,4. The development of such technology has helped to recognize mood, sentiments, speech, facial emotions, and nonverbal cues, and has already made inroads into applications that allow real-time translation2. A multimodal approach could be used to translate Hinglish into English and may be helpful in the future to make Indian cinema more accessible to remote societies5,6. For example, in India, English is often the second language. Research in this context shows that this has improved the quality of English teaching by analyzing Indian speech (mix-code language) for the expressiveness, or degree of feeling and emotion, of each word.
Within this research context, the use of mixed-code language in conjunction with translation has been shown to enhance the quality of English teaching. This is accomplished through the analysis of Indian speech (mixed-code language) to determine the expressiveness, or emotional valence, of each word. Through the application of deep learning to train computers in speech interpretation, this research has already improved the accuracy of computerized speech analysis and facilitated a greater understanding of communication4,5. According to the census results from 2001, Hinglish, a language that is a blend of Hindi and English, is currently used by an estimated 120 million people in India6.
From the contemporary landscape of learning algorithms, it is clear that active learning has emerged as a powerful tool for significantly reducing human effort in annotating large datasets, particularly in the domain of emotion identification and recognition. This iterative approach, which selectively annotates impactful annotations (with appropriate metrics), not only enhances annotation accuracy but also improves efficiency5. Previous studies have demonstrated its effectiveness in achieving substantial reductions in manual annotation workload while maintaining or even improving performance with smaller training datasets and proposing a cluster analysis-based method for informative instance selection7,8. In the specific context of Hinglish emotion recognition, researchers have made valuable contributions through deep learning models and a multi-label annotated dataset9,10,11. Previous studies12,13 have introduced active learning and semi-supervised methods to minimize the dependence on human-labeled data, further enhancing efficiency and reducing annotation costs. Furthermore, active learning has been demonstrated in many projects to boost classification performance, particularly in multi-label emotion classification14.
The efficacy of active learning in improving classifier performance has been recognized across various machine-learning applications. Studies15,16highlighted its crucial role in enhancing performance by focusing on educational applications. Similarly, an early study introduced a novel algorithm for active learning with support vector machines, significantly reducing the need for labeled instances17. Another work also explored its application in tasks involving structured instances, such as text classification18. Active learning's impact on emotion recognition tasks extends beyond efficiency gains, particularly in minimizing reliance on human-labeled data. One study introduced a multi-task framework for emotion classification and regression, surpassing the performance of single-task methods10.
Furthermore, researchers19made significant strides in speech and text emotion recognition using active learning, while demonstrating20 its effectiveness in personalized music emotion classification. However, the process of categorizing and labelling emotions presents a significant challenge, as highlighted21,22, particularly in sentiment analysis contexts. Notes that label usage can significantly influence emotion categorization, particularly for later-learned categories23. To address these challenges, various algorithms, including keyword-based and learning-based methods, have been developed, achieving notable accuracy rates24. Research on emotions based on written utterances and texts has been explored in numerous models, and approaches have implemented a dimensional model using normative databases for effective emotion detection25. In another study26, a cognitive emotion model enhanced a sequential method used for social emotion cause identification. The author provided a computational linguistic interpretation of the OCC emotion model, while a similar study27proposed a system utilizing ontologies for representing word dependency relations and emotions. The authors of one study28discussed the signals that correlate with emotional word processing, highlighting the brain's adaptation in expressing emotions in written language. Annotation of multiple arrays of raw emotions, including that of the multi-model data, is challenging. Nevertheless, investigating emotions related to war and conflict provides a scientific and systematic window into the human psyche under extreme circumstances, allowing us to better understand how individuals and communities cope with trauma, loss, and uncertainty5. Another study found that the annotation technique effectively enhanced genre classification, with the title feature playing a crucial role in the process29. One study created a 44K vision-touch dataset with expert and GPT-4V to train a tactile encoder and a TVL model for text generation30. Another study explored opinion and trend mining on political tweets, focusing on the active learning process to automatically annotate French-language tweets about politicians41. Another study introduced CloudFlows, a cloud-based scientific workflow platform designed for dynamic adaptive central analysis in data streams. It enables active learning to improve sentiment classification, allowing the algorithm to adapt to changes in real-time data42.
There is clear-cut tension between the complexity of human emotion and the desire for automated emotion analysis. An inherent tension exists between the complexity of human emotion and the objective of automated emotion analysis. Most of the contemporary work acknowledges the limitations of manual annotation and emphasizes the need for sophisticated computational methods to tackle the challenges of understanding emotions in diverse forms of communication. This ideal scenario is largely impractical, i.e., getting annotations from the people who wrote or spoke the sentences43. The ideal scenario for obtaining data, specifically getting annotations directly from the individuals who wrote or spoke the sentences, is largely impractical. This impracticality stems from the impossibility of gathering and processing such personalized annotations on a large scale. Therefore, current efforts must rely on expert annotators or automated emotion detection algorithms to analyze and label emotions expressed in text. In this research work, we have attempted to overcome some aspects of these domain challenges. The key contributions in this problem area are presented hereafter44.
Therefore, we need to rely on experts or annotators and emotion detection algorithms to analyze and label the emotions expressed in text. It's impossible to gather and process such personalized annotations on a large scale. Hence, in this research work, we have attempted to overcome some aspects of this domain knowledge. The following are the key contributions in this problem area.
The framework works together with rule-based methods like emotion tagging, code-mix detection, and emoji interpretation with machine learning techniques such as Random Forest and word embeddings, improving annotation accuracy while reducing manual effort. The Iterative Learning of the Classifier employs active learning as well as transfer learning to prioritize ambiguous feature samples, reducing the need for hard labor. This approach reduced operational costs by 40% compared to hard manual labeling.
To handle the nuances of Hinglish at a granular level, a custom context-sensitive tokenization method was developed. This approach processes code-mixed text by accounting for language switching, punctuation, emojis, and subword segmentation, enabling more accurate emotion annotation in mixed Hindi-English text. At a granular level, we developed custom context-sensitive tokenization for Hinglish text. The framework addresses the complexities of code-mixed text by incorporating bilingual emotion dictionaries, subword tokenization, and custom context-sensitive tokenization. Lexical rules resolved 89% of code-switching ambiguities.
Our work is grounded in established psychological theories of emotion, such as Discrete Emotions Theory and Cognitive Appraisal Theory. The research demonstrates the scalability of the approach for crisis response and social media monitoring, providing a blueprint for low-resource, multilingual NLP applications.
Table 1 explains the available studies for the same problem domain. From the literature survey and the tabulated summary, it can be inferred that most of the studies cannot escape doing some initial work on annotation using manual methods. Few researchers are following semi-automatic approaches41. However, the real difference in performance comes from the use of an effective learning model that can automate the process of annotation. The emotional content of the tweets must match theories that explain the pathways of humans' emotions and the organization of sentiments. The next section defines the problem based on limitations of existing approaches and empirical results of the papers.
| Study | Dataset | Emotion | Methods | Domain | Labeling Process | Gaps | Future Scope |
| [31] | 9,000,000 Tweets | tension, depression, anger, vigour, fatigue, | Confusion Profile of Mood States | English | No labeling | The study overlooks subtle emotional differences like surprise, joy, or fear, suggesting that emotion labelling can enhance the interpretability and granularity of sentiment trends, particularly in relation to socioeconomic events. | It could investigate how to better capture and examine a range of emotional expressions in social media data by utilising automated categorisation methods and well-established emotion taxonomies. |
| [32] | 7000 Tweets | anger, disgust, fear, joy, love, sadness, | Support Vector Machine | English | Manual | The generalisability of the dataset is limited due to its topic-specificity and lack of representativeness of overall Twitter usage. Due to subjective interpretation and minimal context, which is shown in modest inter-annotator agreement, it is challenging to annotate emotions in brief, casual tweets. | Future work will focus on developing improved emotion detection models by incorporating distinctions between topic-specific and emotion-specific linguistic styles, enabling more accurate classification in diverse tweet contexts. |
| [33] | 21,000 Tweet | anger, disgust, fear, joy, sadness, surprise | Support Vector Machine | ------ | Using Hashtag | Existing emotion-labeled corpora are limited in size and domain, lacking large, diverse datasets for microblogs. Tweets are short, noisy, and context-limited, making accurate emotion detection and annotation difficult. | In future work study may includes expanding the emotion lexicon with synonyms and additional hashtags to improve coverage and detection accuracy. |
| [34] | 16485 Tweets | anger, disgust, fear, joy, sadness, surprise | Support Vector Regression | Chinese | Manual | Traditional emotion classification methods often overlook the underlying cause of emotions, limiting feature quality.
Accurately extracting emotion causes from short, informal microblog posts requires robust rule-based systems and domain knowledge.
| Further exploration of emotion cause analysis can enhance emotion detection models and open new directions in textual emotion understanding. |
| [35] | 10,040 Tweet | Fear, hope, joy anger, surprise, sad, disgust | LDA, inter-rater agreement | Hinglish | Manual | There is a lack of publicly available, structured datasets for Hinglish, especially those capturing pragmatic and emotional nuances in crisis related content. Hinglish is a non-standard, code-mixed language, and regional variations complicate accurate sentiment analysis and annotation.
|
To expand multimodal datasets, integrate deep pragmatic analysis with machine learning models, and address scalability for real-time emotion tracking in conflict discourse. |
| [36] | 134,000 tweets | active, inactive happy, unhappy | Support Vector Machine and K-Nearest Neighbors | Hinglish | Using Hashtags | Manual emotion labeling of tweets is labor-intensive and inconsistent, limiting large-scale emotion classification efforts
Crowdsourced annotations lack reliability, especially in identifying emotion arousal levels, highlighting subjectivity in emotion interpretation.
| Focus on refining hashtag-based labeling and expanding emotion detection models for improved accuracy and generalizability across diverse emotional contexts. |
| [37] | 3,000 students, psychologists and non-psychologistsfrom 37 countries | joy,fear, anger, sadness, disgust, shame, and guilt. | -- | ----- | Manual | Limited exploration of how cultural factors influence the regulation and expression of specific emotions across diverse societies. Balancing evidence for universal emotional patterns with culturally specific variations in emotion elicitation and interpretation remains complex.
| Further studies should investigate the interaction between biological universality and cultural context in shaping emotional experience and communication |
| [38] | 12000 | Happiness, Sadness and Anger | inter-rater agreement | Hindi+English | Manual | Current research lacks a comprehensive, annotated dataset and standardized models for Hinglish emotion detection. The irregular grammar and code-mixed nature of social media texts make accurate emotion classification difficult.
|
Future work will focus on expanding emotion categories and developing larger, multilingual code-mixed datasets. |
| [39] | 2866 | happiness, sadness, anger, surprise and sadness | Support Vector Machine | Hinglish (Hindi+English) | Manual | Lack of emotion-annotated code-mixed datasets. Emotion expression in code-mixed text varies across languages and scripts, making annotation and classification complex.
|
Future work could expand the corpus to include more emotional diversity, integrate part-of-speech tagging, and explore multi-language code-mixed content. |
| [40] | 13738 | --- | Machine Translation Google Translator | Hinglish | Manual | Existing machine translation systems lack accuracy on code-mixed social media data due to the absence of large, domain-specific parallel corpora. High spelling variation, informal structure, and ambiguity in language identification complicate translation of Romanized Hindi-English text.
| The corpus can support development of code-mixed translation systems and be extended to other low-resource languages and NLP tasks like named entity recognition |
| [41] | 11527 | positive,very positive and negative,very negative | kNN-based classification,BOW representation | French politicians | Manual | Limited availability of high-quality annotated datasets for political opinion mining in non-English languages. Balancing annotation noise reduction with information retention and handling uneven label distribution in large-scale tweet datasets are key difficulties.
|
Future work may refine active learning methods to better preserve critical content while minimizing annotation noise in multilingual political discourse. |
| [42] | 764,416 | --- | Kmeans Clustering, SVM | English | Semi supervised | Real-time labeling and model updating in sentiment analysis is constrained by data stream variability, labeling cost, and system scalability. | Future work will explore multi-class sentiment classification, integrate additional labeling strategies, and expand control over initial model generation |
Table 1: Available studies with corresponding labeling methods. The table provides a complete comparative overview of the existing studies, addressing the emotion annotation and establishing the methodological landscape and conceptualizing the contribution of the present work within existing literature.
Problem statement
The most frequently studied emotions in annotation are heavily influenced by foundational psychological models like Ekman's and Plutchik's, primarily focusing on core categories like anger, fear, happiness, sadness, surprise, and so on44. Hence, in this research work, we intend to work on well-established connotations of emotions. The challenge is to develop a dynamic computational framework, F, capable of accurately annotating Hinglish text instances (tᵢ ) from a corpus T focused on wars and conflicts with emotion labels (eᵢ) from a predefined set E = {e1, e2, ..., e8}. This framework must synthesize principles from the Constructionist Theory of Emotion, Affective Events Theory (AET), Discrete Emotions Theory, and Cognitive Appraisal Theory to model the multifaceted emotional landscape of conflict-related discourse. Each text instance tᵢ in T is linguistically complex, blending Hindi (in Roman script), English, emojis, and symbols, necessitating a multi-layered approach to capture nuanced emotional expressions.
The computational model of emotions related to war (as a case study) may involve a multi-faceted approach, beginning with lexical rules that address Hinglish-based nuances. Tokenization, denoted as T, encompasses Roman scripts (Hindi written in Roman script), along with emojis and punctuation, forming the basis of language processing. emotion dictionaries, represented as D, map words across languages to specific emotions, such as anger, joy, and others, where each emotion_i has associated words_j in language_k. Subword decomposition, S, breaks down compound terms into their constituent subwords, enabling a deeper understanding of complex expressions. Subsequently, machine learning techniques, M, utilize embeddings, E, such as Word2Vec/fastText, to transform tokens into vector representations, vector_v, facilitating numerical analysis. Ensemble classifiers, C, like Random Forest, then predict emotion labels, emotion_label_p, from these vector sets. To iteratively improve the annotation learning model, an active learning mechanism, AL, is employed. Expert feedback, F, refines ambiguous cases, ambiguous_sample_q, by assigning refined_label_r, providing crucial corrections. Sample prioritization, P, focuses on low-confidence samples, low_confidence_sample_s, assigning them annotation_priority_t, thus optimizing the annotation process.
By integrating these components and theories, this framework aims to dynamically process Hinglish text, bridge linguistic and cultural nuances, and adaptively refine emotion annotations, offering a scalable solution for analyzing affective dimensions in conflict discourse.