File size: 128,939 Bytes
65cc963
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
55
56
57
58
59
60
61
62
63
64
65
66
67
68
69
70
71
72
73
74
75
76
77
78
79
80
81
82
83
84
85
86
87
88
89
90
91
92
93
94
95
96
97
98
99
100
101
102
103
104
105
106
107
108
109
110
111
112
113
114
115
116
117
118
119
120
121
122
123
124
125
126
127
128
129
130
131
132
133
134
135
136
137
138
139
140
141
142
143
144
145
146
147
148
149
150
151
152
153
154
155
156
157
158
159
160
161
162
163
164
165
166
167
168
169
170
171
172
173
174
175
176
177
178
179
180
181
182
183
184
185
186
187
188
189
190
191
192
193
194
195
196
197
198
199
200
201
202
203
204
205
206
207
208
209
210
211
212
213
214
215
216
217
218
219
220
221
222
223
224
225
226
227
228
229
230
231
232
233
234
235
236
237
238
239
240
241
242
243
244
245
246
247
248
249
250
251
252
253
254
255
256
257
258
259
260
261
262
263
264
265
266
267
268
269
270
271
272
273
274
275
276
277
278
279
280
281
282
283
284
285
286
287
288
289
290
291
292
293
294
295
296
297
298
299
300
301
302
303
304
305
306
307
308
309
310
311
312
313
314
315
316
317
318
319
320
321
322
323
324
325
326
327
328
329
330
331
332
333
334
335
336
337
338
339
340
341
342
343
344
345
346
347
348
349
350
351
352
353
354
355
356
357
358
359
360
361
362
363
364
365
366
367
368
369
370
371
372
373
374
375
376
377
378
379
380
381
382
383
384
385
386
387
388
389
390
391
392
393
394
395
396
397
398
399
400
401
402
403
404
405
406
407
408
409
410
411
412
413
414
415
416
417
418
419
420
421
422
423
424
425
426
427
428
429
430
431
432
433
434
435
436
437
438
439
440
441
442
443
444
445
446
447
448
449
450
451
452
453
454
455
456
457
458
459
460
461
462
463
464
465
466
467
468
469
470
471
472
473
474
475
476
477
478
479
480
481
482
483
484
485
486
487
488
489
490
491
492
493
494
495
496
497
498
499
500
501
502
503
504
505
506
507
508
509
510
511
512
513
514
515
516
517
518
519
520
521
522
523
524
525
526
527
528
529
530
531
532
533
534
535
536
537
538
539
540
541
542
543
544
545
546
547
548
549
550
551
552
553
554
555
556
557
558
559
560
561
562
563
564
565
566
567
568
569
570
571
572
573
574
575
576
577
578
579
580
581
582
583
584
585
586
587
588
589
590
591
592
593
594
595
596
597
598
599
600
601
602
603
604
605
606
607
608
609
610
611
612
613
614
615
616
617
618
619
620
621
622
623
624
625
626
627
628
629
630
631
632
633
634
635
636
637
638
639
640
641
642
643
644
645
646
647
648
649
650
651
652
653
654
655
656
657
658
659
660
661
662
663
664
665
666
667
668
669
670
671
672
673
674
675
676
677
678
679
680
681
682
683
684
685
686
687
688
689
690
691
692
693
694
695
696
697
698
699
700
701
702
703
704
705
706
707
708
709
710
711
712
713
714
715
716
717
718
719
720
721
722
723
724
725
726
727
728
729
730
731
732
733
734
735
736
737
738
739
740
741
742
743
744
745
746
747
748
749
750
751
752
753
754
755
756
757
758
759
760
761
762
763
764
765
766
767
768
769
770
771
772
773
774
775
776
777
778
779
780
781
782
783
784
785
786
787
788
789
790
791
792
793
794
795
796
797
798
799
800
801
802
803
804
805
806
807
808
809
810
811
812
813
814
815
816
817
818
819
820
821
822
823
824
825
826
827
828
829
830
831
832
833
834
835
836
837
838
839
840
841
842
843
844
845
846
847
848
849
850
851
852
853
854
855
856
857
858
859
860
861
862
863
864
865
866
867
868
869
870
871
872
873
874
875
876
877
878
879
880
881
882
883
884
885
886
887
888
889
890
891
892
893
894
895
896
897
898
899
900
901
902
903
904
905
906
907
908
909
910
911
912
913
914
915
916
917
918
919
920
921
922
923
924
925
926
927
928
929
930
931
932
933
934
935
936
937
938
939
940
941
942
943
944
945
946
947
948
949
950
951
952
953
954
955
956
957
958
959
960
961
962
963
964
965
966
967
968
969
970
971
972
973
974
975
976
977
978
979
980
981
982
983
984
985
986
987
988
989
990
991
992
993
994
995
996
997
998
999
1000
1001
1002
1003
1004
1005
1006
1007
1008
1009
1010
1011
1012
1013
1014
1015
1016
1017
1018
1019
1020
1021
1022
1023
1024
1025
1026
1027
1028
1029
1030
1031
1032
1033
1034
1035
1036
1037
1038
1039
1040
1041
1042
1043
1044
1045
1046
1047
1048
1049
1050
1051
1052
1053
1054
1055
1056
1057
1058
1059
1060
1061
1062
1063
1064
1065
1066
1067
1068
1069
1070
1071
1072
1073
1074
1075
1076
1077
1078
1079
1080
1081
1082
1083
1084
1085
1086
1087
1088
1089
1090
1091
1092
1093
1094
1095
1096
1097
1098
1099
1100
1101
1102
1103
1104
1105
1106
1107
1108
1109
1110
1111
1112
1113
1114
1115
1116
1117
1118
1119
1120
1121
1122
1123
1124
1125
1126
1127
1128
1129
1130
1131
1132
1133
1134
1135
1136
1137
1138
1139
1140
1141
1142
1143
1144
1145
1146
1147
1148
1149
1150
1151
1152
1153
1154
1155
1156
1157
1158
1159
1160
1161
1162
1163
1164
1165
1166
1167
1168
1169
1170
1171
1172
1173
1174
1175
1176
1177
1178
1179
1180
1181
1182
1183
1184
1185
1186
1187
1188
1189
1190
1191
1192
1193
1194
1195
1196
1197
1198
1199
1200
1201
1202
1203
1204
1205
1206
1207
1208
1209
1210
1211
1212
1213
1214
1215
1216
1217
1218
1219
1220
1221
1222
1223
1224
1225
1226
1227
1228
1229
1230
1231
1232
1233
1234
1235
1236
1237
1238
1239
1240
1241
1242
1243
1244
1245
1246
1247
1248
1249
1250
1251
1252
1253
1254
1255
1256
1257
1258
1259
1260
1261
1262
1263
1264
1265
1266
1267
1268
1269
1270
1271
1272
1273
1274
1275
1276
1277
1278
1279
1280
1281
1282
1283
1284
1285
1286
1287
1288
1289
1290
1291
1292
1293
1294
1295
1296
1297
1298
1299
1300
1301
1302
1303
1304
1305
1306
1307
1308
1309
1310
1311
1312
1313
1314
1315
1316
1317
1318
1319
1320
1321
1322
1323
1324
1325
1326
1327
1328
1329
1330
1331
1332
1333
1334
1335
1336
1337
1338
1339
1340
1341
1342
1343
1344
1345
1346
1347
1348
1349
1350
1351
1352
1353
1354
1355
1356
1357
1358
1359
1360
1361
1362
1363
1364
1365
1366
1367
1368
1369
1370
1371
1372
1373
1374
1375
1376
1377
1378
1379
1380
1381
1382
1383
1384
1385
1386
1387
1388
1389
1390
1391
1392
1393
1394
1395
1396
1397
1398
1399
1400
1401
1402
1403
1404
1405
1406
1407
1408
1409
1410
1411
1412
1413
1414
1415
1416
1417
1418
1419
1420
1421
1422
1423
1424
1425
1426
1427
1428
1429
1430
1431
1432
1433
1434
1435
1436
1437
1438
1439
1440
1441
1442
1443
1444
1445
1446
1447
1448
1449
1450
1451
1452
1453
1454
1455
1456
1457
1458
1459
1460
1461
1462
1463
1464
1465
1466
1467
1468
1469
1470
1471
1472
1473
1474
1475
1476
1477
1478
1479
1480
1481
1482
1483
1484
1485
1486
1487
1488
1489
1490
1491
1492
1493
1494
1495
1496
1497
1498
1499
1500
1501
1502
1503
1504
1505
1506
1507
1508
1509
1510
1511
1512
1513
1514
1515
1516
1517
1518
1519
1520
1521
1522
1523
1524
1525
1526
1527
1528
1529
1530
1531
1532
1533
1534
1535
1536
1537
1538
1539
1540
1541
1542
1543
1544
1545
1546
1547
1548
1549
1550
1551
1552
1553
1554
1555
1556
1557
1558
1559
1560
1561
1562
1563
1564
1565
1566
1567
1568
1569
1570
1571
1572
1573
1574
1575
1576
1577
1578
1579
1580
1581
1582
1583
1584
1585
1586
1587
1588
1589
1590
1591
1592
1593
1594
1595
1596
1597
1598
1599
1600
1601
1602
1603
1604
1605
1606
1607
1608
1609
1610
1611
1612
1613
1614
1615
1616
1617
1618
1619
1620
1621
1622
1623
1624
1625
1626
1627
1628
1629
1630
1631
1632
1633
1634
1635
1636
1637
1638
1639
1640
1641
1642
1643
1644
1645
1646
1647
1648
1649
1650
1651
1652
1653
1654
1655
1656
1657
1658
1659
1660
1661
1662
1663
1664
1665
1666
1667
1668
1669
1670
1671
1672
1673
1674
1675
1676
1677
1678
1679
1680
1681
1682
1683
1684
1685
1686
1687
1688
1689
1690
1691
1692
1693
1694
1695
1696
1697
1698
1699
1700
1701
1702
1703
1704
1705
1706
1707
1708
1709
1710
1711
1712
1713
1714
1715
1716
1717
1718
1719
1720
1721
1722
1723
1724
1725
1726
1727
1728
1729
1730
1731
1732
1733
1734
1735
1736
1737
1738
1739
1740
1741
1742
1743
1744
1745
1746
1747
1748
1749
1750
1751
1752
1753
1754
1755
1756
1757
1758
1759
1760
1761
1762
1763
1764
1765
1766
1767
1768
1769
1770
1771
1772
1773
1774
1775
1776
1777
1778
1779
1780
1781
1782
1783
1784
1785
1786
1787
1788
1789
1790
1791
1792
1793
1794
1795
1796
1797
1798
1799
1800
1801
1802
1803
1804
1805
1806
1807
1808
1809
1810
1811
1812
1813
1814
1815
1816
1817
1818
1819
1820
1821
1822
1823
1824
1825
1826
1827
1828
1829
1830
1831
1832
1833
1834
1835
1836
1837
1838
1839
1840
1841
1842
1843
1844
1845
1846
1847
1848
1849
1850
1851
1852
1853
1854
1855
1856
1857
1858
1859
1860
1861
1862
1863
1864
1865
1866
1867
1868
1869
1870
1871
1872
1873
1874
1875
1876
1877
1878
1879
1880
1881
1882
1883
1884
1885
1886
1887
1888
1889
1890
1891
1892
1893
1894
1895
1896
1897
1898
1899
1900
1901
1902
1903
1904
1905
1906
1907
1908
1909
1910
1911
1912
1913
1914
1915
1916
1917
1918
1919
1920
1921
1922
1923
1924
1925
1926
1927
1928
1929
1930
1931
1932
1933
1934
1935
1936
1937
1938
1939
1940
1941
1942
1943
1944
1945
1946
1947
1948
1949
1950
1951
1952
1953
1954
1955
1956
1957
1958
1959
1960
1961
1962
1963
1964
1965
1966
1967
1968
1969
1970
1971
1972
1973
1974
1975
1976
1977
1978
1979
1980
1981
1982
1983
1984
1985
1986
1987
1988
1989
1990
1991
1992
1993
1994
1995
1996
1997
1998
1999
2000
2001
2002
2003
2004
2005
2006
2007
2008
2009
2010
2011
2012
2013
2014
2015
2016
2017
2018
2019
2020
2021
2022
2023
2024
2025
2026
2027
2028
2029
2030
2031
2032
2033
2034
2035
2036
2037
2038
2039
2040
2041
2042
2043
2044
2045
2046
2047
2048
2049
2050
2051
2052
2053
2054
2055
2056
2057
2058
2059
2060
2061
2062
2063
2064
2065
2066
2067
2068
2069
2070
2071
2072
2073
2074
2075
2076
2077
2078
2079
2080
2081
2082
2083
2084
2085
2086
2087
2088
2089
2090
2091
2092
2093
2094
2095
2096
2097
2098
2099
2100
2101
2102
2103
2104
2105
2106
2107
2108
2109
2110
2111
2112
2113
2114
2115
2116
2117
2118
2119
2120
2121
2122
2123
2124
2125
2126
2127
2128
2129
2130
2131
2132
2133
2134
2135
2136
2137
2138
2139
2140
2141
2142
2143
2144
2145
2146
2147
2148
2149
2150
2151
2152
2153
2154
2155
2156
2157
2158
2159
2160
2161
2162
2163
2164
2165
2166
2167
2168
2169
2170
2171
2172
2173
2174
2175
2176
2177
2178
2179
2180
2181
2182
2183
2184
2185
2186
2187
2188
2189
2190
2191
2192
2193
2194
2195
2196
2197
2198
2199
2200
2201
2202
2203
2204
2205
2206
2207
2208
2209
2210
2211
2212
2213
2214
2215
2216
2217
2218
2219
2220
2221
2222
2223
2224
2225
2226
2227
2228
2229
2230
2231
2232
2233
2234
2235
2236
2237
2238
2239
2240
2241
2242
2243
2244
2245
2246
2247
2248
2249
2250
2251
2252
2253
2254
2255
2256
2257
2258
2259
2260
2261
2262
2263
2264
2265
2266
2267
2268
2269
2270
2271
2272
2273
2274
2275
2276
2277
2278
2279
2280
2281
2282
2283
2284
2285
2286
2287
2288
2289
2290
2291
2292
2293
2294
2295
2296
2297
2298
2299
2300
2301
2302
2303
2304
2305
2306
2307
2308
2309
2310
2311
2312
2313
2314
2315
2316
2317
2318
2319
2320
2321
2322
2323
2324
2325
2326
2327
2328
2329
2330
2331
2332
2333
2334
2335
2336
2337
2338
2339
2340
2341
2342
2343
2344
2345
2346
2347
2348
2349
2350
2351
2352
2353
2354
2355
2356
2357
2358
2359
2360
2361
2362
2363
2364
2365
2366
2367
2368
2369
2370
2371
2372
2373
2374
2375
2376
2377
2378
2379
2380
2381
2382
2383
2384
2385
2386
2387
2388
2389
2390
2391
2392
2393
2394
2395
2396
2397
2398
2399
2400
2401
2402
2403
2404
2405
2406
2407
2408
2409
2410
2411
2412
2413
2414
2415
2416
2417
2418
2419
2420
2421
2422
2423
2424
2425
2426
2427
2428
2429
2430
2431
2432
2433
2434
2435
2436
2437
2438
2439
2440
2441
2442
2443
2444
2445
2446
2447
2448
2449
2450
2451
2452
2453
2454
2455
2456
2457
2458
2459
2460
2461
2462
2463
2464
2465
2466
2467
2468
2469
2470
2471
2472
2473
2474
2475
2476
2477
2478
2479
2480
2481
2482
2483
2484
2485
2486
2487
2488
2489
2490
2491
2492
2493
2494
2495
2496
2497
2498
2499
2500
2501
2502
2503
2504
2505
2506
2507
2508
2509
2510
2511
2512
2513
2514
2515
2516
2517
2518
2519
2520
2521
2522
2523
2524
2525
2526
2527
2528
2529
2530
2531
2532
2533
2534
2535
2536
2537
2538
2539
2540
2541
2542
2543
2544
2545
2546
2547
2548
2549
2550
2551
2552
2553
2554
2555
2556
2557
2558
2559
2560
2561
2562
2563
2564
2565
2566
2567
2568
2569
2570
2571
2572
2573
2574
2575
2576
2577
2578
2579
2580
2581
2582
2583
2584
2585
2586
2587
2588
2589
2590
2591
2592
2593
2594
2595
2596
2597
2598
2599
2600
2601
2602
2603
2604
2605
2606
2607
2608
2609
2610
2611
2612
2613
2614
2615
2616
2617
2618
2619
2620
2621
2622
2623
2624
2625
2626
2627
2628
2629
2630
2631
2632
2633
2634
2635
2636
2637
2638
2639
2640
2641
2642
2643
2644
2645
2646
2647
2648
2649
2650
2651
2652
2653
2654
2655
2656
2657
2658
2659
2660
2661
2662
2663
2664
2665
2666
2667
2668
2669
2670
2671
2672
2673
2674
2675
2676
2677
2678
2679
2680
2681
2682
2683
2684
2685
2686
2687
2688
2689
2690
2691
2692
2693
2694
2695
2696
2697
2698
2699
2700
2701
2702
2703
2704
2705
2706
2707
2708
2709
2710
2711
2712
2713
2714
2715
2716
2717
2718
2719
2720
# Full Group Thesis (Reference)

**Download:** [PDF](FinalThesisSP.pdf) | [Word source (.docx)](FinalThesisSP.docx)

> This is the complete group thesis *Developing a Sentiment Analysis Model for Code-Mixed
> Hindi-English (Hinglish) Text*, included as reference. It covers all 17 models; the transformer
> (MuRIL, mBART, HingRoBERTa, MPNet) and Sarvam LLM tracks were built by teammates.
> For **Pankaj Biswas's individual contribution** (the BiLSTM / LSTM track), see the
> [model card / report](README.md).

---

> <img src="thesis_media/media/image1.jpeg" style="width:2.92014in;height:1.03819in" alt="A logo for a university Description automatically generated" />
>
> **DEVELOPING A SENTIMENT ANALYSIS MODEL FOR CODE-MIXED HINDI-ENGLISH (HINGLISH) TEXT**
>
> A Project Report submitted
>
> In fulfillment of the requirements for the degree of

### B.Tech (Computer Science & Engineering)

> Submitted by :-
>
> **Pulakala Prithvi Raj (222025042)**
>
> **Pankaj Biswas (222025043)**
>
> **Pritisha Goswami (222025049)**
>
> **B.Tech Computer Science and Engineering**
>
> **8<sup>th</sup> Semester**
>
> **Royal School of Engineering and Technology (RSET)**
>
> Under the guidance of
>
> **Dr. Dillip Rout**
>
> **Assistant Professor**
>
> **Royal School of Engineering and Technology (RSET)**
>
> **THE ASSAM ROYAL GLOBAL UNIVERSITY**

## GUWAHATI: 781035

> **Session: 2022-2026**
>

# 

#  CERTIFICATE OF APPROVAL

# 

> This is to certify that the project report entitled *"Developing a Sentiment Analysis Model for Code-Mixed Hindi-English (Hinglish) Text"* submitted by **Pulakala Prithvi Raj** (Roll No. 222025042), **Pankaj Biswas** (Roll No. 222025043) and **Pritisha Goswami** (Roll No. 222025049), students of B.Tech, 8th semester in the Department of Computer Science & Engineering, Royal School of Engineering and Technology (RSET), The Assam Royal Global University, Guwahati, Assam, has been completed under my supervision. This work is submitted as part of the requirements for the award of the B.Tech degree in Computer Science & Engineering and has not been submitted elsewhere for a degree.
>
> **Project Guide: Signature of the External**
>
> **Dr. Dillip Rout Name of the External**
>
> **Assistant Professor, CSE, RSET**
>
> **Date***:*
>
> **Place: Guwahati**
>

#  FORWARDING CERTIFICATE

# 

> This is to certify that the project report entitled *"Developing a Sentiment Analysis Model for Code-Mixed Hindi-English (Hinglish) Text"* submitted by **Pulakala Prithvi Raj** (Roll No. 222025042), **Pankaj Biswas** (Roll No. 222025043) and **Pritisha Goswami** (Roll No. 222025049), students of B.Tech, 8th semester in the Department of Computer Science & Engineering at Royal School of Engineering and Technology (RSET), The Assam Royal Global University, Guwahati, Assam, under the guidance of **Dr. Dillip Rout, Assistant Professor**, has been evaluated and deemed satisfactory for submission as a requirement for the degree program.
>
> **Date:**
>
> **Place:** Guwahati

**Dr. Dillip Rout**

**Assistant Professor**

**Department of CSE**

> **Royal School of Engineering &**
>
> **Technology**

#  DECLARATION

We, **Pulakala Prithvi Raj** (Roll No. 222025042), **Pankaj Biswas** (Roll No. 222025043) and **Pritisha Goswami** (Roll No. 222025049), hereby declare that the project work entitled *"Developing a Sentiment Analysis Model for Code-Mixed Hindi-English (Hinglish) Text"* was carried out by us under the guidance and supervision of **Dr. Dillip Rout, Assistant Professor**, **Department of Computer Science & Engineering**. This project is submitted for the academic session 2022-2026. We confirm that this work, or any part of it, has not been submitted elsewhere for any other purpose to date.

### Date:

> **Place:** Guwahati

**Pulakala Prithvi Raj Pankaj Biswas Pritisha Goswami**

**222025042 222025043 222025049**

# ACKNOWLEDGMENT

# 

> We are deeply grateful to Royal School of Engineering and Technology for providing us with the resources and environment needed to complete this project titled, *"Developing a Sentiment Analysis Model for Code-Mixed Hindi-English (Hinglish) Text".*
>
> We would like to extend our heartfelt thanks to our guide, Dr. Dillip Rout, Assistant Professor, Department of CSE, Royal School of Engineering and Technology, whose invaluable guidance, encouragement, and insightful feedback have been crucial throughout the project's development. His support enabled us to navigate complex challenges and explore new dimensions in the field of deep learning.
>
> We are also grateful to all faculty members who offered their support, advice, and assistance, both directly and indirectly. Their guidance has played an essential role in shaping this work.
>
> I wish to express my gratitude to my family and friends, whose constant support, encouragement, and patience have been a source of strength throughout this journey. Without their belief in my abilities, this work would not have been possible.
>
> Thank you

**Pulakala Prithvi Raj Pankaj Biswas Pritisha Goswami**

**222025042 222025043 222025049**

#  ABSTRACT

# Code-mixed languages, such as Hinglishβ€”an informal blend of Hindi and Englishβ€”pose significant challenges for sentiment analysis due to inconsistent grammar, transliteration variations, and limited annotated resources. This study presents a comprehensive comparison of classical machine learning and deep learning architectures, including the transformers for binary sentiment classification of Hinglish text. The analysis utilizes the PRISM dataset, comprising 29,550 Hinglish samples labeled as non-hate (0) or hate (1) sources from Kaggle. The text preprocessing included removing URLs, eliminating mentions and hashtags, and normalizing whitespace. Four models were implemented: MuRIL, GloVe+BiLSTM, FastText, and Word2Vec+Logistic Regression. Evaluation metrics include Accuracy, Precision, Recall, F1-score, Specificity, and AUC-ROC to assess the robustness of the models. Experimental results indicate that MuRIL achieved the highest F1 Score (0.737) and AUC (0.824), highlighting the efficacy of multilingual transformers for modeling code-mixed text. Classical models performed worse, though FastText outperformed the Word2Vec and GloVe baselines. The proposed model also surpasses the projects available on Kaggle. The findings emphasize the importance of multiple model evaluations for robust sentiment classification of low-resource, code-mixed social media data.

# 

# 

# 

# 

# 

# 

#  TABLE OF CONTENTS

**Page No**.

> **Certificate of Approval** i

[Forwarding Certificate ii](#forwarding-certificate)

[Declaration iii](#declaration)

Acknowledgement iv

[Abstract](#abstract) v

[List of Tables vi](#list-of-tables)

[Chapter 1. Introduction 1- 4](#_TOC_250042)

1.  Background Study 2

2.  [Problem Statement 2](#problem-statement)

3.  [Motivation 3](#motivation)

4.  Objective 3

5.  [Contributions 4](#_TOC_250039)

Chapter 2. Literature Survey 5-17

1.  [Related work 6](#_TOC_250038)

[Chapter 3. Methodology 18-3](#section-15)2

1.  [Introduction 19](#methodology)

2.  [Block Diagram 2](#the-overall-workflow-involves-of-three-proposed-methodologies-that-is-illustrated-in-fig.-3.1-fig.3.2-and-fig3.3-which-provides-a-high-level-view-of-the-data-preprocessing-feature-extraction-model-training-and-evaluation-process.-it-visually-summarizes-the-pipeline-from-raw-dataset-input-to-performance-evaluation-across-mentioned-models.-the-project-workflow-begins-with-the-primary-input-which-is-raw-sentiment-data-of-english-in-latin-text-hindi-in-devnagiri-texts-and-often-consisting-of-code-mixed-sentences-hindi-english-hinglish-in-latin-text.-this-dataset-is-unstructured-and-not-immediately-suitable-for-tokenization-and-further-word-embeddings.-therefore-the-first-crucial-step-is-text-pre-processing-that-involves-url-removal-lowercasing-whitespace-removal-tokenization-and-embeddings.-this-text-pre-processing-is-critical-because-deep-learning-models-such-as-bert-transformer-models-assign-different-vectors-for-the-same-word-with-different-casings-to-solve-this-we-use-lowercasing-urls-which-do-not-provide-any-context-so-we-remove-urls-and-space-normalization-as-tokens-of-extra-spaces-are-also-created.-after-preprocessing-the-dataset-is-divided-into-training-60-validation-10-and-testing-30-sets-same-for-all-methodologies.-the-processed-data-is-then-fed-into-respective-sentiment-analysis-model-where-the-embedding-techniques-vary-according-to-the-methodology-being-implemented.)0

3.  [Data Collection 2](#_TOC_250032)2

4.  [Data Preprocessing 24](#_TOC_250031)

5.  [Model Selection and development 25](#_TOC_250027)

    1.  ASR Model 25

    2.  [Summarization Model 26](#_TOC_250025)

6.  Evaluation Metrics 27

> 3.6.1 WER 28
>
> 3.6.2 ROUGE Score 28

7.  [Tools and Framework 29](#_TOC_250022)

8.  Setting Up Environment 30

9.  

[Summary 32[Chapter 4. Results and Discussion 33-4](#section-15)0](#_TOC_250015)

[4.1 Result 40](#_TOC_250015)

> [4.1.1 Regular Training 42-43](#_TOC_250015)
>
> [4.1.2 Multi-Stage Training 46-67](#_TOC_250015)

[4.2 Discussion 68](#_TOC_250015)

[[Chapter 5. Conclusion Future scope 41-43](#section-15)](#_TOC_250015)

# LIST OF TABLES

<table>
<colgroup>
<col style="width: 11%" />
<col style="width: 59%" />
<col style="width: 29%" />
</colgroup>
<thead>
<tr>
<th style="text-align: center;"><strong>Table no</strong></th>
<th><blockquote>
<p><strong>Table name</strong></p>
</blockquote></th>
<th style="text-align: left;"><strong>Page No.</strong></th>
</tr>
</thead>
<tbody>
<tr>
<td style="text-align: center;">2.1</td>
<td><blockquote>
<p>Literature Survey</p>
</blockquote></td>
<td style="text-align: left;">9-16</td>
</tr>
<tr>
<td style="text-align: center;">3.1</td>
<td>Dataset Details</td>
<td style="text-align: left;">21</td>
</tr>
<tr>
<td style="text-align: center;">3.2</td>
<td><blockquote>
<p>List Count of URL Noise</p>
</blockquote></td>
<td style="text-align: left;">23</td>
</tr>
<tr>
<td style="text-align: center;">3.3</td>
<td><blockquote>
<p>Data Metrics Before and After Cleaning</p>
</blockquote></td>
<td style="text-align: left;">23</td>
</tr>
<tr>
<td style="text-align: center;">3.4</td>
<td><blockquote>
<p>Dataset Splits Table</p>
</blockquote></td>
<td style="text-align: left;">27</td>
</tr>
<tr>
<td style="text-align: center;">3.5</td>
<td><blockquote>
<p>Model Algorithms</p>
</blockquote></td>
<td style="text-align: left;">30-32</td>
</tr>
<tr>
<td style="text-align: center;">3.6</td>
<td><blockquote>
<p>Model Hyperparameters</p>
</blockquote></td>
<td style="text-align: left;">33-36</td>
</tr>
<tr>
<td style="text-align: center;">4.1</td>
<td><blockquote>
<p>Regular Result</p>
</blockquote></td>
<td style="text-align: left;">40</td>
</tr>
<tr>
<td style="text-align: center;">4.2</td>
<td><blockquote>
<p>English Strategy for Language-wise Regular Training</p>
</blockquote></td>
<td style="text-align: left;">41</td>
</tr>
<tr>
<td style="text-align: center;">4.3</td>
<td><blockquote>
<p>Hindi Strategy for Language-wise Regular Training</p>
</blockquote></td>
<td style="text-align: left;">41</td>
</tr>
<tr>
<td style="text-align: center;">4.4</td>
<td><blockquote>
<p>Hinglish Strategy for Language-wise Regular Training</p>
</blockquote></td>
<td style="text-align: left;">42</td>
</tr>
<tr>
<td style="text-align: center;">4.5</td>
<td><blockquote>
<p>Combined Strategy for Language-wise Regular Training</p>
</blockquote></td>
<td style="text-align: left;">43</td>
</tr>
<tr>
<td style="text-align: center;">4.6</td>
<td><blockquote>
<p>English Strategy for Multistage Language Training</p>
</blockquote></td>
<td style="text-align: left;">43</td>
</tr>
<tr>
<td style="text-align: center;">4.7</td>
<td><blockquote>
<p>Hindi Strategy for Multistage Language Training</p>
</blockquote></td>
<td style="text-align: left;">43</td>
</tr>
<tr>
<td style="text-align: center;">4.8</td>
<td><blockquote>
<p>Hinglish Strategy for Multistage Language Training</p>
</blockquote></td>
<td style="text-align: left;">44</td>
</tr>
<tr>
<td style="text-align: center;">4.9</td>
<td><blockquote>
<p>Combined Strategy for Multistage Language Training</p>
</blockquote></td>
<td style="text-align: left;">45</td>
</tr>
<tr>
<td style="text-align: center;">4.10</td>
<td><blockquote>
<p>Combined Strategy for Multistage Language Taraining</p>
</blockquote></td>
<td style="text-align: left;">45-47</td>
</tr>
<tr>
<td style="text-align: center;">4.11</td>
<td><blockquote>
<p>Performance Evaluation for Sarvam Model</p>
</blockquote></td>
<td style="text-align: left;">47</td>
</tr>
</tbody>
</table>

# LIST OF FIGURES

<table style="width:99%;">
<colgroup>
<col style="width: 11%" />
<col style="width: 70%" />
<col style="width: 16%" />
</colgroup>
<thead>
<tr>
<th><strong>Fig no</strong></th>
<th><blockquote>
<p><strong>Fig Name</strong></p>
</blockquote></th>
<th style="text-align: left;"><strong>Page no</strong></th>
</tr>
</thead>
<tbody>
<tr>
<td>3.1-3.2</td>
<td><blockquote>
<p>Workflow of the proposed methodology for sentiment analysis</p>
</blockquote></td>
<td style="text-align: left;"><blockquote>
<p>19-21</p>
</blockquote></td>
</tr>
<tr>
<td>3.4</td>
<td><blockquote>
<p>Hate vs Non-Hate Class distribution</p>
</blockquote></td>
<td style="text-align: left;"><blockquote>
<p>24</p>
</blockquote></td>
</tr>
<tr>
<td>3.5</td>
<td><blockquote>
<p>Language Distribution</p>
</blockquote></td>
<td style="text-align: left;"><blockquote>
<p>25</p>
</blockquote></td>
</tr>
<tr>
<td>3.6</td>
<td><blockquote>
<p>Feature Correlation Heatmap</p>
</blockquote></td>
<td style="text-align: left;"><blockquote>
<p>25</p>
</blockquote></td>
</tr>
<tr>
<td>3.7</td>
<td><blockquote>
<p>Text length Statistics and Word count Statistics</p>
</blockquote></td>
<td style="text-align: left;"><blockquote>
<p>26</p>
</blockquote></td>
</tr>
<tr>
<td>4.1-4.10</td>
<td><blockquote>
<p>Training vs Validation Loss and Accuracy for Regular Training</p>
</blockquote></td>
<td style="text-align: left;"><blockquote>
<p>48-51</p>
</blockquote></td>
</tr>
<tr>
<td>4.11-4.20</td>
<td><blockquote>
<p>ROC-AUC Curve for Regular Training</p>
</blockquote></td>
<td style="text-align: left;"><blockquote>
<p>51-52</p>
</blockquote></td>
</tr>
<tr>
<td>4.21-4.30</td>
<td><blockquote>
<p>Confusion Matrix for Regular Training</p>
</blockquote></td>
<td style="text-align: left;"><blockquote>
<p>53-54</p>
</blockquote></td>
</tr>
<tr>
<td>4.31-4.39</td>
<td><blockquote>
<p>t-SNE for Regular Training</p>
</blockquote></td>
<td style="text-align: left;"><blockquote>
<p>54-56</p>
</blockquote></td>
</tr>
<tr>
<td>4.40-4.46</td>
<td><blockquote>
<p>Training vs Validation Loss and Accuracy for Language-wise Regular Training</p>
</blockquote></td>
<td style="text-align: left;"><blockquote>
<p>56-59</p>
</blockquote></td>
</tr>
<tr>
<td>4.47-4.53</td>
<td><blockquote>
<p>ROC-AUC Curve for Language-wise Regular Training</p>
</blockquote></td>
<td style="text-align: left;"><blockquote>
<p>59-60</p>
</blockquote></td>
</tr>
<tr>
<td>4.54-4.59</td>
<td><blockquote>
<p>Confusion Matrix for Language-wise Regular Training</p>
</blockquote></td>
<td style="text-align: left;"><blockquote>
<p>60-61</p>
</blockquote></td>
</tr>
<tr>
<td>4.60-4.67</td>
<td><blockquote>
<p>Training vs Validation Loss and Accuracy for Multistage Language Training</p>
</blockquote></td>
<td style="text-align: left;"><blockquote>
<p>62-64</p>
</blockquote></td>
</tr>
<tr>
<td>4.68-4.73</td>
<td><blockquote>
<p>ROC-AUC Curve for Multistage Language Training</p>
</blockquote></td>
<td style="text-align: left;"><blockquote>
<p>64-65</p>
</blockquote></td>
</tr>
<tr>
<td>4.74-4.80</td>
<td><blockquote>
<p>Confusion Matrix for Multistage Language Training</p>
</blockquote></td>
<td style="text-align: left;"><blockquote>
<p>66-67</p>
</blockquote></td>
</tr>
<tr>
<td>4.81</td>
<td><blockquote>
<p>Confusion Matrix for Sarvam Model</p>
</blockquote></td>
<td style="text-align: left;"><blockquote>
<p>67</p>
</blockquote></td>
</tr>
</tbody>
</table>

**CHAPTER 1**

**INTRODUCTION**

## BACKGROUND

## 

Hinglish, which is a blend of English (*Latin*) and Hindi (*Latin*), mainly used in India as informal conversations, which presents unique challenges for Natural Language Processing (NLP) because of spelling variations and informal grammar, sarcastic contexts and frequent code-switching within sentences. Preliminary research was majorly focused on creating annotating the corpuses for Hinglish to assist supervised learning approaches. These datasets commonly included social media posts, its comments, and chat messages mainly informal ones showing code-mixing at the lexical and syntactic levels. Researchers investigated language identification as an initial step, distinguishing English, Hindi, and mixed tokens, which is essential for efficient sentiment analysis. Earlier studies also used Classical Machine Learning classifier models such as NaΓ―ve Bayes, Decision Trees, and SVM using features like n-grams, part-of-speech tags, and lexicons optimized for code-mixed text. Recent studies have applied deep learning models like Long Short-Term Memory (LSTM) models, Recurrent Neural Networks (RNNs) models and Transformer-based models which are more relevant for contextual understanding in comparison to Classical Machine Learning classifier models in code-mixed sentences.

Many works have also shown the challenges in Code-Mixed categories that widely included transliteration and normalization, about Hindi words written in Latin script with irregular spelling as usually there are multiple spelling variations for single Hindi word.

Various techniques such as embedding-based representations and pronunciation-based comparison have been put forward to deal with these variations effectively. Overall, the background research highlights the complexity of sentiment analysis in Hinglish due to linguistic divergence, lack of normalized spelling system, and the dynamic nature of code-switching. This has influenced the development of specialized datasets, feature extraction methods, and model architectures optimized to the distinctions of code-mixed language.

## PROBLEM STATEMENT

## 

The swift growth of social media platforms has resulted in a notable rise in user contents, which is generally written in code-mixed languages that blend multiple languages within a single sentence. Hinglish, a mix of English and Hindi written in Latin script, is widely seen on platforms like Facebook, YouTube comments, Twitter comments, and Instagram comments. Analysing the sentiment of such code-mixed text is inherently difficult due to irregular grammar, inconsistencies in transliteration, spelling variations, and a lack of annotated datasets. Classical Machine Learning models, such as Word2Vec and FastText embeddings paired with linear classifiers, that provide computational efficiency but often struggle to capture the complex contextual information in code-mixed text as these types of models are usually of Static embeddings technique. Whereas, Deep Learning models, including recurrent neural networks (RNNs), bidirectional long short-term memory networks (BiLSTM), and transformer-based multilingual encoders (MuRIL), are proficient at complex contextual understandings.

## 

## MOTIVATION

## 

This research offers an in-depth comparative study of machine learning and deep learning techniques applied to Hinglish code-mixing for the classification of hate and non-hate speech. It leverages the publicly accessible PRISM dataset, which includes 29950 entries \[16\], for hate-speech detection. The methodology proposed in this study fills the gap in comparing multiple models using various performance metrics such as Accuracy, Balanced Accuracy, Precision, Recall, Specificity, F1 Score, and ROC-AUC Score. The research involves four models that encompass both traditional and deep learning frameworks. The goal is to pinpoint models that exhibit strong performance in binary sentiment classification of Hinglish code-mixed text and to shed light on the benefits of combining different modeling approaches. Additionally, the study aims to evaluate the transformer model without training and to assess the significance of preprocessing.

## OBJECTIVES

## 

To develop and optimize NLP architectures for accurate Hinglish sentiment analysis, begin with preparing a comprehensive annotated dataset and generating multiple embedding representations, including pretrained transformer embeddings (MuRIL, mBART, HingRoBERTa, MPNet) , traditional embeddings combined with BiLSTM (Word2Vec, GloVe, FastText) and embeddings combined with LSTM (Word2Vec, Glove, FastText, Elmo, USE) . Consistent preprocessing steps such as tokenization, transliteration normalization, and language identification are essential. Establish a baseline framework using simpler models like Word2Vec+BiLSTM or GloVe+BiLSTM alongside classical classifiers such as LIGHTGBM to set reference performance metrics. Fine-tune transformer-based models on the Hinglish dataset while optimizing hyperparameters, and similarly train and tune BiLSTM layers with traditional embeddings. For LIGHTGBM, we have used embedding features (like Word2Vec,FastText,Glove,Elmo,USE)Β  as input and optimize boosting the parameters.

Implement a multi-phase language-specific training strategy by splitting the dataset into English, Hindi, and Hinglish subsets, sequentially training and fine-tuning models across these phases with independent validation to monitor performance shifts. Analyze cross-language adaptation by evaluating performance consistency, knowledge transfer, and adaptation capability after each retraining phase. Following multi-phase training, assess robustness and generalization on a unified test corpus, including noisy and informal samples and varied sentiment-emotion categories. Use iterative refinement to adjust architectures, embeddings, and training strategies, considering ensemble approaches that combine transformer LSTM and BiLSTM-based models to enhance accuracy and build a robust Hinglish sentiment-emotion analysis system.

**CHAPTER 2**

**\
LITERATURE SURVEY**

**2.1 RELATED WORK**

Research in code-mixed sentiment analysis, particularly for Hinglish (Hindi–English mixed text), has witnessed substantial growth over the last decade, evolving from simple lexicon-based techniques to sophisticated transformer-driven architectures. Code-mixed text poses unique challenges due to transliteration variations, inconsistent grammar, and a lack of large annotated corpora. Earlier studies primarily relied on lexicons and statistical models, while recent approaches emphasize deep learning and multilingual transformer models that better capture bilingual context and semantics.

The earliest work in this domain explored classical machine learning approaches using handcrafted linguistic and statistical features. Singh analyzed code-mixed social media text using NaΓ―ve Bayes and SVM classifiers, revealing that token-level language identification and transliteration inconsistencies greatly impacted sentiment accuracy \[1\]. Similarly, Thakur et al. outlined that while traditional methods offered moderate accuracy, they lacked scalability and contextual depth \[2\]. These early models were limited in their ability to handle non-standardized language usage and failed to capture more complex syntactic relationships.

A significant shift occurred with the adoption of embedding-based representations that moved beyond sparse lexical features. Techniques such as Word2Vec and GloVe introduced dense vector representations capable of encoding semantic similarity. The study by Agarwal (2024) demonstrated that integrating CNN and BiLSTM with pretrained embeddings significantly improved sentiment detection accuracy for Hinglish text \[3\]. However, these embeddings were static and unable to account for polysemy or word sense variations, leading to limited performance in diverse contexts.

In subsequent years, deep neural architectures like RNNs and LSTMs became prominent for modeling sequential dependencies within sentences. However, these models struggled with long-term contextual understanding and required large labeled corpora for practical training. The introduction of transformer-based architectures revolutionized natural language processing by replacing sequential recurrence with self-attention, enabling parallel processing of long-range dependencies. This innovation paved the way for transformer-based contextual embeddings such as BERT and MuRIL, which excel at handling code-mixed and multilingual data.

Singh et al. presented a study on sentiments in Code-Mixed texts, validating the effectiveness of transformer-based architectures in understanding sentiment polarity and emotion intensity in bilingual texts \[4\]. Similarly, a hybrid attention-based mechanism that outperformed CNN and RNN baselines was proposed, proving that contextualized embeddings substantially improve cross-lingual generalization \[5\]. These findings confirmed that contextual modeling plays a vital role in decoding the semantics of Hinglish text, where literal translations are insufficient for accurate sentiment recognition.

The challenge of data scarcity and domain imbalance in Hinglish corpora was addressed by Yadav et al. (2024), who employed weak supervision and semi-supervised techniques to enhance dataset diversity and reduce annotation cost \[6\]. Similarly, Aggarwal et al. showcased how deep contextual encoders successfully captured sarcasm and implicit sentiment polarity in Hinglish \[7\]. Generally, traditional sentiment models often misclassify or are inefficient at capturing these aspects of sentiment analysis; however, deep learning models achieve satisfactory results. These studies highlight the evolution from sentiment-level to emotion and sarcasm-level understanding in code-mixed research.

The rise of ensemble-based architectures has further pushed the boundaries of Hinglish sentiment and hate-speech classification. Gupta et al. (2021) illustrated that integrating outputs from multiple deep learning and transformer models achieved higher recall and robustness than individual networks \[8\]. Similarly, a combination of several transformer models was deployed to capture the varied contextual nuances and linguistic cues, resulting in improved detection accuracy \[9\]. Moreover, Aloria et al. (2023) further emphasized that attention-based transformers outperform traditional CNN or RNN frameworks, underscoring the superiority of contextual understanding in handling humor, irony, and sarcasm \[10\].

Recent studies have extended Hinglish sentiment analysis to broader tasks, such as multilingual emotion recognition and affective computing. A study used a multilingual transformer pipeline to analyze complex code-mixed expressions, reporting enhanced accuracy through contextual embeddings and domain adaptation \[11\]. Similarly, Baruah et al. developed a BiLSTM-based architecture optimized for mixed-script input, achieving notable improvements in recognizing emotion intensity and polarity \[12\]. Both studies underline that domain-specific pretraining and attention-based architectures significantly improve emotion recognition in low-resource settings.

Furthermore, Paul et al. introduced a sentiment dynamics framework that integrates attention layers to visualize the flow of emotions in bilingual texts \[17\]. This research demonstrated that attention mechanisms not only improve model interpretability but also help localize sentiment-bearing tokens in code-mixed data. Likewise, a review of Code-Mixed Sentiment Analysis shows that a multi-layered CNN-BiLSTM model achieves higher accuracy on benchmark Hinglish datasets by leveraging word embeddings and sentiment lexicons \[18\]. Additionally, recent efforts have focused on addressing linguistic variability and transliteration inconsistencies. A study on challenges in Code-Mixed NLP highlighted the limitations of tokenization, spelling variation, and the representation of Romanized Hindi, underscoring the need for data normalization before model training \[19\]. This work provides an essential foundation for preprocessing strategies in Hinglish NLP pipelines.

From the literature reviewed, it is evident that the field has undergone a clear methodological evolution from feature-engineered machine learning models to embedding-based, deep learning, and transformer-driven architectures. Early models offered interpretability but struggled with the complex semantics of bilingual text. Embedding models improved word-level representation but lacked contextual flexibility. Deep neural networks, such as CNN–BiLSTM, enhanced sequential understanding but required substantial labeled data. In contrast, transformer-based multilingual encoders such as MuRIL deliver superior contextual sensitivity, cross-lingual adaptability, and robustness to transliteration noise. Furthermore, ensemble and hybrid architectures combining these models continue to outperform standalone systems, offering a comprehensive solution to the nuances of Hinglish text processing.

However, despite notable progress, a gap persists in standardized benchmarking and comparative evaluation across models. Hence, the present research aims to bridge this gap by systematically evaluating diverse architectures β€” including MuRIL,Β  CNN–BiLSTM, GloVe–BiLSTM, and Word2Vec–Logistic Regression β€” on a unified Hinglish sentiment dataset to propose a robust multimodel framework for effective sentiment and emotion classification.

**Table 2.1: Literature Survey**

| Sl.NO | Title of the Article | Name of the Author | Name of the Journal | Year of publish | Findings | Research Gap |
|----|----|----|----|----|----|----|
| 1 | Sentiment Analysis of Code-Mixed Social Media Text (Hinglish) | Gaurav Singh\[1\] | arXiv (Preprint) | 2021 | Early Hinglish sentiment analysis; accuracy affected by transliteration inconsistency and token errors. | The necessity for generic sentiment analysis models specifically designed to address the grammatical and lexical intricacies of code-mixed Hinglish material. |
| 2 | Current State ofΒ  Hinglish | Varsha Thakur, Roshani Sahu and Somya Omer\[2\] | SSRN Electronic Journal | Β 2020 | Traditional models showed limited contextual depth and poor scalability | The lack of a thorough assessment and organization of existing approaches, problems, and the latest technological advancements in Hinglish sentiment analysis. |

<table>
<colgroup>
<col style="width: 12%" />
<col style="width: 14%" />
<col style="width: 13%" />
<col style="width: 20%" />
<col style="width: 12%" />
<col style="width: 1%" />
<col style="width: 12%" />
<col style="width: 14%" />
</colgroup>
<thead>
<tr>
<th>Sl.NO</th>
<th>Title of the Article</th>
<th>Name of the Author</th>
<th>Name of the Journal</th>
<th>Year of publish</th>
<th colspan="2">Findings</th>
<th>Research Gap</th>
</tr>
</thead>
<tbody>
<tr>
<td>3</td>
<td>Improving Sentiment Analysis</td>
<td>Prof Neha Agarwal, Viraj Shah , Rishikesh Sharma , Himanshu Yadav, Vaibhav Shah[3]</td>
<td>Educational Administration: Theory and Practice</td>
<td colspan="2">2024</td>
<td>Hybrid deep learning enhanced accuracy compared to classical ML</td>
<td>Current models exhibit constrained accuracy in processing Hinglish; thus, there is a necessity for hybrid deep learning architectures to enhance performance.</td>
</tr>
<tr>
<td>4</td>
<td><h1 id="predicting-multi-label-emojisemotions-and-sentiments-in-code-mixed-texts-using-an-emojifying"><strong>Predicting Multi Label emojis,Emotions, and sentiments in code-mixed Texts using an emojifying</strong></h1></td>
<td>Gopendra Vikram Singh, Soumitra Ghosh, Mauajana Firdaus, Asif Ekbal, Pushpak Bhattacharya[4]</td>
<td>Scientific Reports</td>
<td colspan="2">Β 2024</td>
<td>Multilabel emotion and sentiment analysis improved via transformer-based contextual modeling.</td>
<td>Typical models usually forecast only one label (for instance, sentiment alone). There exists a deficiency in the ability to concurrently predict the interrelated aspects of emojis, various emotions, and sentiments in code-mixed language.</td>
</tr>
</tbody>
</table>

<table>
<colgroup>
<col style="width: 12%" />
<col style="width: 13%" />
<col style="width: 13%" />
<col style="width: 20%" />
<col style="width: 12%" />
<col style="width: 12%" />
<col style="width: 14%" />
</colgroup>
<thead>
<tr>
<th>Sl.NO</th>
<th>Title of the Article</th>
<th>Name of the Author</th>
<th>Name of the Journal</th>
<th>Year of publish</th>
<th>Findings</th>
<th>Research Gap</th>
</tr>
</thead>
<tbody>
<tr>
<td>5</td>
<td><p>A self-Attention hybrid emoji prediction model for code-mixedΒ </p>
<p>language</p></td>
<td>Gadde SatyaΒ  Sai Naga Himabindu,Rajat Rao, Divyasikha Sethiya[5]</td>
<td>Social Network Analysis and Mining</td>
<td>2022</td>
<td>Contextualized embeddings outperformed CNN/RNN; effective in emoji prediction for Hinglish.</td>
<td>Predicting emojis in code-mixed Hinglish with precision is challenging, necessitating specific models that implement self-attention strategies to grasp combined semantic meanings.</td>
</tr>
<tr>
<td>6</td>
<td>Leveraging weakly annotated data for hate speech detection in code-mixed Hinglish: A feasibility-driven transfer learning approach with Large Language Models. In arXiv [cs.CL].</td>
<td>Sargam Yadav, Abishek Kaushik, Kevin McDaid,[6]</td>
<td>arXiv (Preprint)</td>
<td>2024</td>
<td>Used weak supervision to improve dataset diversity and mitigate annotation cost</td>
<td>The critical shortage of well-labeled datasets for Hinglish hate speech requires strategies that utilize poorly annotated data through Large Language Models (LLMs) and transfer learning.</td>
</tr>
</tbody>
</table>

<table style="width:100%;">
<colgroup>
<col style="width: 12%" />
<col style="width: 14%" />
<col style="width: 12%" />
<col style="width: 20%" />
<col style="width: 12%" />
<col style="width: 13%" />
<col style="width: 14%" />
</colgroup>
<thead>
<tr>
<th>Sl.NO</th>
<th>Title of the Article</th>
<th>Name of the Author</th>
<th>Name of the Journal</th>
<th>Year of publish</th>
<th>Findings</th>
<th>Research Gap</th>
</tr>
</thead>
<tbody>
<tr>
<td>7</td>
<td><p>β€œDid you really mean what you said?” :</p>
<p>Sarcasm Detection in Hindi-English Code-Mixed Data using Bilingual Word Embeddings. In arXiv [cs.CL].</p></td>
<td>Akshita Agarwal,Anshul Wadhawan, Ashima Choudhury, Kavita Mourya, [7]</td>
<td><p>Β </p>
<p>arXiv (Preprint)</p></td>
<td>2020</td>
<td>Captured sarcasm and implicit polarity in bilingual data effectively.</td>
<td>Recognition of sarcasm in Hinglish frequently does not succeed with typical monolingual word representations, necessitating the use of tailored bilingual word embeddings to understand irony across languages.</td>
</tr>
<tr>
<td>8</td>
<td>Ensemble based hinglish hate speech detection. 2021 5th International Conference on Intelligent Computing and Control Systems (ICICCS).</td>
<td><p>Β Rahul, Gupta, V., Sehra, V., &amp; Vardhan, Y. R.Β </p>
<p>[8]</p></td>
<td>ICICCS 2021 (Conference)</td>
<td>2021</td>
<td>Ensemble boosted recall and robustness compared to single models.</td>
<td><p>.</p>
<p>Independent models show limited effectiveness in identifying hate speech within noisy Hinglish datasets; collective techniques are necessary to combine their predictive strengths.</p></td>
</tr>
</tbody>
</table>

<table>
<colgroup>
<col style="width: 11%" />
<col style="width: 14%" />
<col style="width: 12%" />
<col style="width: 18%" />
<col style="width: 9%" />
<col style="width: 16%" />
<col style="width: 16%" />
</colgroup>
<thead>
<tr>
<th>Sl.NO</th>
<th>Title of the Article</th>
<th>Name of the Author</th>
<th>Name of the Journal</th>
<th>Year of publish</th>
<th>Findings</th>
<th>Research Gap</th>
</tr>
</thead>
<tbody>
<tr>
<td>9</td>
<td>Ensemble learning-based sarcasm detection in hinglish tweets using Word2Vec embedding. 2025 IEEE International Conference on Interdisciplinary Approaches in Technology and Management for Social Innovation (IATMSI), 1–6.</td>
<td><p>Acharya, A., &amp; Goyal, R.</p>
<p>[9]</p></td>
<td>IATMSI 2025 (Conference)</td>
<td>2025</td>
<td>Combined transformer outputs achieved improved sarcasm detection accuracy.</td>
<td>The requirement to integrate word-level meaning models (Word2Vec) with ensemble machine learning techniques to more effectively understand sarcastic subtleties in Hinglish tweets.</td>
</tr>
<tr>
<td>10</td>
<td>Hilarious or hidden? Detecting sarcasm in hinglish tweets using BERT-GRU. 2023 14th International Conference on Computing Communication and Networking Technologies (ICCCNT)</td>
<td><p>Aloria, S., Aggarwal, I., Baliyan, N., &amp; Ghosh, M.</p>
<p>[10]</p></td>
<td>ICCCNT 2023 (Conference)</td>
<td>Β 2023</td>
<td>Demonstrated attention-based transformer superiority for humor and irony detection</td>
<td>The sequential context and deep semantics of sarcasm in Hinglish are not fully captured by traditional models, requiring advanced hybrid deep learning architectures like BERT-GRU.</td>
</tr>
</tbody>
</table>

<table>
<colgroup>
<col style="width: 10%" />
<col style="width: 16%" />
<col style="width: 11%" />
<col style="width: 15%" />
<col style="width: 10%" />
<col style="width: 17%" />
<col style="width: 17%" />
</colgroup>
<thead>
<tr>
<th>Sl.NO</th>
<th>Title of the Article</th>
<th>Name of the Author</th>
<th>Name of the Journal</th>
<th>Year of publish</th>
<th>Findings</th>
<th>Research Gap</th>
</tr>
</thead>
<tbody>
<tr>
<td>11</td>
<td>Β A comparative study of machine learning and deep learning approaches for identifying Assamese abusive comments on social media. Procedia Computer Science, 258, 981–992.Β </td>
<td><p>Chutia, T., Baruah, N., &amp; Sonowal, P.</p>
<p>[11]</p></td>
<td>Procedia Computer Science</td>
<td>2025</td>
<td>compared traditional ML and DL; BiLSTM outperformed SVM in contextual text understanding.</td>
<td>There is an absence of strict comparative standards between conventional machine learning methods and deep learning approaches specifically aimed at identifying abusive language in Assamese.</td>
</tr>
<tr>
<td>12</td>
<td>Named Entity Recognition in Assamese Language using two separate models: BiLSTM and BERT. Procedia Computer Science, 258, 242–251.</td>
<td><p>Baruah, P., Dutta, B., Sarma, S. K., &amp; Talukdar, K</p>
<p>[12]</p></td>
<td>Procedia Computer Science</td>
<td>2025</td>
<td>Implemented dual-model NER; contributed insights for low-resource Indian languages.</td>
<td>The scarcity of efficient Named Entity Recognition (NER) solutions for the under-resourced Assamese language has led to the investigation of BiLSTM and BERT functionalities.</td>
</tr>
</tbody>
</table>

<table>
<colgroup>
<col style="width: 10%" />
<col style="width: 13%" />
<col style="width: 16%" />
<col style="width: 16%" />
<col style="width: 10%" />
<col style="width: 17%" />
<col style="width: 15%" />
</colgroup>
<thead>
<tr>
<th>Sl.NO</th>
<th>Title of the Article</th>
<th>Name of the Author</th>
<th>Name of the Journal</th>
<th>Year of publish</th>
<th>Findings</th>
<th>Research Gap</th>
</tr>
</thead>
<tbody>
<tr>
<td>13</td>
<td>Β Sentiment analysis of Mizo using lexical features in low resource based models. Natural Language Processing Journal, 13(100181), 100181[13]</td>
<td><p>Lalthangmawii, M., &amp; Singh, T. D.</p>
<p>[13]</p></td>
<td>Natural Language Processing Journal</td>
<td>2025</td>
<td>Proposed lexical sentiment analysis for low-resource Mizo; highlighted linguistic diversity challenges</td>
<td>A significant deficiency of natural language processing tools and sentiment analysis frameworks for the severely low-resource Mizo language exists, necessitating customized models that utilize fundamental lexical characteristics.</td>
</tr>
<tr>
<td>14</td>
<td>Bidirectional LSTM-based sentiment analysis for Assamese text. American Journal of Computer Science and Technology, 7(2), 29–37</td>
<td>Talukdar, M., &amp; Sarma, S. [14]</td>
<td>American Journal of Computer Science and Technology</td>
<td>2024</td>
<td>Achieved strong sequential understanding for sentiment prediction using BiLSTM.</td>
<td>The necessity for models that can grasp two-way sequential context (BiLSTM) to enhance the precision of sentiment polarity classification in Assamese language content.</td>
</tr>
</tbody>
</table>

<table>
<colgroup>
<col style="width: 7%" />
<col style="width: 17%" />
<col style="width: 15%" />
<col style="width: 12%" />
<col style="width: 8%" />
<col style="width: 19%" />
<col style="width: 18%" />
</colgroup>
<thead>
<tr>
<th>Sl.NO</th>
<th>Title of the Article</th>
<th>Name of the Author</th>
<th>Name of the Journal</th>
<th>Year of publish</th>
<th>Findings</th>
<th>Research Gap</th>
</tr>
</thead>
<tbody>
<tr>
<td>15</td>
<td>Β CodemixedNLP: An Extensible and Open NLP Toolkit for Code-Mixing. In arXiv [cs.CL].Β </td>
<td><p>Jayanthi, S. M., Nerella, K., Chandu, K. R., &amp; Black, A. W.Β </p>
<p>[15]</p></td>
<td>arXiv (Preprint)</td>
<td>2021</td>
<td>Introduced open-source NLP toolkit for processing code-mixed languages like Hinglish</td>
<td>The lack of a comprehensive, open-source, and adaptable NLP framework specifically created to manage the preprocessing, modeling, and assessment of languages that are mixed with code.</td>
</tr>
<tr>
<td>16</td>
<td>Code-Mixed Hinglish Hate Speech Detection Dataset. Kaggle.com</td>
<td><p>Dhekane, S.Β </p>
<p>[16]</p></td>
<td>Kaggle (Dataset Repository)</td>
<td>2025</td>
<td>Publicly available dataset enabling research on Hinglish hate-speech classification.</td>
<td>There is a considerable shortage of publicly available, high-quality, and uniform datasets that are essential for training and evaluating Hinglish hate speech detection models.</td>
</tr>
</tbody>
</table>

# 

# 

# 

# 

**CHAPTER 3**

# **METHODOLOGY**

# 

# 3.1 INTRODUCTION

# 

## The overall workflow involves of three proposed methodologies that is illustrated in Fig. 3.1, Fig.3.2, and Fig3.3 which provides a high-level view of the data preprocessing, feature extraction, model training, and evaluation process. It visually summarizes the pipeline from raw dataset input to performance evaluation across mentioned models. The project workflow begins with the primary input which is raw sentiment data of English in Latin text, Hindi in Devnagiri texts and often consisting of code-mixed sentences Hindi-English (Hinglish) in Latin text. This dataset is unstructured and not immediately suitable for Tokenization and further Word Embeddings. Therefore, the first crucial step is Text Pre-processing that involves URL Removal, Lowercasing, Whitespace removal, Tokenization and Embeddings. This Text Pre-processing is critical because deep learning models such as BERT Transformer Models assign different vectors for the same word with different casings, to solve this we use Lowercasing, URLs which do not provide any context so we remove URLs and space normalization as tokens of extra spaces are also created. After preprocessing, the dataset is divided into training (60%), validation (10%), and testing (30%) sets (same for all methodologies). The processed data is then fed into respective sentiment analysis model, where the embedding techniques vary according to the methodology being implemented.

<img src="thesis_media/media/image2.jpg" style="width:5.62651in;height:4.21988in" />

**Fig 3.1: Β Workflow of the proposed methodology for sentiment analysis.**

<span id="_TOC_250032" class="anchor"></span>After preprocessing and dataset splitting, the training set is used to develop multiple hybrid sentiment classification models by combining various word embedding techniques with machine learning and deep learning algorithms. Specifically, Word2Vec, GloVe, FastText, Universal Sentence Encoder (USE), and ELMo embeddings are integrated with LSTM and LightGBM classifiers to capture the semantic and contextual information present in code-mixed text. The validation set is used for model tuning and performance optimization, while the test set is employed for final evaluation. The predicted sentiment labels generated by these embedding-classifier combinations are then analyzed to compare their effectiveness and identify the best-performing model for code-mixed sentiment analysis.

## 

<img src="thesis_media/media/image3.jpg" style="width:6.26667in;height:4.7in" />

**Fig 3.2: Β Workflow of the proposed methodology for sentiment analysis.**

In the above proposed methodology Fig3.2: Two training strategies are applied. In Strategy-1, each language corpus and the combined dataset are independently trained and evaluated on the test set. In Strategy-2, multi-stage language training is performed sequentially using English, Hinglish, Hindi, and combined corpora to improve cross-lingual understanding. Finally, predicted sentiment labels for code-mixed texts are compared through result analysis to evaluate model effectiveness and across multilingual strategies.

<img src="thesis_media/media/image4.jpg" style="width:6.37751in;height:4.78313in" />

**Fig 3.3: Β Workflow of the proposed methodology for sentiment analysis.**

In the proposed methodology fig3, we performed multi-stage language training that consists of six various strategies of training done sequentially using (English, Hinglish, Hindi, all combined), (English, Hindi, Hinglish, all combined ), (Hinglish, English, Hindi, all combined), (Hinglish, Hindi, English, all combined), (Hindi, Hinglish, English, all combined), and (Hindi, English, Hinglish, all combined) corpora to improve cross-lingual understanding. Finally, predicted sentiment labels for code-mixed texts are compared through result analysis to evaluate model effectiveness and across multilingual strategies.

## 3.3 DATA COLLECTION

## 

The research utilized the "combined_hate_speech_dataset" that is publically available on Kaggle. This dataset includes 29,550 labeled text entries, predominantly featuring code-mixed Hinglish sentences. For the purpose of analysis, two main columns were preserved: the text content and the hate_label. The target variable is binary, with 0 indicating non-hate content and 1 signifying hate content. The dataset shows a slight imbalance in class distribution, with non-hate samples being more prevalent. The following table clearly provides data description.

Table 3.1: Dataset Details

<table style="width:89%;">
<colgroup>
<col style="width: 25%" />
<col style="width: 63%" />
</colgroup>
<thead>
<tr>
<th style="text-align: left;"><p>Β </p>
<p><strong>Dataset Name</strong></p></th>
<th style="text-align: left;"><strong>combined_hate_speech_dataset (PRISM) – Kaggle</strong></th>
</tr>
</thead>
<tbody>
<tr>
<td style="text-align: left;">Total Samples</td>
<td style="text-align: left;">29,550</td>
</tr>
<tr>
<td style="text-align: left;">Classification Type</td>
<td style="text-align: left;">Binary (0 = Non-hate, 1 = Hate)</td>
</tr>
<tr>
<td style="text-align: left;">Languages Covered</td>
<td style="text-align: left;">English, Hindi (Devanagari), Hinglish</td>
</tr>
<tr>
<td style="text-align: left;">Data Composition</td>
<td style="text-align: left;">English: 15,000</td>
</tr>
<tr>
<td style="text-align: left;">Β </td>
<td style="text-align: left;">Hindi: 9,767</td>
</tr>
<tr>
<td style="text-align: left;">Β </td>
<td style="text-align: left;">Hinglish: 4,783</td>
</tr>
<tr>
<td style="text-align: left;">Label Distribution</td>
<td style="text-align: left;">Non-hate: 15,825</td>
</tr>
<tr>
<td style="text-align: left;">Β </td>
<td style="text-align: left;">Hate: 13,725</td>
</tr>
<tr>
<td style="text-align: left;">Profanity Lexicon</td>
<td style="text-align: left;">209 offensive terms with severity scores</td>
</tr>
<tr>
<td style="text-align: left;">Application</td>
<td style="text-align: left;">Hate-speech / Toxicity detection in multilingual text</td>
</tr>
</tbody>
</table>

## 3.4 DATA PREPROCESSING

The dataset originally consisted of 29,550 entries, with 9 attributes such as text, hate_label, source, profanity_score, language, dataset_version, combined_date, text_length and word_count. This dataset included samples in English (Latin Script), Hindi(Devnagiri Script), and Hinglish(Latin Script), for the purpose of classifying hate speech. To enhance the quality of the data, thorough preprocessing and noise analysis were conducted prior to tokenization and embedding. During the analysis, it was found that duplicate rows, repeated texts, URLs, mentions, hashtags, elongated words, and social media attachments were significant sources of noise. Regex-based techniques were employed to identify and eliminate these noisy elements. Further preprocessing involved converting text to lowercase, removing extra spaces, normalizing elongated words, and deleting URLs and HTML entities. The following table shows the general noise statistics that are found during preprocessing.

Table 3.2: List Count of URL Noise

| **URL / Social Noise Type** | **Count** |
|:----------------------------|:----------|
| HTTP/HTTPS Links            | 356       |
| WWW Links                   | 7         |
| pic.twitter Links           | 106       |
| YouTube Links               | 22        |
| Facebook Links              | 0         |
| Hungama Links               | 1         |
| Attached URLs               | 33        |

The preprocessing stage significantly improved the overall quality of the dataset by removing noisy and redundant textual patterns. Duplicate text entries, URLs, and elongated words were identified as major sources of inconsistency that could negatively affect model learning and classification performance. Cleaning operations such as duplicate removal, URL elimination, and text normalization helped create a more consistent and standardized corpus for embedding. As a result, the dataset size was slightly reduced while preserving meaningful information required for hate speech detection. The following table shows the comparison of the dataset before and after cleaning.

Table 3.3: Data Metrics Before and After Cleaning

| **Metric**                 | **Before Cleaning** | **After Cleaning** |
|:---------------------------|:--------------------|:-------------------|
| Total Rows                 | 29,550              | 29,506             |
| Duplicate Texts            | 11                  | 0                  |
| Texts with URLs            | 458                 | 0                  |
| Texts with Elongated Words | 2,176               | 0                  |

After preprocessing, essential features were extracted to prepare the dataset for hate speech classification. The refined dataset retained only relevant attributes required for model training and analysis. The clean_text feature contains normalized textual content, while hate_label represents the target classification variable. The language feature identifies the language category of each sample, and additional statistical features such as text_length and word_count were included to capture textual characteristics. These extracted features help improve data representation and support effective downstream modeling. The final processed dataset consisted of 29,506 samples with 5 important features.

## 3.5 EDA

<img src="thesis_media/media/image5.png" style="width:3.66234in;height:3.45649in" />

Fig:3.4 Hate vs Non-Hate Class distribution

The following pie chart illustrates the distribution of hate and non-hate samples in the PRISM dataset. The dataset contains approximately balanced class representations, where non-hate samples account for **53.5%** and hate samples account for **46.5%** of the total data. Maintaining a balanced class distribution is important for reducing model bias and improving classification performance.

<img src="thesis_media/media/image6.png" style="width:3.64935in;height:3.60256in" />

Fig:3.5 Language Distribution

The following pie chart presents the language distribution of the dataset across English, Hindi, and Hinglish corpora. English samples constitute **50.8%**, Hindi samples represent **33.0%**, and Hinglish samples contribute **16.2%** of the dataset. This multilingual distribution enables the models to learn diverse linguistic patterns for multilingual hate speech detection.

<img src="thesis_media/media/image7.png" style="width:4.27923in;height:3.25974in" />

Fig:3.6 Feature Correlation Heatmap

The correlation heatmap illustrates the relationship among numerical features such as hate_label, text_length, and word_count. A very strong positive correlation (0.99) is observed between text length and word count, indicating that longer texts generally contain more words. In contrast, hate_label shows a weak negative correlation with both text length and word count, suggesting that text size has minimal influence on hate speech classification.

<img src="thesis_media/media/image8.png" style="width:3.19849in;height:2.51948in" /><img src="thesis_media/media/image9.png" style="width:3.00752in;height:2.58956in" />

Fig:3.7 Text length Statistics and Word count Statisitics

The following figures present the statistical distribution of text length and word count in the final cleaned dataset. The text length analysis shows a mean of 150.52 characters and a median of 94 characters, with some samples reaching a maximum length of 1926 characters. Similarly, the word count distribution shows an average of 28.38 words and a median of 18 words, while the maximum word count reaches 300 words. The 95th percentile values of 480 characters and 90 words indicate the presence of long textual samples and high variance within the dataset. These observations are important for determining appropriate padding and truncation limits in transformer-based models to ensure efficient training and balanced sequence representation.

## 3.6 DATASET SPLITTING

The cleaned PRISM dataset was analyzed to understand the distribution of hate labels and language categories before model training. The dataset contains English, Hindi, and Hinglish samples with a relatively balanced hate speech distribution across languages. To ensure reliable model evaluation, the dataset was divided into training, validation, and testing subsets using a stratified splitting approach. Initially, 70% of the data was reserved as the training pool and 30% as the independent test set. The training pool was further divided into 60% training data and 10% validation data. Only the clean_text and hate_label columns were used during model training. Language-wise class distributions were maintained across all subsets to preserve dataset balance and reduce sampling bias.

The following table shows the language wise data splitting.

Table 3.4: Dataset Splits Table

<table>
<colgroup>
<col style="width: 15%" />
<col style="width: 9%" />
<col style="width: 10%" />
<col style="width: 15%" />
<col style="width: 9%" />
<col style="width: 8%" />
<col style="width: 7%" />
<col style="width: 7%" />
<col style="width: 7%" />
<col style="width: 7%" />
</colgroup>
<thead>
<tr>
<th colspan="3" style="text-align: left;"><p>Β </p>
<p><strong>Category</strong></p></th>
<th style="text-align: left;"><strong>Combined</strong></th>
<th colspan="2" style="text-align: left;"><strong>English</strong></th>
<th colspan="2" style="text-align: left;"><strong>Hindi</strong></th>
<th colspan="2" style="text-align: left;"><strong>Hinglish</strong></th>
</tr>
</thead>
<tbody>
<tr>
<td colspan="3" style="text-align: left;"><strong>Total Samples</strong></td>
<td style="text-align: left;">29506</td>
<td colspan="2" style="text-align: left;">14994</td>
<td colspan="2" style="text-align: left;">9738</td>
<td colspan="2" style="text-align: left;">4774</td>
</tr>
<tr>
<td rowspan="2" style="text-align: left;"><strong>Label Distribution</strong></td>
<td colspan="2" style="text-align: left;">Non-hate(0)</td>
<td style="text-align: left;">15,799</td>
<td colspan="2" style="text-align: left;">7495</td>
<td colspan="2" style="text-align: left;">5393</td>
<td colspan="2" style="text-align: left;">2911</td>
</tr>
<tr>
<td colspan="2" style="text-align: left;">Hate(1)</td>
<td style="text-align: left;">13,707</td>
<td colspan="2" style="text-align: left;">7499</td>
<td colspan="2" style="text-align: left;">4345</td>
<td colspan="2" style="text-align: left;">1863</td>
</tr>
<tr>
<td rowspan="6" style="text-align: left;"><strong>Split</strong></td>
<td rowspan="2" style="text-align: left;"><p><strong>Train</strong></p>
<p><strong>(60%)</strong></p></td>
<td style="text-align: left;">Non-hate(0)</td>
<td rowspan="2" style="text-align: left;">18617</td>
<td style="text-align: left;">4478</td>
<td rowspan="2" style="text-align: left;">9446</td>
<td style="text-align: left;">3237</td>
<td rowspan="2" style="text-align: left;">6143</td>
<td style="text-align: left;">1764</td>
<td rowspan="2" style="text-align: left;">3011</td>
</tr>
<tr>
<td style="text-align: left;">Hate(1)</td>
<td style="text-align: left;">4485</td>
<td style="text-align: left;">2622</td>
<td style="text-align: left;">1117</td>
</tr>
<tr>
<td rowspan="2" style="text-align: left;"><p><strong>Val</strong></p>
<p><strong>(10%)</strong></p></td>
<td style="text-align: left;">Non-hate(0)</td>
<td rowspan="2" style="text-align: left;">2086</td>
<td style="text-align: left;">525</td>
<td rowspan="2" style="text-align: left;">1050</td>
<td style="text-align: left;">378</td>
<td rowspan="2" style="text-align: left;">683</td>
<td style="text-align: left;">204</td>
<td rowspan="2" style="text-align: left;">335</td>
</tr>
<tr>
<td style="text-align: left;">Hate(1)</td>
<td style="text-align: left;">525</td>
<td style="text-align: left;">305</td>
<td style="text-align: left;">131</td>
</tr>
<tr>
<td rowspan="2" style="text-align: left;"><p><strong>Test</strong></p>
<p><strong>(30%)</strong></p></td>
<td style="text-align: left;">Non-hate(0)</td>
<td rowspan="2" style="text-align: left;">8865</td>
<td style="text-align: left;">2297</td>
<td rowspan="2" style="text-align: left;">4499</td>
<td style="text-align: left;">1597</td>
<td rowspan="2" style="text-align: left;">2926</td>
<td style="text-align: left;">846</td>
<td rowspan="2" style="text-align: left;">1434</td>
</tr>
<tr>
<td style="text-align: left;">Hate(1)</td>
<td style="text-align: left;">2248</td>
<td style="text-align: left;">1303</td>
<td style="text-align: left;">561</td>
</tr>
</tbody>
</table>

## 

## 3.7 MODEL SELECTION 

##  

##  3.7.1 MODEL DESCRIPTION

> 3.7.1.1 TRANSFORMER BASED MODELS

1.  **MuRIL:** MuRIL is a multilingual transformer model developed by Google for Indian languages and code-mixed text. It generates contextual embeddings that capture semantic relationships across multiple languages. MuRIL was selected because the dataset contains Hindi, English, and Hinglish text, making it highly effective for multilingual hate speech classification.

2.  **mBART:** mBART is a multilingual encoder–decoder transformer model designed for cross-lingual understanding and contextual language representation. It was used because of its strong capability to learn multilingual semantic patterns from diverse textual inputs.

3.  **HingRoBERTa:** HingRoBERTa is a RoBERTa-based transformer specifically adapted for Hinglish and code-mixed language processing. It was selected because it effectively handles transliterated and mixed-language text commonly found in social media hate speech datasets.

4.  **MPNet:** MPNet combines masked language modeling with permuted positional encoding to improve contextual understanding. It was chosen because of its strong sentence representation capability and effectiveness in capturing complex hate speech semantics.

> 3.7.1.2 DEEP LEARNING MODELS

5.  **FastText + BiLSTM:** This model combines FastText embeddings with a Bidirectional LSTM network. FastText captures subword information and spelling variations, while BiLSTM learns contextual dependencies from both forward and backward directions. It was selected because Hinglish and social media text often contain noisy and misspelled words.

6.  **Word2Vec + BiLSTM:** Word2Vec provides semantic word embeddings, and BiLSTM captures sequential contextual information from text. This model was used to evaluate the effectiveness of predictive word embeddings for multilingual hate speech detection.

7.  **GloVe + BiLSTM:** GloVe embeddings capture global word co-occurrence information, while BiLSTM models sequential text patterns. This architecture was selected to compare static embedding-based contextual learning against transformer models.

8.  **Word2Vec + LSTM:** This model combines Word2Vec embeddings with LSTM networks for sequence learning. It was selected because LSTM effectively captures long-term dependencies in textual data.

9.  **GloVe + LSTM:** GloVe embeddings with LSTM were used to analyze the effectiveness of global semantic representations in hate speech classification tasks.

10. **FastText + LSTM:** FastText embeddings combined with LSTM were selected because FastText handles subword-level variations effectively, which is useful for multilingual and code-mixed text.

11. **USE+LSTM:** This model uses Universal Sentence Encoder embeddings with LSTM networks. USE captures sentence-level semantic meaning, making it useful for understanding contextual hate speech patterns.

12. **ELMo+LSTM:** ELMo in conjunction with LSTM integrates advanced contextualized word representations with a sequential deep learning architecture, which enhances its effectiveness in grasping intricate syntactic and semantic frameworks throughout a text.

> 3.7.1.3 MACHINE LEARNING MODELS

13. **Word2Vec + LightGBM:** This model combines Word2Vec embeddings with the LightGBM classifier. It was selected because LightGBM efficiently handles vectorized textual features using gradient boosting techniques.

14. **GloVe + LightGBM:** GloVe embeddings were used as feature inputs for LightGBM to evaluate how global semantic vectors perform with boosting-based classification.

15. **FastText + LightGBM:** FastText embeddings combined with LightGBM were selected because FastText captures subword semantics effectively while LightGBM provides efficient classification performance.

16. **USE + LightGBM:** This model uses USE sentence embeddings with LightGBM classification. It was selected to evaluate sentence-level semantic representations using boosting methods.

17. **ELMo+ LightGBM:** ELMo combined with LightGBM derives constant contextual feature vectors from text through ELMo and feeds them into a highly efficient, gradient-boosted decision tree classifier, providing a resource-saving method for classification tasks that resemble tabular data.

## 3.7.2 MODEL ALGORITHMS

The selected models use supervised learning for hate speech classification. Transformer models employ fine-tuning with contextual embeddings and attention mechanisms for sequence understanding. BiLSTM and LSTM architectures process text sequentially to capture contextual dependencies from embeddings. LightGBM models use gradient boosting decision trees for efficient feature-based classification, while Logistic Regression applies a linear decision boundary over Word2Vec embeddings. Cross-Entropy and Binary Cross-Entropy losses were used for binary classification tasks, and optimizers such as Adam, AdamW, and Gradient Boosting strategies were applied for stable convergence and improved learning performance.

Table 3.5: Model Algorithms

<table>
<colgroup>
<col style="width: 21%" />
<col style="width: 22%" />
<col style="width: 24%" />
<col style="width: 12%" />
<col style="width: 19%" />
</colgroup>
<thead>
<tr>
<th><p>Β </p>
<p><strong>Model</strong></p></th>
<th><strong>Core Concept</strong></th>
<th><strong>Training Method</strong></th>
<th><strong>Loss function</strong></th>
<th><strong>Optimizer</strong></th>
</tr>
</thead>
<tbody>
<tr>
<td>MuRIL</td>
<td>Transformer-based model, contextual multilingual embeddings</td>
<td>Supervised sequence fine-tuning + linear warmup–decay scheduler</td>
<td>Cross-Entropy Loss</td>
<td>AdamW</td>
</tr>
<tr>
<td>mBART</td>
<td>Transformer-based encoder–decoder model with fully contextual multilingual embeddings</td>
<td>Supervised sequence fine-tuning with linear warmup–decay scheduler</td>
<td>Cross-Entropy Loss</td>
<td>Adam Optimizer</td>
</tr>
<tr>
<td>HingRoBERTa</td>
<td>Transformer-based model (RoBERTa variant), contextual embeddings adapted for Hinglish/code-mixed text</td>
<td>Supervised sequence fine-tuning + linear warmup–decay scheduler</td>
<td>Cross-Entropy Loss</td>
<td>AdamW</td>
</tr>
<tr>
<td><p>Β </p>
<p>MPNet</p></td>
<td>Transformer-based model using masked language modeling with permuted position encoding for enhanced contextual representation</td>
<td>Supervised sequence fine-tuning + linear warmup–decay scheduler</td>
<td>Cross-Entropy Loss</td>
<td>AdamW</td>
</tr>
<tr>
<td>FastText + BiLSTM</td>
<td>Subword (character n-gram) embeddings + Bidirectional LSTM contextual modeling</td>
<td>Supervised mini-batch training with backpropagation</td>
<td>Cross-Entropy Loss</td>
<td>Adam Optimizer</td>
</tr>
<tr>
<td>Word2Vec + BiLSTM</td>
<td>Predictive word embeddings (CBOW/Skip-gram) + Bidirectional LSTM contextual modeling</td>
<td>Supervised mini-batch training with backpropagation</td>
<td>Cross-Entropy Loss</td>
<td>Adam Optimizer</td>
</tr>
<tr>
<td>GloVe + BiLSTM</td>
<td>Static global co-occurrence embeddings + Bidirectional LSTM contextual modeling</td>
<td>Supervised mini-batch training with backpropagation</td>
<td>Cross-Entropy Loss</td>
<td>Adam Optimizer</td>
</tr>
<tr>
<td><p>Β </p>
<p>Word2Vec + LSTM</p></td>
<td>Uses Word2Vec embeddings (semantic similarity) + LSTM to capture sequential dependencies in text</td>
<td>Pretrained Word2Vec embeddings fed into LSTM, trained end-to-end on labeled data</td>
<td>Binary Cross-Entropy</td>
<td>Adam</td>
</tr>
<tr>
<td>GloVe + LSTM</td>
<td>Uses GloVe embeddings (global word co-occurrence statistics) + LSTM for sequence learning</td>
<td>Pretrained GloVe vectors used as embedding layer, then LSTM training on dataset</td>
<td>Binary Cross-Entropy</td>
<td>Adam</td>
</tr>
<tr>
<td>FastText + LSTM</td>
<td>FastText captures subword information (handles misspellings, Hinglish variations) + LSTM</td>
<td>Pretrained FastText embeddings β†’ LSTM trained on sequence data</td>
<td>Binary Cross-Entropy</td>
<td>Adam</td>
</tr>
<tr>
<td><p>Β </p>
<p>USE + LSTM</p></td>
<td>Universal Sentence Encoder provides sentence-level embeddings + LSTM for deeper sequence modeling</td>
<td>USE embeddings generated β†’ passed to LSTM for classification training</td>
<td>Binary Cross-Entropy</td>
<td>Adam</td>
</tr>
<tr>
<td>Word2Vec + LightGBM</td>
<td>Gradient Boosting decision trees + Word2Vec feature vectors</td>
<td>Word2Vec embeddings averaged or pooled β†’ fed into LightGBM classifier</td>
<td>Binary Log Loss</td>
<td>Gradient Boosting Decision Trees</td>
</tr>
<tr>
<td>GloVe + LightGBM</td>
<td>GloVe embeddings used as input features + LightGBM for classification</td>
<td>GloVe vectors aggregated β†’ used to train LightGBM model</td>
<td>Binary Log Loss</td>
<td><p>Stochastic Gradient Descent</p>
<p>(+ Negative Sampling)</p></td>
</tr>
<tr>
<td>Β FastText + LightGBM</td>
<td>FastText embeddings (handles subwords well) + LightGBM classifier</td>
<td>FastText vectors β†’ feature input β†’ LightGBM training</td>
<td>Binary Log Loss</td>
<td><p>Stochastic Gradient Descent</p>
<p>(+ Negative Sampling + Subword learning)</p></td>
</tr>
<tr>
<td>USE + LightGBM</td>
<td>Sentence-level embeddings from USE + LightGBM classification</td>
<td>USE embeddings β†’ directly fed to LightGBM</td>
<td>Binary Log Loss</td>
<td>AdaGrad</td>
</tr>
<tr>
<td>ELMo+ LSTM</td>
<td>ELMo gives GBDT</td>
<td>Supervised DL (hybrid )</td>
<td><p>Binary Cross</p>
<p>entropy</p></td>
<td>SGD</td>
</tr>
<tr>
<td>ELMo+ LightGBM</td>
<td>Word2Vec embeddings + Logistic Regression classifier</td>
<td>Supervised ML (linear model)</td>
<td>Binary Log Loss</td>
<td>Gradient-Based One-Side Sampling (GOSS) and Exclusive Feature Bundling (EFB)</td>
</tr>
</tbody>
</table>

## 3.7.3 MODEL HYPERPARAMETERS

Hyperparameters were selected to balance computational efficiency and model performance. For Transformer models we set max sequence length = 128 to capture sufficient contextual information while maintaining manageable memory usage. A learning rate of 2e-5 with AdamW optimizer and weight decay of 0.01 was used to ensure stable fine-tuning and prevent overfitting. For BiLSTM and LSTM-based models, embedding dimensions of 100–300 and hidden dimensions up to 256 were chosen to learn rich semantic representations. Dropout values between 0.2–0.5 were applied to reduce overfitting. Batch sizes of 16 and 32 were selected for balanced training stability and GPU utilization. LightGBM hyperparameters such as max_depth, num_leaves, and n_estimators were tuned to improve classification performance while controlling model complexity. Lower learning rates and regularization parameters were used to achieve stable gradient updates and better generalization.

Table 3.6: Model Hyperparameters

<table>
<colgroup>
<col style="width: 18%" />
<col style="width: 20%" />
<col style="width: 23%" />
<col style="width: 17%" />
<col style="width: 10%" />
<col style="width: 9%" />
</colgroup>
<thead>
<tr>
<th><p>Β </p>
<p><strong>Model</strong></p></th>
<th><strong>Core Components / Embeddings</strong></th>
<th><strong>Key Hyperparameters</strong></th>
<th><strong>Optimizer</strong></th>
<th><strong>Epochs</strong></th>
<th><strong>Batch Size</strong></th>
</tr>
</thead>
<tbody>
<tr>
<td><strong>MuRIL</strong></td>
<td>Transformer (google/muril-base-cased)</td>
<td>max_seq_len=128, weight_decay=0.01, warmup_ratio=0.1</td>
<td>AdamW, lr=2e-5</td>
<td>8</td>
<td>16</td>
</tr>
<tr>
<td><strong>mBART</strong></td>
<td>Transformer (facebook/mbart-large-50)</td>
<td>max_seq_len = 128, weight_decay = 0.01, warmup_ratio = 0.1</td>
<td>AdamW, lr = 2e-5</td>
<td>8</td>
<td>16</td>
</tr>
<tr>
<td><strong>HingRoBERTa</strong></td>
<td>Transformer (RoBERTa-based Hinglish-adapted model)</td>
<td>max_seq_len = 128, weight_decay = 0.01, warmup_ratio = 0.1</td>
<td>AdamW, lr = 2e-5</td>
<td>8</td>
<td>16</td>
</tr>
<tr>
<td><strong>MPNet</strong></td>
<td>Transformer (microsoft/mpnet-base)</td>
<td>max_seq_len = 128, weight_decay = 0.01, warmup_ratio = 0.1</td>
<td>AdamW, lr = 2e-5</td>
<td>8</td>
<td>16</td>
</tr>
<tr>
<td><p><strong>Β </strong></p>
<p><strong>FastText + BiLSTM</strong></p></td>
<td>BiLSTM + Pre-trained FastText (subword) embeddings</td>
<td>max_seq_len = 128, embedding_dim = 300, hidden_dim = 256, dropout = 0.5</td>
<td>Adam, lr = 1e-3, weight_decay = 0</td>
<td>10</td>
<td>32</td>
</tr>
<tr>
<td><strong>Word2Vec + BiLSTM</strong></td>
<td>BiLSTM + Pre-trained Word2Vec embeddings</td>
<td>max_seq_len = 128, embedding_dim = 300, hidden_dim = 256, dropout = 0.5</td>
<td>Adam, lr = 1e-3, weight_decay = 0</td>
<td>10</td>
<td>32</td>
</tr>
<tr>
<td><strong>GloVe + BiLSTM</strong></td>
<td>BiLSTM + Pre-trained GloVe embeddings</td>
<td>max_seq_len = 128, embedding_dim = 300, hidden_dim = 256, dropout = 0.5</td>
<td>Adam, lr = 1e-3, weight_decay = 0</td>
<td>10Β </td>
<td>32</td>
</tr>
<tr>
<td><p><strong>Β </strong></p>
<p><strong>LSTMΒ </strong></p></td>
<td>Word2vec</td>
<td>max_seq_len =100 , embedding_dim =100 , hidden_dim = 128, dropout = 0.2</td>
<td>Adam, lr = 1e-2, weight_decay = 0</td>
<td>30</td>
<td>16</td>
</tr>
<tr>
<td><strong>LSTMΒ </strong></td>
<td>GLOVEΒ </td>
<td>max_seq_len =100 , embedding_dim =100 , hidden_dim = 128, dropout = 0.2</td>
<td>Adam, lr = 1e-2, weight_decay = 0</td>
<td>30</td>
<td>16</td>
</tr>
<tr>
<td><strong>LSTMΒ </strong></td>
<td>FASTTEXTΒ </td>
<td>max_seq_len =100 , embedding_dim =100 , hidden_dim = 128, dropout = 0.2</td>
<td>Adam, lr = 1e-2, weight_decay = 0</td>
<td>30</td>
<td>16</td>
</tr>
<tr>
<td><p><strong>Β </strong></p>
<p><strong>LSTMΒ </strong></p></td>
<td>USE</td>
<td>max_seq_len =100 , embedding_dim =100 , hidden_dim = 128, dropout = 0.2</td>
<td>Adam, lr = 1e-2, weight_decay = 0</td>
<td>30</td>
<td>16</td>
</tr>
<tr>
<td><strong>LightGBM</strong></td>
<td>Word2vec</td>
<td><p>learning_rate = 0.01</p>
<p>n_estimators = 500</p>
<p>max_depth = 8</p>
<p>num_leaves = 63</p>
<p>subsample = 0.8</p>
<p>colsample_bytree = 0.8</p>
<p>reg_alpha = 0.1</p>
<p>reg_lambda = 0.2</p></td>
<td>Gradient Boosting Decision Trees</td>
<td>20</td>
<td>16</td>
</tr>
<tr>
<td><strong>LightGBM</strong></td>
<td>GLOVEΒ </td>
<td><p>learning_rate = 0.01</p>
<p>n_estimators = 500</p>
<p>max_depth = 8</p>
<p>num_leaves = 63</p>
<p>subsample = 0.8</p>
<p>colsample_bytree = 0.8</p>
<p>reg_alpha = 0.1</p>
<p>reg_lambda = 0.2</p></td>
<td><p>Stochastic Gradient Descent</p>
<p>(+ Negative Sampling)</p></td>
<td>20</td>
<td>16</td>
</tr>
<tr>
<td><p><strong>Β </strong></p>
<p><strong>LightGBM</strong></p></td>
<td>FASTTEXTΒ </td>
<td><p>learning_rate = 0.01</p>
<p>n_estimators = 500</p>
<p>max_depth = 8</p>
<p>num_leaves = 63</p>
<p>subsample = 0.8</p>
<p>colsample_bytree = 0.8</p>
<p>reg_alpha = 0.1</p>
<p>reg_lambda = 0.2</p></td>
<td><p>Stochastic Gradient Descent</p>
<p>(+ Negative Sampling + Subword learning)</p></td>
<td>20</td>
<td>16</td>
</tr>
<tr>
<td><strong>LightGBM</strong></td>
<td>USE</td>
<td><p>learning_rate = 0.01</p>
<p>n_estimators = 500</p>
<p>max_depth = 8</p>
<p>num_leaves = 63</p>
<p>subsample = 0.8</p>
<p>colsample_bytree = 0.8</p>
<p>reg_alpha = 0.1</p>
<p>reg_lambda = 0.2</p></td>
<td>AdaGrad</td>
<td>20</td>
<td>16</td>
</tr>
<tr>
<td><strong>LightGBM</strong></td>
<td>ELMo</td>
<td><p>learning_rate = 0.01</p>
<p>n_estimators = 500</p>
<p>max_depth = 8</p>
<p>num_leaves = 63</p>
<p>subsample = 0.8</p>
<p>colsample_bytree = 0.8</p>
<p>reg_alpha = 0.1</p>
<p>reg_lambda = 0.2</p></td>
<td>Gradient-Based One-Side Sampling (GOSS) and Exclusive Feature Bundling (EFB)</td>
<td>30</td>
<td>16</td>
</tr>
<tr>
<td><p><strong>Β </strong></p>
<p><strong>LSTM</strong></p></td>
<td>ELMo</td>
<td>max_seq_len =200 , embedding_dim =200 , hidden_dim = 128, dropout = 0.2</td>
<td>SGD (Stochastic Gradient Descent)., lr = 1e-2, weight_decay = 0</td>
<td>30</td>
<td>16</td>
</tr>
</tbody>
</table>

## 

## **3.6 EVALUATION METRICS**

## 

## As the Dataset consists of two classes i.e, 0 (non-hate) and 1 (hate) The performance of the proposed system was rigorously evaluated using 7 well-established metrics that includes Accuracy, Balanced Accuracy, Precision, Recall, Specificity, F1 and AUC-ROC.

**Accuracy:** Accuracy measures the overall percentage of correctly classified samples among the total predictions. It evaluates how well the model performs on both hate and non-hate classes collectively. Accuracy was used to measure the general classification performance of the models.

``` math
Accuracy = \frac{TP + TN}{TP + TN + FP + FN}
```

**Balanced Accuracy:** Balanced Accuracy computes the average recall obtained for each class and is useful for handling class imbalance. It ensures that both hate and non-hate classes contribute equally to evaluation. This metric was used to provide unbiased performance measurement across classes.

``` math
Balanced\, Accuracy = \frac{Recall\  + Specificity}{2}
```

**Precision:** Precision measures the proportion of correctly predicted hate samples among all samples predicted as hate. It evaluates the model’s ability to reduce false positive predictions. Precision was used because false hate predictions can negatively affect classification reliability.

``` math
Precision = \frac{TP}{TP + FP}
```

**Recall (Sensitivity):** Recall measures the proportion of actual hate samples correctly identified by the model. It evaluates the model’s ability to detect hate speech effectively. Recall was important because missing harmful content may reduce system effectiveness.

``` math
Recall = \frac{TP}{TP + FN}
```

**Specificity:** Specificity measures the proportion of correctly identified non-hate samples. It evaluates how effectively the model avoids false hate predictions for normal text. This metric was used to ensure balanced non-hate classification performance.

``` math
Specificity = \frac{TN}{TN + FP}
```

**F1-Score:** F1-Score is the harmonic mean of Precision and Recall. It provides a balanced evaluation when both false positives and false negatives are important. F1-score was used because hate speech datasets often require balanced detection capability.

``` math
F1 = \frac{2 \times Precision \times Recall}{Precision + Recall}
```

**AUC-ROC:** AUC-ROC measures the model’s ability to distinguish between hate and non-hate classes across different classification thresholds. Higher AUC values indicate better discrimination capability. This metric was used to evaluate overall classification robustness and threshold-independent performance.

3.7 TOOLS AND FRAMEWORKS

The development and evaluation of the Hindi-English code-mixed involved a combination of tools, libraries, and frameworks from both speech processing and natural language processing domains. The following are the major tools and frameworks utilized throughout the project:

1.  Python was used as the primary programming language for dataset preprocessing, model implementation, training, and evaluation.

2.  Google Colab was used for executing experiments with GPU support and cloud-based computation.

3.  Jupyter Notebook was used for interactive coding, experimentation, and result visualization.

4.  Pandas was used for data loading, preprocessing, cleaning, and tabular data manipulation.

5.  NumPy was used for numerical computations and array-based operations.

6.  Regex was used for detecting and removing URLs, mentions, hashtags, and noisy textual patterns.

7.  NLTK was used for tokenization and text preprocessing operations.

8.  Scikit-learn was used for dataset splitting, evaluation metrics, and machine learning utilities.

9.  PyTorch was used for implementing transformer-based and deep learning models.

10. TensorFlow and Keras were used for implementing LSTM, BiLSTM, and neural network architectures.

11. Transformers was used for loading and fine-tuning transformer models such as MuRIL, mBART, MPNet, and HingRoBERTa.

12. Gensim was used for generating Word2Vec and FastText embeddings.

13. LightGBM was used for machine learning-based classification using boosted decision trees.

14. Universal Sentence Encoder was used for generating sentence-level semantic embeddings.

15. Matplotlib and Seaborn were used for plotting graphs, pie charts, and correlation heatmaps for dataset analysis and visualization.

## 

## 

## 

## 

## **CHAPTER 4**

## 

## 

## 

## 

## 

## 

## 

## 

## 

## **RESULTS AND DISCUSSION**

## 

## **4.1 RESULTS**

##  **4.1.1 REGULAR TRAINING**

## 

## The performance of the Code-Mixed dataset model was evaluated using Accuracy, Balanced Accuracy, Precision, Recall, Specificity, F1-Score and AUC-ROC across different hybrid models for comparison for the ground truth. The USE+LSTM model initially performed highest on the code-mixed dataset with 68% of accuracy and 59%approx in f1-score. GloVe+LightGBM scored the least in the scale comparison to other hybrid models with 65% accuracy and 57% f1 score. 

Table 4.1: Regular Result

| **Models** | **accuracy** | **bal_acc** | **precision** | **recall** | **specificity** | **f1_score** | **auc_roc** |
|----|---:|---:|---:|---:|---:|---:|---:|
| **Word2vec+LSTM** | **0.6675** | **0.6573** | **0.6914** | **0.5133** | **0.8012** | **0.5892** | **0.7294** |
| **GloVe+ LSTM** | **0.6779** | **0.6627** | **0.7591** | **0.4491** | **0.8763** | **0.5644** | **0.7532** |
| **FastText+LSTM** | **0.661** | **0.6496** | **0.6906** | **0.4895** | **0.8097** | **0.5729** | **0.7208** |
| **USE+LSTM** | **0.6805** | **0.6684** | **0.7283** | **0.4981** | **0.8388** | **0.5916** | **0.7555** |
| **ELMo+LSTM** | **0.6632** | **0.6503** | **0.7084** | **0.4674** | **0.8331** | **0.5632** | **0.7327** |
| **Word2Vec+LightGBM** | **0.6665** | **0.6548** | **0.7025** | **0.4893** | **0.8203** | **0.5768** | **0.7397** |
| **GloVe+LightGBM** | **0.6527** | **0.6403** | **0.6863** | **0.465** | **0.8156** | **0.5544** | **0.717** |
| **FastText+LightGBM** | **0.676** | **0.6643** | **0.7173** | **0.4993** | **0.8293** | **0.5888** | **0.7478** |
| **USE+LightGBM** | **0.6739** | **0.6619** | **0.716** | **0.4937** | **0.8302** | **0.5844** | **0.7462** |
| **ELMo+LightGBM** | **0.6744** | **0.6636** | **0.6808** | **0.6744** | **0.8156** | **0.6658** | **0.7469** |

## 

##  **4.1.2 LANUAGE WISE REGULAR TRAINING** 

## This is a one of a kind strategy technique that involves Language wise regular finetuning and it used the same performance metrics that is being used in Table 4.1. This work consists of various deep learning models with different background architectures and hybrid models. Here MPNet achieved highest accuracy of 74% and highest f1-score with 72% when all languages combined.

## **Table 4.2: English Strategy**

| **Model Name** | **Accuracy** | **Bal Acc** | **Precision** | **Recall** | **Specificity** | **F1 Score** | **AUC-ROC** |
|:--:|:--:|:--:|:--:|:--:|:--:|:--:|:--:|
| **mBART** | **0.835074** | **0.835074** | **0.834813** | **0.835556** | **0.834593** | **0.835184** | **0.912011** |
| **MuRIL** | **0.814625** | **0.81463** | **0.828996** | **0.792889** | **0.836372** | **0.810541** | **0.896677** |
| **HingRoBERTa** | **0.835519** | **0.835526** | **0.858159** | **0.804** | **0.867052** | **0.830197** | **0.917912** |
| **MPNet** | **0.822183** | **0.822181** | **0.817426** | **0.829778** | **0.814584** | **0.823555** | **0.899569** |
| **GloVe+BiLSTM** | **0.776238** | **0.776579** | **0.756138** | **0.808274** | **0.744885** | **0.781337** | **0.854612** |
| **Word2Vec+BiLSTM** | **0.719049** | **0.719050** | **0.722072** | **0.712444** | **0.725656** | **0.717226** | **0.797342** |
| **FastText+BiLSTM** | **0.752355** | **0.750867** | **0.740920** | **0.798956** | **0.702778** | **0.768844** | **0.825146** |

## **Table 4.3: Hindi Strategy**

| **Model Name** | **Accuracy** | **Bal Acc** | **Precision** | **Recall** | **Specificity** | **F1 Score** | **AUC-ROC** |
|:--:|:--:|:--:|:--:|:--:|:--:|:--:|:--:|
| **mBART** | **0.567762** | **0.590282** | **0.510024** | **0.799847** | **0.380717** | **0.622872** | **0.622511** |
| **MuRIL** | **0.607803** | **0.608876** | **0.554258** | **0.618865** | **0.598888** | **0.584783** | **0.67028** |
| **HingRoBERTa** | **0.605407** | **0.605746** | **0.55254** | **0.608896** | **0.602596** | **0.579351** | **0.631553** |
| **MPNet** | **0.599932** | **0.597676** | **0.549306** | **0.576687** | **0.618665** | **0.562664** | **0.611597** |
| **GloVe+BiLSTM** | **0.550690** | **0.500000** | **0.000000** | **0.000000** | **1.000000** | **0.000000** | **0.486025** |
| **Word2Vec+BiLSTM** | **0.600274** | **0.579456** | **0.578161** | **0.385736** | **0.773177** | **0.462741** | **0.604954** |
| **FastText+BiLSTM** | **0.645194** | **0.622978** | **0.612500** | **0.467780** | **0.778175** | **0.530447** | **0.673262** |

## 

## **Table 4.4: Hinglish Strategy**

| **Model** | **Accuracy** | **Bal Acc** | **Precision** | **Recall** | **Specificity** | **F1 Score** | **AUC-ROC** |
|:--:|:--:|:--:|:--:|:--:|:--:|:--:|:--:|
| **mBART** | **0.706909** | **0.671073** | **0.662005** | **0.50805** | **0.834096** | **0.574899** | **0.732349** |
| **MuRIL** | **0.679693** | **0.669716** | **0.583612** | **0.624329** | **0.715103** | **0.603284** | **0.745** |
| **HingRoBERTa** | **0.73552** | **0.709035** | **0.688285** | **0.588551** | **0.829519** | **0.634523** | **0.778005** |
| **MPNet** | **0.717376** | **0.681911** | **0.679907** | **0.520572** | **0.843249** | **0.589666** | **0.753845** |
| **GloVe+BiLSTM** | **0.697939** | **0.638326** | **0.772000** | **0.344029** | **0.932624** | **0.475956** | **0.705756** |
| **Word2Vec+BiLSTM** | **0.706211** | **0.662119** | **0.682540** | **0.461538** | **0.862700** | **0.550694** | **0.740799** |
| **FastText+BiLSTM** | **0.691358** | **0.619592** | **0.710843** | **0.318919** | **0.920266** | **0.440299** | **0.688354** |

## 

## **Table 4.5: Combined Strategy**

| **Model** | **Accuracy** | **Bal Acc** | **Precision** | **Recall** | **Specificity** | **F1 Score** | **AUC-ROC** |
|:--:|:--:|:--:|:--:|:--:|:--:|:--:|:--:|
| **mBART** | **0.734184** | **0.734024** | **0.706504** | **0.731761** | **0.736287** | **0.718911** | **0.827426** |
| **MuRIL** | **0.721532** | **0.721742** | **0.690934** | **0.724708** | **0.718776** | **0.707418** | **0.804714** |
| **HingRoBERTa** | **0.748531** | **0.7432** | **0.761364** | **0.668045** | **0.818354** | **0.711658** | **0.83382** |
| **MPNet** | **0.74164** | **0.741131** | **0.716694** | **0.733949** | **0.748312** | **0.725219** | **0.82881** |
| **GloVe+BiLSTM** | **0.683009** | **0.677217** | **0.681793** | **0.595574** | **0.758861** | **0.635774** | **0.763671** |
| **Word2Vec+BiLSTM** | **0.670357** | **0.662825** | **0.676418** | **0.556663** | **0.768987** | **0.610726** | **0.736175** |
| **FastText+BiLSTM** | **0.676610** | **0.665650** | **0.710953** | **0.511679** | **0.819620** | **0.595076** | **0.754570** |

## **4.1.3 MULTI-STAGE LANGUAGE TRAINING**

## This work includes the multi-stage language training with different models by which it means that instead of regular fine tuning this training goes through sequential starts initial from English then Hinglish then Hindi and at last all combined. GloVe+BiLSTM has gained the highest accuracy with 82% and 80% f1-score.

## **Table 4.6: English Strategy**

| **Model Name** | **Accuracy** | **Bal Acc** | **Precision** | **Recall** | **Specificity** | **F1 Score** | **AUC-ROC** |
|:--:|:--:|:--:|:--:|:--:|:--:|:--:|:--:|
| **mBART** | **0.829483** | **0.829266** | **0.840185** | **0.809164** | **0.849369** | **0.824383** | **0.897722** |
| **MuRIL** | **0.796480** | **0.796017** | **0.820650** | **0.753114** | **0.838920** | **0.785433** | **0.880163** |
| **HingRoBERTa** | **0.832563** | **0.832665** | **0.823401** | **0.842082** | **0.823248** | **0.832637** | **0.906317** |
| **MPNet** | **0.816502** | **0.816566** | **0.809545** | **0.822509** | **0.810623** | **0.815975** | **0.893683** |
| **GloVe+BiLSTM** | **0.6106** | **0.6258** | **0.5532** | **0.8407** | **0.4110** | **0.6673** | **0.6250** |
| **Word2Vec+BiLSTM** | **0.587438** | **0.588573** | **0.550975** | **0.60457** | **0.572574** | **0.576531** | **0.636518** |

## 

## **Table 4.7: Hindi Strategy**

| **Model Name** | **Accuracy** | **Bal Acc** | **Precision** | **Recall** | **Specificity** | **F1 Score** | **AUC-ROC** |
|:--:|:--:|:--:|:--:|:--:|:--:|:--:|:--:|
| **mBART** | **0.600345** | **0.594817** | **0.556962** | **0.540292** | **0.649343** | **0.5485** | **0.645617** |
| **MuRIL** | **0.598621** | **0.588589** | **0.561126** | **0.489639** | **0.687539** | **0.522951** | **0.637667** |
| **HingRoBERTa** | **0.598966** | **0.595967** | **0.552395** | **0.566385** | **0.625548** | **0.559303** | **0.631097** |
| **MPNet** | **0.605172** | **0.602944** | **0.55826** | **0.580967** | **0.624922** | **0.569387** | **0.632868** |
| **GloVe+BiLSTM** | **0.5276** | **0.5382** | **0.4939** | **0.6885** | **0.3880** | **0.5752** | **0.5192** |
| **Word2Vec+BiLSTM** | **0.624492** | **0.624139** | **0.591543** | **0.619163** | **0.629114** | **0.605038** | **0.673203** |
| **FastText+BiLSTM** | **0.585889** | **0.585743** | **0.514705** | **0.584725** | **0.586762** | **0.547486** | **0.612223** |

## 

## **Table 4.8: Hinglish Strategy**

| **Model** | **Accuracy** | **Bal Acc** | **Precision** | **Recall** | **Specificity** | **F1 Score** | **AUC-ROC** |
|:--:|:--:|:--:|:--:|:--:|:--:|:--:|:--:|
| **mBART** | **0.687278** | **0.660987** | **0.627368** | **0.531194** | **0.79078** | **0.57529** | **0.696596** |
| **MuRIL** | **0.697939** | **0.657843** | **0.678947** | **0.459893** | **0.855792** | **0.548353** | **0.691189** |
| **HingRoBERTa** | **0.707889** | **0.702147** | **0.623762** | **0.673797** | **0.730496** | **0.647815** | **0.750452** |
| **MPNet** | **0.685856** | **0.676619** | **0.601019** | **0.631016** | **0.722222** | **0.615652** | **0.738968** |
| **GloVe+BiLSTM** | **0.5198** | **0.5147** | **0.4816** | **0.4431** | **0.5863** | **0.4616** | **0.5225** |
| **Word2Vec+BiLSTM** | **0.566539** | **0.536258** | **0.72** | **0.109436** | **0.96308** | **0.189994** | **0.577701** |
| **FastText+BiLSTM** | **0.728395** | **0.651575** | **0.884057** | **0.329729** | **0.973421** | **0.480314** | **0.687276** |

## 

## **Table 4.9: Combined Strategy**

| **Model** | **Accuracy** | **Bal Acc** | **Precision** | **Recall** | **Specificity** | **F1 Score** | **AUC-ROC** |
|:--:|:--:|:--:|:--:|:--:|:--:|:--:|:--:|
| **mBART** | **0.731812** | **0.72878** | **0.722592** | **0.686041** | **0.771519** | **0.703842** | **0.802136** |
| **MuRIL** | **0.715996** | **0.710274** | **0.723184** | **0.629621** | **0.790928** | **0.673167** | **0.795080** |
| **HingRoBERTa** | **0.736218** | **0.735923** | **0.709502** | **0.731761** | **0.740084** | **0.720460** | **0.820400** |
| **MPNet** | **0.726502** | **0.726061** | **0.699929** | **0.719844** | **0.732278** | **0.709747** | **0.810777** |
| **GloVe+BiLSTM** | **0.8204** | **0.8186** | **0.8149** | **0.7935** | **0.8437** | **0.8041** | **0.9139** |
| **Word2Vec+BiLSTM** | **0.730456** | **0.726676** | **0.72639** | **0.673395** | **0.779958** | **0.698889** | **0.809126** |
| **FastText+BiLSTM** | **0.701762** | **0.700688** | **0.676668** | **0.685554** | **0.715822** | **0.681082** | **0.770482** |

## **4.1.4 MULTI-STAGE LANGUAGE TRAINING WITH SIX VARIATIONS ON MuRIL AND GloVe + BiLSTM**

## This work involves two models MuRIL and Glove+BiLSTM for model training that includes multi-stage language training of six variations using (English, Hinglish, Hindi, all combined), (English, Hindi, Hinglish, all combined ), (Hinglish, English, Hindi, all combined), (Hinglish, Hindi, English, all combined), (Hindi, Hinglish, English, all combined), and (Hindi, English, Hinglish, all combined) to check which variation performs well on which model. So, we have achieved the highest 72% accuracy with Hinglish, Hindi, English, all combined variation from MuRIL model and highest accuracy with 66% in Hindi, Hinglish, English, all combined variation using GloVe+BiLSTM model.

## **Table 4.10: Combined Strategy**

<table>
<colgroup>
<col style="width: 16%" />
<col style="width: 12%" />
<col style="width: 6%" />
<col style="width: 9%" />
<col style="width: 11%" />
<col style="width: 9%" />
<col style="width: 6%" />
<col style="width: 10%" />
<col style="width: 6%" />
<col style="width: 8%" />
</colgroup>
<thead>
<tr>
<th><strong>Model</strong></th>
<th><strong>Strategy</strong></th>
<th><strong>Phase</strong></th>
<th><strong>Accuracy</strong></th>
<th><strong>Balanced Acc</strong></th>
<th><strong>Precision</strong></th>
<th><strong>Recall</strong></th>
<th><strong>Specificity</strong></th>
<th><strong>F1</strong></th>
<th><strong>ROC-AUC</strong></th>
</tr>
</thead>
<tbody>
<tr>
<td rowspan="24"><strong>BiLSTM+GloVe</strong></td>
<td rowspan="4"><strong>E -&gt; HG -&gt; H -&gt; F</strong></td>
<td><strong>E</strong></td>
<td><strong>0.7498</strong></td>
<td><strong>0.7503</strong></td>
<td><strong>0.7243</strong></td>
<td><strong>0.7980</strong></td>
<td><strong>0.7027</strong></td>
<td><strong>0.7594</strong></td>
<td><strong>0.8190</strong></td>
</tr>
<tr>
<td><strong>HG</strong></td>
<td><strong>0.4222</strong></td>
<td><strong>0.4862</strong></td>
<td><strong>0.3906</strong></td>
<td><strong>0.8021</strong></td>
<td><strong>0.1702</strong></td>
<td><strong>0.5254</strong></td>
<td><strong>0.4919</strong></td>
</tr>
<tr>
<td><strong>H</strong></td>
<td><strong>0.4759</strong></td>
<td><strong>0.5007</strong></td>
<td><strong>0.4498</strong></td>
<td><strong>0.7460</strong></td>
<td><strong>0.2555</strong></td>
<td><strong>0.5612</strong></td>
<td><strong>0.5025</strong></td>
</tr>
<tr>
<td><strong>F</strong></td>
<td><strong>0.6080</strong></td>
<td><strong>0.6195</strong></td>
<td><strong>0.5554</strong></td>
<td><strong>0.7821</strong></td>
<td><strong>0.4570</strong></td>
<td><strong>0.6496</strong></td>
<td><strong>0.7023</strong></td>
</tr>
<tr>
<td rowspan="4"><strong>E -&gt; H -&gt; HG -&gt; F</strong></td>
<td><strong>E</strong></td>
<td><strong>0.7391</strong></td>
<td><strong>0.7392</strong></td>
<td><strong>0.7301</strong></td>
<td><strong>0.7496</strong></td>
<td><strong>0.7288</strong></td>
<td><strong>0.7397</strong></td>
<td><strong>0.8147</strong></td>
</tr>
<tr>
<td><strong>HG</strong></td>
<td><strong>0.4890</strong></td>
<td><strong>0.5081</strong></td>
<td><strong>0.4053</strong></td>
<td><strong>0.6025</strong></td>
<td><strong>0.4137</strong></td>
<td><strong>0.4846</strong></td>
<td><strong>0.4989</strong></td>
</tr>
<tr>
<td><strong>H</strong></td>
<td><strong>0.5731</strong></td>
<td><strong>0.5428</strong></td>
<td><strong>0.5569</strong></td>
<td><strong>0.2441</strong></td>
<td><strong>0.8416</strong></td>
<td><strong>0.3394</strong></td>
<td><strong>0.5863</strong></td>
</tr>
<tr>
<td><strong>F</strong></td>
<td><strong>0.6449</strong></td>
<td><strong>0.6399</strong></td>
<td><strong>0.6305</strong></td>
<td><strong>0.5693</strong></td>
<td><strong>0.7105</strong></td>
<td><strong>0.5983</strong></td>
<td><strong>0.7163</strong></td>
</tr>
<tr>
<td rowspan="4"><strong>HG -&gt; E -&gt; H -&gt; F</strong></td>
<td><strong>E</strong></td>
<td><strong>0.7393</strong></td>
<td><strong>0.7392</strong></td>
<td><strong>0.7370</strong></td>
<td><strong>0.7353</strong></td>
<td><strong>0.7431</strong></td>
<td><strong>0.7361</strong></td>
<td><strong>0.8203</strong></td>
</tr>
<tr>
<td><strong>HG</strong></td>
<td><strong>0.5942</strong></td>
<td><strong>0.5202</strong></td>
<td><strong>0.4728</strong></td>
<td><strong>0.1551</strong></td>
<td><strong>0.8853</strong></td>
<td><strong>0.2336</strong></td>
<td><strong>0.5822</strong></td>
</tr>
<tr>
<td><strong>H</strong></td>
<td><strong>0.6079</strong></td>
<td><strong>0.5937</strong></td>
<td><strong>0.5817</strong></td>
<td><strong>0.4536</strong></td>
<td><strong>0.7339</strong></td>
<td><strong>0.5097</strong></td>
<td><strong>0.6320</strong></td>
</tr>
<tr>
<td><strong>F</strong></td>
<td><strong>0.6732</strong></td>
<td><strong>0.6661</strong></td>
<td><strong>0.6770</strong></td>
<td><strong>0.5669</strong></td>
<td><strong>0.7654</strong></td>
<td><strong>0.6171</strong></td>
<td><strong>0.7438</strong></td>
</tr>
<tr>
<td rowspan="4"><strong>HG -&gt; H -&gt; E -&gt; F</strong></td>
<td><strong>E</strong></td>
<td><strong>0.6854</strong></td>
<td><strong>0.6855</strong></td>
<td><strong>0.6754</strong></td>
<td><strong>0.7006</strong></td>
<td><strong>0.6704</strong></td>
<td><strong>0.6878</strong></td>
<td><strong>0.7586</strong></td>
</tr>
<tr>
<td><strong>HG</strong></td>
<td><strong>0.6347</strong></td>
<td><strong>0.5581</strong></td>
<td><strong>0.6516</strong></td>
<td><strong>0.1800</strong></td>
<td><strong>0.9362</strong></td>
<td><strong>0.2821</strong></td>
<td><strong>0.5664</strong></td>
</tr>
<tr>
<td><strong>H</strong></td>
<td><strong>0.6372</strong></td>
<td><strong>0.6248</strong></td>
<td><strong>0.6187</strong></td>
<td><strong>0.5019</strong></td>
<td><strong>0.7477</strong></td>
<td><strong>0.5542</strong></td>
<td><strong>0.6598</strong></td>
</tr>
<tr>
<td><strong>F</strong></td>
<td><strong>0.6615</strong></td>
<td><strong>0.6553</strong></td>
<td><strong>0.6574</strong></td>
<td><strong>0.5666</strong></td>
<td><strong>0.7439</strong></td>
<td><strong>0.6087</strong></td>
<td><strong>0.7110</strong></td>
</tr>
<tr>
<td rowspan="4"><strong>H -&gt; HG -&gt; E -&gt; F</strong></td>
<td><strong>E</strong></td>
<td><strong>0.5963</strong></td>
<td><strong>0.5965</strong></td>
<td><strong>0.5878</strong></td>
<td><strong>0.6152</strong></td>
<td><strong>0.5777</strong></td>
<td><strong>0.6012</strong></td>
<td><strong>0.6335</strong></td>
</tr>
<tr>
<td><strong>HG</strong></td>
<td><strong>0.6084</strong></td>
<td><strong>0.5362</strong></td>
<td><strong>0.5260</strong></td>
<td><strong>0.1800</strong></td>
<td><strong>0.8924</strong></td>
<td><strong>0.2683</strong></td>
<td><strong>0.5194</strong></td>
</tr>
<tr>
<td><strong>H</strong></td>
<td><strong>0.6331</strong></td>
<td><strong>0.6257</strong></td>
<td><strong>0.5995</strong></td>
<td><strong>0.5526</strong></td>
<td><strong>0.6988</strong></td>
<td><strong>0.5751</strong></td>
<td><strong>0.6660</strong></td>
</tr>
<tr>
<td><strong>F</strong></td>
<td><strong>0.6103</strong></td>
<td><strong>0.6053</strong></td>
<td><strong>0.5884</strong></td>
<td><strong>0.5360</strong></td>
<td><strong>0.6747</strong></td>
<td><strong>0.5610</strong></td>
<td><strong>0.6455</strong></td>
</tr>
<tr>
<td rowspan="4"><strong>H -&gt; E -&gt; HG -&gt; F</strong></td>
<td><strong>E</strong></td>
<td><strong>0.7193</strong></td>
<td><strong>0.7188</strong></td>
<td><strong>0.7359</strong></td>
<td><strong>0.6744</strong></td>
<td><strong>0.7632</strong></td>
<td><strong>0.7038</strong></td>
<td><strong>0.8011</strong></td>
</tr>
<tr>
<td><strong>HG</strong></td>
<td><strong>0.5572</strong></td>
<td><strong>0.4940</strong></td>
<td><strong>0.3835</strong></td>
<td><strong>0.1818</strong></td>
<td><strong>0.8061</strong></td>
<td><strong>0.2467</strong></td>
<td><strong>0.5129</strong></td>
</tr>
<tr>
<td><strong>H</strong></td>
<td><strong>0.6272</strong></td>
<td><strong>0.6165</strong></td>
<td><strong>0.6002</strong></td>
<td><strong>0.5104</strong></td>
<td><strong>0.7226</strong></td>
<td><strong>0.5516</strong></td>
<td><strong>0.6606</strong></td>
</tr>
<tr>
<td><strong>F</strong></td>
<td><strong>0.6634</strong></td>
<td><strong>0.6562</strong></td>
<td><strong>0.6648</strong></td>
<td><strong>0.5552</strong></td>
<td><strong>0.7572</strong></td>
<td><strong>0.6051</strong></td>
<td><strong>0.7297</strong></td>
</tr>
<tr>
<td rowspan="24"><strong>MuRIL</strong></td>
<td rowspan="4"><strong>E -&gt; HG -&gt; H -&gt; F</strong></td>
<td><strong>E</strong></td>
<td><strong>0.7965</strong></td>
<td><strong>0.7960</strong></td>
<td><strong>0.8206</strong></td>
<td><strong>0.7531</strong></td>
<td><strong>0.8389</strong></td>
<td><strong>0.7854</strong></td>
<td><strong>0.8802</strong></td>
</tr>
<tr>
<td><strong>HG</strong></td>
<td><strong>0.6979</strong></td>
<td><strong>0.6578</strong></td>
<td><strong>0.6789</strong></td>
<td><strong>0.4599</strong></td>
<td><strong>0.8558</strong></td>
<td><strong>0.5484</strong></td>
<td><strong>0.6912</strong></td>
</tr>
<tr>
<td><strong>H</strong></td>
<td><strong>0.5986</strong></td>
<td><strong>0.5886</strong></td>
<td><strong>0.5611</strong></td>
<td><strong>0.4896</strong></td>
<td><strong>0.6875</strong></td>
<td><strong>0.5230</strong></td>
<td><strong>0.6377</strong></td>
</tr>
<tr>
<td><strong>F</strong></td>
<td><strong>0.7160</strong></td>
<td><strong>0.7103</strong></td>
<td><strong>0.7232</strong></td>
<td><strong>0.6296</strong></td>
<td><strong>0.7909</strong></td>
<td><strong>0.6732</strong></td>
<td><strong>0.7951</strong></td>
</tr>
<tr>
<td rowspan="4"><strong>E -&gt; H -&gt; HG -&gt; F</strong></td>
<td><strong>E</strong></td>
<td><strong>0.7985</strong></td>
<td><strong>0.7977</strong></td>
<td><strong>0.8422</strong></td>
<td><strong>0.7291</strong></td>
<td><strong>0.8663</strong></td>
<td><strong>0.7816</strong></td>
<td><strong>0.8645</strong></td>
</tr>
<tr>
<td><strong>HG</strong></td>
<td><strong>0.6645</strong></td>
<td><strong>0.6478</strong></td>
<td><strong>0.5817</strong></td>
<td><strong>0.5651</strong></td>
<td><strong>0.7305</strong></td>
<td><strong>0.5732</strong></td>
<td><strong>0.7036</strong></td>
</tr>
<tr>
<td><strong>H</strong></td>
<td><strong>0.5793</strong></td>
<td><strong>0.5803</strong></td>
<td><strong>0.5285</strong></td>
<td><strong>0.5902</strong></td>
<td><strong>0.5704</strong></td>
<td><strong>0.5577</strong></td>
<td><strong>0.6034</strong></td>
</tr>
<tr>
<td><strong>F</strong></td>
<td><strong>0.7054</strong></td>
<td><strong>0.7025</strong></td>
<td><strong>0.6906</strong></td>
<td><strong>0.6627</strong></td>
<td><strong>0.7424</strong></td>
<td><strong>0.6763</strong></td>
<td><strong>0.7782</strong></td>
</tr>
<tr>
<td rowspan="4"><strong>HG -&gt; E -&gt; H -&gt; F</strong></td>
<td><strong>E</strong></td>
<td><strong>0.8090</strong></td>
<td><strong>0.8089</strong></td>
<td><strong>0.8103</strong></td>
<td><strong>0.8016</strong></td>
<td><strong>0.8163</strong></td>
<td><strong>0.8059</strong></td>
<td><strong>0.8318</strong></td>
</tr>
<tr>
<td><strong>HG</strong></td>
<td><strong>0.6660</strong></td>
<td><strong>0.6318</strong></td>
<td><strong>0.6061</strong></td>
<td><strong>0.4635</strong></td>
<td><strong>0.8002</strong></td>
<td><strong>0.5253</strong></td>
<td><strong>0.6622</strong></td>
</tr>
<tr>
<td><strong>H</strong></td>
<td><strong>0.5931</strong></td>
<td><strong>0.5868</strong></td>
<td><strong>0.5494</strong></td>
<td><strong>0.5249</strong></td>
<td><strong>0.6487</strong></td>
<td><strong>0.5369</strong></td>
<td><strong>0.6083</strong></td>
</tr>
<tr>
<td><strong>F</strong></td>
<td><strong>0.7155</strong></td>
<td><strong>0.7124</strong></td>
<td><strong>0.7045</strong></td>
<td><strong>0.6678</strong></td>
<td><strong>0.7570</strong></td>
<td><strong>0.6856</strong></td>
<td><strong>0.7442</strong></td>
</tr>
<tr>
<td rowspan="4"><strong>HG -&gt; H -&gt; E -&gt; F</strong></td>
<td><strong>E</strong></td>
<td><strong>0.7985</strong></td>
<td><strong>0.7977</strong></td>
<td><strong>0.8422</strong></td>
<td><strong>0.7291</strong></td>
<td><strong>0.8663</strong></td>
<td><strong>0.7816</strong></td>
<td><strong>0.8779</strong></td>
</tr>
<tr>
<td><strong>HG</strong></td>
<td><strong>0.6830</strong></td>
<td><strong>0.6547</strong></td>
<td><strong>0.6242</strong></td>
<td><strong>0.5152</strong></td>
<td><strong>0.7943</strong></td>
<td><strong>0.5645</strong></td>
<td><strong>0.7030</strong></td>
</tr>
<tr>
<td><strong>H</strong></td>
<td><strong>0.6279</strong></td>
<td><strong>0.6165</strong></td>
<td><strong>0.6029</strong></td>
<td><strong>0.5035</strong></td>
<td><strong>0.7295</strong></td>
<td><strong>0.5487</strong></td>
<td><strong>0.6506</strong></td>
</tr>
<tr>
<td><strong>F</strong></td>
<td><strong>0.7242</strong></td>
<td><strong>0.7179</strong></td>
<td><strong>0.7389</strong></td>
<td><strong>0.6284</strong></td>
<td><strong>0.8074</strong></td>
<td><strong>0.6792</strong></td>
<td><strong>0.7981</strong></td>
</tr>
<tr>
<td rowspan="4"><strong>H -&gt; HG -&gt; E -&gt; F</strong></td>
<td><strong>E</strong></td>
<td><strong>0.7901</strong></td>
<td><strong>0.7894</strong></td>
<td><strong>0.8304</strong></td>
<td><strong>0.7233</strong></td>
<td><strong>0.8555</strong></td>
<td><strong>0.7732</strong></td>
<td><strong>0.8516</strong></td>
</tr>
<tr>
<td><strong>HG</strong></td>
<td><strong>0.6802</strong></td>
<td><strong>0.6590</strong></td>
<td><strong>0.6086</strong></td>
<td><strong>0.5544</strong></td>
<td><strong>0.7636</strong></td>
<td><strong>0.5802</strong></td>
<td><strong>0.6945</strong></td>
</tr>
<tr>
<td><strong>H</strong></td>
<td><strong>0.6238</strong></td>
<td><strong>0.6179</strong></td>
<td><strong>0.5849</strong></td>
<td><strong>0.5602</strong></td>
<td><strong>0.6756</strong></td>
<td><strong>0.5723</strong></td>
<td><strong>0.6334</strong></td>
</tr>
<tr>
<td><strong>F</strong></td>
<td><strong>0.7181</strong></td>
<td><strong>0.7135</strong></td>
<td><strong>0.7175</strong></td>
<td><strong>0.6486</strong></td>
<td><strong>0.7785</strong></td>
<td><strong>0.6813</strong></td>
<td><strong>0.7770</strong></td>
</tr>
<tr>
<td rowspan="4"><strong>H -&gt; E -&gt; HG -&gt; F</strong></td>
<td><strong>E</strong></td>
<td><strong>0.7668</strong></td>
<td><strong>0.7654</strong></td>
<td><strong>0.8502</strong></td>
<td><strong>0.6415</strong></td>
<td><strong>0.8894</strong></td>
<td><strong>0.7312</strong></td>
<td><strong>0.8083</strong></td>
</tr>
<tr>
<td><strong>HG</strong></td>
<td><strong>0.6915</strong></td>
<td><strong>0.6486</strong></td>
<td><strong>0.6749</strong></td>
<td><strong>0.4367</strong></td>
<td><strong>0.8605</strong></td>
<td><strong>0.5303</strong></td>
<td><strong>0.6898</strong></td>
</tr>
<tr>
<td><strong>H</strong></td>
<td><strong>0.6193</strong></td>
<td><strong>0.5971</strong></td>
<td><strong>0.6268</strong></td>
<td><strong>0.3776</strong></td>
<td><strong>0.8165</strong></td>
<td><strong>0.4713</strong></td>
<td><strong>0.6455</strong></td>
</tr>
<tr>
<td><strong>F</strong></td>
<td><strong>0.7065</strong></td>
<td><strong>0.6948</strong></td>
<td><strong>0.7662</strong></td>
<td><strong>0.5299</strong></td>
<td><strong>0.8597</strong></td>
<td><strong>0.6265</strong></td>
<td><strong>0.7508</strong></td>
</tr>
</tbody>
</table>

## **4.1.5 SARVAM MODEL \***

We have used SARVAM AI Model which is an Indian startup specially made on and made for Indian languages. It is a LLM capable of speech-to-text, and text-to-speech. SARVAM has different models with different parameters but we have implemented Sarvam-1 that consists of 2 Billion parameters. It supports 22+ Indian languages with different scripts. Sarvam can also well handle the code-mixed texts so we implemented our dataset with Sarvam for model training an evaluation as an experimental work with llms and we have achieved highest performance in all 6 metrics.

Table 4.11: Performance Evaluation

<table style="width:54%;">
<colgroup>
<col style="width: 37%" />
<col style="width: 16%" />
</colgroup>
<thead>
<tr>
<th><p><strong>Β </strong></p>
<p><strong>Metrics</strong></p></th>
<th><strong>Value</strong></th>
</tr>
</thead>
<tbody>
<tr>
<td><strong>Accuracy</strong></td>
<td><strong>0.9373</strong></td>
</tr>
<tr>
<td><strong>Balanced Accuracy</strong></td>
<td><strong>0.9284</strong></td>
</tr>
<tr>
<td><strong>Precision</strong></td>
<td><strong>0.9486</strong></td>
</tr>
<tr>
<td><strong>Recall</strong></td>
<td><strong>0.8877</strong></td>
</tr>
<tr>
<td><strong>Specificity</strong></td>
<td><strong>0.9691</strong></td>
</tr>
<tr>
<td><strong>F1-Score</strong></td>
<td><strong>0.9171</strong></td>
</tr>
<tr>
<td><strong>ROC-AUC</strong></td>
<td><strong>0.9326</strong></td>
</tr>
</tbody>
</table>

## **4.2 VISUALIZATION** 

## **4.2.1 REGULAR TRAINING**

This work involves total 39 figures for visualisation that uses two variations of hybrid models one is (Word2Vec, GloVe, FastText, USE and ELMo) with LSTM and another with LightGBM. Fig:4.1 – Fig: 4.10 consists of Train vs Validation Accuracy Curve and Train vs Validation Loss Curve, Fig:4.11 – Fig: 4.20 consists of AUC-ROC Curves, FIG:4.21 – Fig:4.30 consists of Confusion Matrix and Fig:4.31-Fig:4.39 is t-SNE Visualisation.

<img src="thesis_media/media/image10.png" style="width:3.06944in;height:2.39296in" /><img src="thesis_media/media/image11.png" style="width:3.00717in;height:2.375in" />

Fig:4.1 Training vs Validation Loss and Accuracy Word2Vec+LSTM

<img src="thesis_media/media/image12.png" style="width:2.9021in;height:2.2724in" /><img src="thesis_media/media/image13.png" style="width:2.83916in;height:2.25189in" />

Fig:4.2 Training vs Validation Loss and Accuracy GloVe+LSTM

<img src="thesis_media/media/image14.png" style="width:3.15385in;height:2.36538in" /> <img src="thesis_media/media/image15.png" style="width:3.1477in;height:2.36078in" />

Fig:4.3 Training vs Validation Loss and Accuracy FastText+LSTM

<img src="thesis_media/media/image16.png" style="width:2.82517in;height:2.20261in" /> <img src="thesis_media/media/image17.png" style="width:2.7972in;height:2.20917in" />

Fig:4.4 Training vs Validation Loss and Accuracy USE+LSTM

<img src="thesis_media/media/image18.png" style="width:3.04196in;height:2.37153in" /><img src="thesis_media/media/image19.png" style="width:2.94406in;height:2.32515in" />

Fig:4.5 Training vs Validation Loss and Accuracy ELMo+LSTM

<img src="thesis_media/media/image20.png" style="width:3.07639in;height:2.46732in" /><img src="thesis_media/media/image21.png" style="width:3.16084in;height:2.63424in" />

Fig:4.6 Training vs Validation Loss and Accuracy Word2Vec+LightGBM

<img src="thesis_media/media/image22.png" style="width:3.03497in;height:2.37735in" /><img src="thesis_media/media/image23.png" style="width:3.04196in;height:2.38282in" />

Fig:4.7Training vs Validation Loss and Accuracy GloVe+LightGBM

<img src="thesis_media/media/image24.png" style="width:3.22378in;height:2.40165in" /><img src="thesis_media/media/image25.png" style="width:3.2028in;height:2.38699in" />

Fig:4.8 Training vs Validation Loss and Accuracy FastText+LightGBM

<img src="thesis_media/media/image26.png" style="width:3.11806in;height:2.32662in" /><img src="thesis_media/media/image27.png" style="width:3.1049in;height:2.31745in" />

Fig:4.9 Training vs Validation Loss and Accuracy USE+LightGBM

<img src="thesis_media/media/image28.png" style="width:3.17361in;height:2.47532in" /><img src="thesis_media/media/image29.png" style="width:3.0979in;height:2.485in" />

Fig:4.10 Training vs Validation Loss and Accuracy ELMo+LightGBM

<img src="thesis_media/media/image30.png" style="width:3.11806in;height:2.46239in" /> <img src="thesis_media/media/image31.png" style="width:3.08392in;height:2.44602in" />

Fig:4.11 ROC-AUC Curve Word2Vec+LSTM Fig:4.12 ROC-AUC Curve GloVe+LSTM

<img src="thesis_media/media/image32.png" style="width:3.35664in;height:2.51767in" /><img src="thesis_media/media/image33.png" style="width:2.99301in;height:2.36381in" />

Fig:4.13 ROC-AUC Curve FastText+LSTM Fig:4.14 ROC-AUC Curve USE+LSTM

<img src="thesis_media/media/image34.png" style="width:3.09973in;height:2.44792in" /><img src="thesis_media/media/image35.png" style="width:3.07692in;height:2.45962in" />

Fig:4.15 ROC-AUC Curve ELMo+ Fig:4.16 ROC-AUC Curve Word2Vec+LightGBM

<img src="thesis_media/media/image36.png" style="width:3.02083in;height:2.39598in" /><img src="thesis_media/media/image37.png" style="width:3.13287in;height:2.33487in" />

Fig:4.17 ROC-AUC Curve GloVe+LightGBM Fig:4.18 ROC-AUC Curve FastText+LightGBM

<img src="thesis_media/media/image38.png" style="width:3.07639in;height:2.29621in" /><img src="thesis_media/media/image39.png" style="width:2.93747in;height:2.28365in" />

Fig:4.19 ROC-AUC Curve USE+LightGBM Fig:4.20 ROC-AUC Curve ELMo+LightGBM

<img src="thesis_media/media/image40.png" style="width:3.07639in;height:2.55116in" /><img src="thesis_media/media/image41.png" style="width:3.02083in;height:2.51613in" />

Fig:4.21 Confusion Matrix Word2Vec+LSTM Fig:4.22 Confusion Matrix GloVe+LSTM

<img src="thesis_media/media/image42.png" style="width:3.20347in;height:2.40278in" /><img src="thesis_media/media/image43.png" style="width:2.90972in;height:2.41295in" />

Fig:4.23 Confusion Matrix FastText+LSTM Fig:4.24 Confusion Matrix USE+LSTM

<img src="thesis_media/media/image44.png" style="width:3.22917in;height:2.48323in" /><img src="thesis_media/media/image45.png" style="width:2.81476in;height:2.49792in" />

Fig:4.25 Confusion Matrix ELMo+LSTM Fig:4.26 Confusion Matrix Word2Vec+LightGBM

<img src="thesis_media/media/image46.png" style="width:3.05556in;height:2.54505in" /><img src="thesis_media/media/image47.png" style="width:2.80556in;height:2.53666in" />

Fig:4.27 Confusion Matrix GloVe+LightGBM Fig:4.28 Confusion Matrix FastText+LightGBM

<img src="thesis_media/media/image48.png" style="width:3.04787in;height:2.39176in" /><img src="thesis_media/media/image49.png" style="width:2.79861in;height:2.39181in" />

Fig:4.29 Confusion Matrix USE+LightGBM Fig:4.30 Confusion Matrix ELMo+LightGBM

<img src="thesis_media/media/image50.png" style="width:3.07639in;height:2.41551in" /><img src="thesis_media/media/image51.png" style="width:3.1111in;height:2.40278in" />

Fig: 4.31 t-SNE for GloVe + LSTM Fig: 4.32 t-SNE for ELMo + LSTM

<img src="thesis_media/media/image52.png" style="width:2.99696in;height:2.47168in" /> <img src="thesis_media/media/image53.png" style="width:3.23077in;height:2.5304in" />

Fig 4.33: t-SNE for Word2Vec+ LightGBM Fig:4.34: t-SNE for GloVe + LightGBM

<img src="thesis_media/media/image54.png" style="width:5.37063in;height:2.60737in" />

Fig:4.35 t-SNE for Comparison FastText + LSTM

<img src="thesis_media/media/image55.png" style="width:2.95624in;height:2.20556in" /> <img src="thesis_media/media/image56.png" style="width:3.0979in;height:2.29918in" />

Fig 4.36: t-SNE for FastText + LSTM Fig:4.37 t-SNE for Use + LightGBM

<img src="thesis_media/media/image57.png" style="width:6.26806in;height:2.75347in" />

Fig:4.38 t-SNE for cluster comparison USE + LightGBM

<img src="thesis_media/media/image58.png" style="width:3.18486in;height:2.45833in" />

Fig: 4.39 t-SNE for FastText + LSTM

## 

## **4.2.2 LANGUAGE WISE REGULAR TRAINING** 

This work involves total 19 figures for visualisation that uses multiplevariations of hybrid models (GloVe+BiLSTM, FastText+BiLSTM, Word2Vec+BiLSTM) and deep learning models(Mbart, MuRIL, HingRoBERTa-mixed, MPNet).but we have used only all combined(Full dataset) visualisations Fig:4.40 – Fig: 4.46 consists of Train vs Validation Accuracy Curve and Train vs Validation Loss Curve, Fig:4.47 – Fig: 4.53 consists of AUC-ROC Curves, FIG:4.54 – Fig:4.59 consists of Confusion Matrix.

<img src="thesis_media/media/image59.png" style="width:3.18182in;height:2.3198in" /><img src="thesis_media/media/image60.png" style="width:3.13287in;height:2.29218in" />

Fig:4.40 Training vs Validation Loss and Accuracy MuRIL Full dataset

<img src="thesis_media/media/image61.png" style="width:3.20979in;height:2.38357in" /><img src="thesis_media/media/image62.png" style="width:3.18881in;height:2.36185in" />

Fig:4.41 Training vs Validation Loss and Accuracy mBART Full dataset

<img src="thesis_media/media/image63.png" style="width:3.22378in;height:2.39362in" /><img src="thesis_media/media/image64.png" style="width:3.23077in;height:2.37606in" />

Fig:4.42 Training vs Validation Loss and Accuracy HingRoBERTa Full dataset

<img src="thesis_media/media/image65.png" style="width:3.12587in;height:2.31344in" /><img src="thesis_media/media/image66.png" style="width:3.13287in;height:2.31758in" />

Fig:4.43 Training vs Validation Loss and Accuracy MPNet Full dataset

<img src="thesis_media/media/image67.png" style="width:3.21727in;height:2.41295in" /><img src="thesis_media/media/image68.png" style="width:3.32407in;height:2.49306in" />

Fig:4.44 Training vs Validation Loss and Accuracy GloVe+BiLSTM Full dataset

<img src="thesis_media/media/image69.png" style="width:3.30417in;height:2.20278in" /> <img src="thesis_media/media/image70.png" style="width:3.25175in;height:2.16783in" />

Fig:4.45 Training vs Validation Loss and Accuracy Word2Vec+BiLSTM Full dataset

<img src="thesis_media/media/image71.png" style="width:3.23889in;height:2.42917in" /><img src="thesis_media/media/image72.png" style="width:3.32361in;height:2.49271in" />

Fig:4.46 Training vs Validation Loss and Accuracy FastText +BiLSTM Full dataset

<img src="thesis_media/media/image73.png" style="width:3.02797in;height:2.26424in" /> <img src="thesis_media/media/image74.png" style="width:3.03497in;height:2.26983in" />

Fig:4.47 ROC-AUC Curve MuRIL Fig:4.48 ROC-AUC Curve mBART

<img src="thesis_media/media/image75.png" style="width:3.07692in;height:2.27583in" /><img src="thesis_media/media/image76.png" style="width:3.00699in;height:2.22326in" />

Fig 4.49 ROC-AUC Curve HingRoBERTa Fig: 4.50 ROC-AUC Curve MPNet

<img src="thesis_media/media/image77.png" style="width:3.41259in;height:2.27506in" /><img src="thesis_media/media/image78.png" style="width:3.12587in;height:2.34441in" />

Fig:4.51 ROC-AUC Curve GloVe+BiLSTM Fig:4.52 ROC-AUC Curve Word2Vec+BiLSTM

<img src="thesis_media/media/image79.png" style="width:3.36364in;height:2.52273in" />

Fig:4.53 ROC-AUC Curve FastText +BiLSTM

<img src="thesis_media/media/image80.png" style="width:3.29371in;height:2.62563in" /> <img src="thesis_media/media/image81.png" style="width:3.2028in;height:2.55349in" />

Fig:4.54 Confusion Matrix MuRIL Fig:4.55 Confusion Matrix mBART

<img src="thesis_media/media/image82.png" style="width:3.45764in;height:2.7125in" /><img src="thesis_media/media/image83.png" style="width:3.24606in;height:2.62284in" />

Fig 4.56: Confusion Matrix HingRoBERTa Fig:4.57 Confusion Matrix MPNet

<img src="thesis_media/media/image84.png" style="width:3.01458in;height:2.5125in" /><img src="thesis_media/media/image85.png" style="width:3.21678in;height:2.41259in" />

Fig:4.58 Confusion Matrix GloVe + BiLSTM Fig:4.58 Confusion Matrix Word2Vec + BiLSTM

<img src="thesis_media/media/image86.png" style="width:3.39583in;height:2.54688in" />

Fig:4.59 Confusion Matrix FastText + BiLSTM

## **4.2.3 MULTI-STAGE LANGUAGE TRAINING**

This work involves total 13 figures for visualisation that uses multiple variations of hybrid models (GloVe+BiLSTM, FastText+BiLSTM, Word2Vec+BiLSTM) and deep learning models (Mbart, MuRIL, HingRoBERTa-mixed, MPNet). For multistage language training on English-\> Hinglish -\> Hindi -\> All combined, and we have used only all combined(Full dataset) visualisations Fig:4.60 – Fig: 4.66 consists of Train vs Validation Accuracy Curve and Train vs Validation Loss Curve, Fig:4.67 – Fig: 4.73 consists of AUC-ROC Curves.

<img src="thesis_media/media/image87.png" style="width:3.16783in;height:2.33624in" /><img src="thesis_media/media/image88.png" style="width:3.12587in;height:2.34252in" />

Fig:4.60 Training vs Validation Loss and Accuracy MuRIL Full dataset

<img src="thesis_media/media/image89.png" style="width:3.26573in;height:2.31164in" /><img src="thesis_media/media/image90.png" style="width:3.18182in;height:2.35762in" />

Fig:4.61 Training vs Validation Loss and Accuracy mBART Full dataset

<img src="thesis_media/media/image91.png" style="width:3.26528in;height:2.27365in" /><img src="thesis_media/media/image92.png" style="width:3.28671in;height:2.31297in" />

Fig:4.62 Training vs Validation Loss and Accuracy HingRoBERTa Full dataset

<img src="thesis_media/media/image93.png" style="width:3.11888in;height:2.23829in" /><img src="thesis_media/media/image94.png" style="width:3.18125in;height:2.2279in" />

Fig:4.63 Training vs Validation Loss and Accuracy MPNet Full dataset

<img src="thesis_media/media/image95.png" style="width:3.27972in;height:2.23462in" /> <img src="thesis_media/media/image96.png" style="width:3.3986in;height:2.29075in" />

Fig:4.64 Training vs Validation Loss and Accuracy GloVe + BiLSTM Full dataset

<img src="thesis_media/media/image97.png" style="width:3.23077in;height:2.19592in" /> <img src="thesis_media/media/image98.png" style="width:3.41958in;height:2.27404in" />

Fig:4.65 Training vs Validation Loss and Accuracy Word2Vec + BiLSTM Full dataset

<img src="thesis_media/media/image99.png" style="width:3.04895in;height:2.05446in" /><img src="thesis_media/media/image100.png" style="width:3.16469in;height:2.10577in" />

Fig:4.66 Training vs Validation Loss and Accuracy FastText + BiLSTM Full dataset

<img src="thesis_media/media/image101.png" style="width:3.07692in;height:2.54811in" /><img src="thesis_media/media/image102.png" style="width:3.12587in;height:2.53985in" />

Fig:4.67 ROC-AUC Curve MuRIL Full Dataset Fig:4.68 ROC-AUC Curve mBART Full Dataset

<img src="thesis_media/media/image103.png" style="width:3.46154in;height:2.46232in" /><img src="thesis_media/media/image104.png" style="width:3.02797in;height:2.55086in" />

Fig:4.69 ROC-AUC Curve HingRoBERTa Full Dataset Fig: 4.70 ROC-AUC Curve MPNet Full Dataset

<img src="thesis_media/media/image105.png" style="width:3.41944in;height:2.38472in" /><img src="thesis_media/media/image106.png" style="width:3.38894in;height:2.2856in" />

Fig4.71: ROC-AUC Curve GloVe + BiLSTM Fig:4.72 ROC-AUC Curve Word2Vec + BiLSTM

<img src="thesis_media/media/image107.png" style="width:3.35664in;height:2.28167in" />

Fig:4.73 ROC-AUC Curve FastText + BiLSTM Full Dataset

<img src="thesis_media/media/image108.png" style="width:3.26812in;height:2.70922in" /> <img src="thesis_media/media/image109.png" style="width:3.23932in;height:2.7029in" />

Fig:4.74 Confusion Matrix MuRIL Fig:4.75 Confusion Matrix mBART

<img src="thesis_media/media/image110.png" style="width:3.09643in;height:2.68653in" /> <img src="thesis_media/media/image111.png" style="width:3.34108in;height:2.45652in" />

Fig:4.76 Confusion Matrix HingRoBERTa Fig:4.77 Confusion Matrix MPNet

<img src="thesis_media/media/image112.png" style="width:3.09583in;height:2.85069in" /> <img src="thesis_media/media/image113.png" style="width:3.07659in;height:2.82464in" />

Fig:4.78 Confusion Matrix Word2Vec + BiLSTM Fig:4.79 Confusion Matrix FastText + BiLSTM

<img src="thesis_media/media/image114.png" style="width:3.70498in;height:3.42424in" />

Fig:4.80 Confusion Matrix GloVe + BiLSTM

<img src="thesis_media/media/image115.png" style="width:4.99624in;height:3.23021in" />

Fig: 4.81 Confusion Matrix for Sarvam \*

Β 

## **4.3 DISCUSSION**

## 

## The overall workflow consists of three methodologies in total, one with regular fine-tuning for hybrid models, second methodology involves of two different strategies one with regular fine-tuning and one with sequential training (English-\> Hinglish-\> Hindi-\> Full) on both hybrid models and deep learning models and the third methodology involves of multiple variations sequential training (English, Hinglish, Hindi, all combined), (English, Hindi, Hinglish, all combined ), (Hinglish, English, Hindi, all combined), (Hinglish, Hindi, English, all combined), (Hindi, Hinglish, English, all combined), and (Hindi, English, Hinglish, all combined) corpora to improve cross-lingual understanding. So, in our entire workflow we observed that third methodology took more training time in comparison to other two methodologies and gained very good accuracy and f1-score overall which was achieved by GloVe+BiLSTM model which is a classical hybrid model but performed better than deep learning models we used. But, in regular trainings deep learning models outperformed every other hybrid models. We have also observed Hindi(devnagiri) made overall performance degrade due to its less contextual understanding and tokens mishandle. For an extra activity we worked with Sarvam AI which is a LLM and we achieved highest of all metrics in our all workflows combined.

## 

## 

## 

## **CHAPTER 5**

**CONCLUSION AND FUTURE SCOPE**

This project successfully addressed the critical need for an effective and domain-specific Code-mixed sentiment analysis system, which is an obvious unexplored area in the field of natural language processing and speech technology. The outcomes of this project open up several promising opportunities for future research and development in better contextual understandings and natural language processing for code-mixed. One of the most immediate areas of expansion is the enhancement of the dataset. Our primary aim was that a good LLM or a model with a lot of billions of parameter can handle, can understand the sarcastic contexts, but the computational cost of these models are very height. We tried to make a small differentiate with developing a sentimental analysis prediction system with those models who has less parameters and achieve a near good prediction with good performance our overall works has gained the accuracy between 72%-80% which is a good start but we will move further and works with different collections of architectures to understand the depth of mechanism for enhancements. In our study we were also introduced with Explainable AI which is a modern trend model that tends to discover the happenings inside the black box. We have worked with some types of XAI that includes Shap, Lime, Captum, Integrated Gradients and all these are good powerful models. Our future scope is to study more on hybrid models as they somehow manage to perform well then deep learning models if selected smartly and we will combine the explainable AI in order to for analyzing and handling of misclassifications happening in training.

**REFERENCE**

\[1\] Singh, G. (2021). Sentiment analysis of code-mixed social media text (Hinglish). In *arXiv \[cs.CL\]*. https://doi.org/10.48550/ARXIV.2102.12149\
\[2\] Thakur, V., Sahu, R., & Omer, S. (2020). Current state of hinglish text sentiment analysis. *SSRN Electronic Journal*. <https://doi.org/10.2139/ssrn.3614442>

\[3\] Agarwal, P. N. (2024). Improving sentiment analysis accuracy in hinglish text using hybrid deep learning approaches. *Educational Administration: Theory and Practice*, 741–750. https://doi.org/10.53555/kuey.v30i11.8739

\[4\] Singh, G. V., Ghosh, S., Firdaus, M., Ekbal, A., & Bhattacharyya, P. (2024). Predicting multi-label emojis, emotions, and sentiments in code-mixed texts using an emojifying sentiments framework. *Scientific Reports*, *14*(1), 12204. https://doi.org/10.1038/s41598-024-58944-5

\[5\] Himabindu, G. S. S. N., Rao, R., & Sethia, D. (2022). A self-attention hybrid emoji prediction model for code-mixed language: (Hinglish). *Social Network Analysis and Mining*, *12*(1). https://doi.org/10.1007/s13278-022-00961-1

\[6\] Yadav, S., Kaushik, A., & McDaid, K. (2024). Leveraging weakly annotated data for hate speech detection in code-mixed Hinglish: A feasibility-driven transfer learning approach with Large Language Models. In *arXiv \[cs.CL\]*. http://arxiv.org/abs/2403.02121

\[7\] Aggarwal, A., Wadhawan, A., Chaudhary, A., & Maurya, K. (2020). β€œDid you really mean what you said?” : Sarcasm Detection in Hindi-English Code-Mixed Data using Bilingual Word Embeddings. In *arXiv \[cs.CL\]*. https://doi.org/10.48550/ARXIV.2010.00310

\[8\] Rahul, Gupta, V., Sehra, V., & Vardhan, Y. R. (2021). Ensemble based hinglish hate speech detection. *2021 5th International Conference on Intelligent Computing and Control Systems (ICICCS)*.

\[9\] Acharya, A., & Goyal, R. (2025). Ensemble learning-based sarcasm detection in hinglish tweets using Word2Vec embedding. *2025 IEEE International Conference on Interdisciplinary Approaches in Technology and Management for Social Innovation (IATMSI)*, 1–6.

\[10\] Aloria, S., Aggarwal, I., Baliyan, N., & Ghosh, M. (2023). Hilarious or hidden? DetectiSarcasmasm Hinglish Tweets Usinging BERT-GRU. *2023 14th International Conference on Computing Communication and Networking Technologies (ICCCNT)*.

\[11\] Chutia, T., Baruah, N., & Sonowal, P. (2025). A comparative study of machine learning and deep learning approaches for identifying Assamese abusive comments on social media. *Procedia Computer Science*, *258*, 981–992. https://doi.org/10.1016/j.procs.2025.04.335

\[12\] Baruah, P., Dutta, B., Sarma, S. K., & Talukdar, K. (2025). Named Entity Recognition in Assamese Language using two separate models: BiLSTM and BERT. *Procedia Computer Science*, *258*, 242–251.Β  https://doi.org/10.1016/j.procs.2025.04.262

\[13\] Lalthangmawii, M., & Singh, T. D. (2025). Sentiment analysis of Mizo using lexical features in low resource based models. *Natural Language Processing Journal*, *13*(100181), 100181. https://doi.org/10.1016/j.nlp.2025.100181

\[14\] Talukdar, M., & Sarma, S. (2024). Bidirectional LSTM-based sentiment analysis for Assamese text. *American Journal of Computer Science and Technology*, *7*(2), 29–37. <https://doi.org/10.11648/j.ajcst.20240702.11>

\[15\] Jayanthi, S. M., Nerella, K., Chandu, K. R., & Black, A. W. (2021). CodemixedNLP: An Extensible and Open NLP Toolkit for Code-Mixing. In *arXiv \[cs.CL\]*.\
<https://doi.org/10.48550/ARXIV.2106.06004>

\[16\] Dhekane, S. (2025, August). *Code-Mixed Hinglish Hate Speech Detection Dataset*. Kaggle.com. <https://www.kaggle.com/datasets/sharduldhekane/code-mixed-hinglish-hate-speech-detection-dataset>

\[17\] Paul, K., Wankhade, M., & Dutta, S. C. (2025). Dynamic multi-attention fusion for joint intent detection and slot filling in code-mixed language understanding. 2025 6th International Conference on Recent Advances in Information Technology (RAIT), 1–6.

\[18\] Tho, C., Warnars, H. L. H. S., Soewito, B., & Gaol, F. L. (2020). Code-mixed sentiment analysis using machine learning approach – A systematic literature review. 2020 4th International Conference on Informatics and Computational Sciences (ICICoS), 1–6.

\[19\] Srivastava, V., & Singh, M. (2021). Challenges and considerations with code-mixed NLP for multilingual societies. In arXiv \[cs.CL\]. http://arxiv.org/abs/2106.07823<u>\
</u>