File size: 128,939 Bytes
65cc963 | 1 2 3 4 5 6 7 8 9 10 11 12 13 14 15 16 17 18 19 20 21 22 23 24 25 26 27 28 29 30 31 32 33 34 35 36 37 38 39 40 41 42 43 44 45 46 47 48 49 50 51 52 53 54 55 56 57 58 59 60 61 62 63 64 65 66 67 68 69 70 71 72 73 74 75 76 77 78 79 80 81 82 83 84 85 86 87 88 89 90 91 92 93 94 95 96 97 98 99 100 101 102 103 104 105 106 107 108 109 110 111 112 113 114 115 116 117 118 119 120 121 122 123 124 125 126 127 128 129 130 131 132 133 134 135 136 137 138 139 140 141 142 143 144 145 146 147 148 149 150 151 152 153 154 155 156 157 158 159 160 161 162 163 164 165 166 167 168 169 170 171 172 173 174 175 176 177 178 179 180 181 182 183 184 185 186 187 188 189 190 191 192 193 194 195 196 197 198 199 200 201 202 203 204 205 206 207 208 209 210 211 212 213 214 215 216 217 218 219 220 221 222 223 224 225 226 227 228 229 230 231 232 233 234 235 236 237 238 239 240 241 242 243 244 245 246 247 248 249 250 251 252 253 254 255 256 257 258 259 260 261 262 263 264 265 266 267 268 269 270 271 272 273 274 275 276 277 278 279 280 281 282 283 284 285 286 287 288 289 290 291 292 293 294 295 296 297 298 299 300 301 302 303 304 305 306 307 308 309 310 311 312 313 314 315 316 317 318 319 320 321 322 323 324 325 326 327 328 329 330 331 332 333 334 335 336 337 338 339 340 341 342 343 344 345 346 347 348 349 350 351 352 353 354 355 356 357 358 359 360 361 362 363 364 365 366 367 368 369 370 371 372 373 374 375 376 377 378 379 380 381 382 383 384 385 386 387 388 389 390 391 392 393 394 395 396 397 398 399 400 401 402 403 404 405 406 407 408 409 410 411 412 413 414 415 416 417 418 419 420 421 422 423 424 425 426 427 428 429 430 431 432 433 434 435 436 437 438 439 440 441 442 443 444 445 446 447 448 449 450 451 452 453 454 455 456 457 458 459 460 461 462 463 464 465 466 467 468 469 470 471 472 473 474 475 476 477 478 479 480 481 482 483 484 485 486 487 488 489 490 491 492 493 494 495 496 497 498 499 500 501 502 503 504 505 506 507 508 509 510 511 512 513 514 515 516 517 518 519 520 521 522 523 524 525 526 527 528 529 530 531 532 533 534 535 536 537 538 539 540 541 542 543 544 545 546 547 548 549 550 551 552 553 554 555 556 557 558 559 560 561 562 563 564 565 566 567 568 569 570 571 572 573 574 575 576 577 578 579 580 581 582 583 584 585 586 587 588 589 590 591 592 593 594 595 596 597 598 599 600 601 602 603 604 605 606 607 608 609 610 611 612 613 614 615 616 617 618 619 620 621 622 623 624 625 626 627 628 629 630 631 632 633 634 635 636 637 638 639 640 641 642 643 644 645 646 647 648 649 650 651 652 653 654 655 656 657 658 659 660 661 662 663 664 665 666 667 668 669 670 671 672 673 674 675 676 677 678 679 680 681 682 683 684 685 686 687 688 689 690 691 692 693 694 695 696 697 698 699 700 701 702 703 704 705 706 707 708 709 710 711 712 713 714 715 716 717 718 719 720 721 722 723 724 725 726 727 728 729 730 731 732 733 734 735 736 737 738 739 740 741 742 743 744 745 746 747 748 749 750 751 752 753 754 755 756 757 758 759 760 761 762 763 764 765 766 767 768 769 770 771 772 773 774 775 776 777 778 779 780 781 782 783 784 785 786 787 788 789 790 791 792 793 794 795 796 797 798 799 800 801 802 803 804 805 806 807 808 809 810 811 812 813 814 815 816 817 818 819 820 821 822 823 824 825 826 827 828 829 830 831 832 833 834 835 836 837 838 839 840 841 842 843 844 845 846 847 848 849 850 851 852 853 854 855 856 857 858 859 860 861 862 863 864 865 866 867 868 869 870 871 872 873 874 875 876 877 878 879 880 881 882 883 884 885 886 887 888 889 890 891 892 893 894 895 896 897 898 899 900 901 902 903 904 905 906 907 908 909 910 911 912 913 914 915 916 917 918 919 920 921 922 923 924 925 926 927 928 929 930 931 932 933 934 935 936 937 938 939 940 941 942 943 944 945 946 947 948 949 950 951 952 953 954 955 956 957 958 959 960 961 962 963 964 965 966 967 968 969 970 971 972 973 974 975 976 977 978 979 980 981 982 983 984 985 986 987 988 989 990 991 992 993 994 995 996 997 998 999 1000 1001 1002 1003 1004 1005 1006 1007 1008 1009 1010 1011 1012 1013 1014 1015 1016 1017 1018 1019 1020 1021 1022 1023 1024 1025 1026 1027 1028 1029 1030 1031 1032 1033 1034 1035 1036 1037 1038 1039 1040 1041 1042 1043 1044 1045 1046 1047 1048 1049 1050 1051 1052 1053 1054 1055 1056 1057 1058 1059 1060 1061 1062 1063 1064 1065 1066 1067 1068 1069 1070 1071 1072 1073 1074 1075 1076 1077 1078 1079 1080 1081 1082 1083 1084 1085 1086 1087 1088 1089 1090 1091 1092 1093 1094 1095 1096 1097 1098 1099 1100 1101 1102 1103 1104 1105 1106 1107 1108 1109 1110 1111 1112 1113 1114 1115 1116 1117 1118 1119 1120 1121 1122 1123 1124 1125 1126 1127 1128 1129 1130 1131 1132 1133 1134 1135 1136 1137 1138 1139 1140 1141 1142 1143 1144 1145 1146 1147 1148 1149 1150 1151 1152 1153 1154 1155 1156 1157 1158 1159 1160 1161 1162 1163 1164 1165 1166 1167 1168 1169 1170 1171 1172 1173 1174 1175 1176 1177 1178 1179 1180 1181 1182 1183 1184 1185 1186 1187 1188 1189 1190 1191 1192 1193 1194 1195 1196 1197 1198 1199 1200 1201 1202 1203 1204 1205 1206 1207 1208 1209 1210 1211 1212 1213 1214 1215 1216 1217 1218 1219 1220 1221 1222 1223 1224 1225 1226 1227 1228 1229 1230 1231 1232 1233 1234 1235 1236 1237 1238 1239 1240 1241 1242 1243 1244 1245 1246 1247 1248 1249 1250 1251 1252 1253 1254 1255 1256 1257 1258 1259 1260 1261 1262 1263 1264 1265 1266 1267 1268 1269 1270 1271 1272 1273 1274 1275 1276 1277 1278 1279 1280 1281 1282 1283 1284 1285 1286 1287 1288 1289 1290 1291 1292 1293 1294 1295 1296 1297 1298 1299 1300 1301 1302 1303 1304 1305 1306 1307 1308 1309 1310 1311 1312 1313 1314 1315 1316 1317 1318 1319 1320 1321 1322 1323 1324 1325 1326 1327 1328 1329 1330 1331 1332 1333 1334 1335 1336 1337 1338 1339 1340 1341 1342 1343 1344 1345 1346 1347 1348 1349 1350 1351 1352 1353 1354 1355 1356 1357 1358 1359 1360 1361 1362 1363 1364 1365 1366 1367 1368 1369 1370 1371 1372 1373 1374 1375 1376 1377 1378 1379 1380 1381 1382 1383 1384 1385 1386 1387 1388 1389 1390 1391 1392 1393 1394 1395 1396 1397 1398 1399 1400 1401 1402 1403 1404 1405 1406 1407 1408 1409 1410 1411 1412 1413 1414 1415 1416 1417 1418 1419 1420 1421 1422 1423 1424 1425 1426 1427 1428 1429 1430 1431 1432 1433 1434 1435 1436 1437 1438 1439 1440 1441 1442 1443 1444 1445 1446 1447 1448 1449 1450 1451 1452 1453 1454 1455 1456 1457 1458 1459 1460 1461 1462 1463 1464 1465 1466 1467 1468 1469 1470 1471 1472 1473 1474 1475 1476 1477 1478 1479 1480 1481 1482 1483 1484 1485 1486 1487 1488 1489 1490 1491 1492 1493 1494 1495 1496 1497 1498 1499 1500 1501 1502 1503 1504 1505 1506 1507 1508 1509 1510 1511 1512 1513 1514 1515 1516 1517 1518 1519 1520 1521 1522 1523 1524 1525 1526 1527 1528 1529 1530 1531 1532 1533 1534 1535 1536 1537 1538 1539 1540 1541 1542 1543 1544 1545 1546 1547 1548 1549 1550 1551 1552 1553 1554 1555 1556 1557 1558 1559 1560 1561 1562 1563 1564 1565 1566 1567 1568 1569 1570 1571 1572 1573 1574 1575 1576 1577 1578 1579 1580 1581 1582 1583 1584 1585 1586 1587 1588 1589 1590 1591 1592 1593 1594 1595 1596 1597 1598 1599 1600 1601 1602 1603 1604 1605 1606 1607 1608 1609 1610 1611 1612 1613 1614 1615 1616 1617 1618 1619 1620 1621 1622 1623 1624 1625 1626 1627 1628 1629 1630 1631 1632 1633 1634 1635 1636 1637 1638 1639 1640 1641 1642 1643 1644 1645 1646 1647 1648 1649 1650 1651 1652 1653 1654 1655 1656 1657 1658 1659 1660 1661 1662 1663 1664 1665 1666 1667 1668 1669 1670 1671 1672 1673 1674 1675 1676 1677 1678 1679 1680 1681 1682 1683 1684 1685 1686 1687 1688 1689 1690 1691 1692 1693 1694 1695 1696 1697 1698 1699 1700 1701 1702 1703 1704 1705 1706 1707 1708 1709 1710 1711 1712 1713 1714 1715 1716 1717 1718 1719 1720 1721 1722 1723 1724 1725 1726 1727 1728 1729 1730 1731 1732 1733 1734 1735 1736 1737 1738 1739 1740 1741 1742 1743 1744 1745 1746 1747 1748 1749 1750 1751 1752 1753 1754 1755 1756 1757 1758 1759 1760 1761 1762 1763 1764 1765 1766 1767 1768 1769 1770 1771 1772 1773 1774 1775 1776 1777 1778 1779 1780 1781 1782 1783 1784 1785 1786 1787 1788 1789 1790 1791 1792 1793 1794 1795 1796 1797 1798 1799 1800 1801 1802 1803 1804 1805 1806 1807 1808 1809 1810 1811 1812 1813 1814 1815 1816 1817 1818 1819 1820 1821 1822 1823 1824 1825 1826 1827 1828 1829 1830 1831 1832 1833 1834 1835 1836 1837 1838 1839 1840 1841 1842 1843 1844 1845 1846 1847 1848 1849 1850 1851 1852 1853 1854 1855 1856 1857 1858 1859 1860 1861 1862 1863 1864 1865 1866 1867 1868 1869 1870 1871 1872 1873 1874 1875 1876 1877 1878 1879 1880 1881 1882 1883 1884 1885 1886 1887 1888 1889 1890 1891 1892 1893 1894 1895 1896 1897 1898 1899 1900 1901 1902 1903 1904 1905 1906 1907 1908 1909 1910 1911 1912 1913 1914 1915 1916 1917 1918 1919 1920 1921 1922 1923 1924 1925 1926 1927 1928 1929 1930 1931 1932 1933 1934 1935 1936 1937 1938 1939 1940 1941 1942 1943 1944 1945 1946 1947 1948 1949 1950 1951 1952 1953 1954 1955 1956 1957 1958 1959 1960 1961 1962 1963 1964 1965 1966 1967 1968 1969 1970 1971 1972 1973 1974 1975 1976 1977 1978 1979 1980 1981 1982 1983 1984 1985 1986 1987 1988 1989 1990 1991 1992 1993 1994 1995 1996 1997 1998 1999 2000 2001 2002 2003 2004 2005 2006 2007 2008 2009 2010 2011 2012 2013 2014 2015 2016 2017 2018 2019 2020 2021 2022 2023 2024 2025 2026 2027 2028 2029 2030 2031 2032 2033 2034 2035 2036 2037 2038 2039 2040 2041 2042 2043 2044 2045 2046 2047 2048 2049 2050 2051 2052 2053 2054 2055 2056 2057 2058 2059 2060 2061 2062 2063 2064 2065 2066 2067 2068 2069 2070 2071 2072 2073 2074 2075 2076 2077 2078 2079 2080 2081 2082 2083 2084 2085 2086 2087 2088 2089 2090 2091 2092 2093 2094 2095 2096 2097 2098 2099 2100 2101 2102 2103 2104 2105 2106 2107 2108 2109 2110 2111 2112 2113 2114 2115 2116 2117 2118 2119 2120 2121 2122 2123 2124 2125 2126 2127 2128 2129 2130 2131 2132 2133 2134 2135 2136 2137 2138 2139 2140 2141 2142 2143 2144 2145 2146 2147 2148 2149 2150 2151 2152 2153 2154 2155 2156 2157 2158 2159 2160 2161 2162 2163 2164 2165 2166 2167 2168 2169 2170 2171 2172 2173 2174 2175 2176 2177 2178 2179 2180 2181 2182 2183 2184 2185 2186 2187 2188 2189 2190 2191 2192 2193 2194 2195 2196 2197 2198 2199 2200 2201 2202 2203 2204 2205 2206 2207 2208 2209 2210 2211 2212 2213 2214 2215 2216 2217 2218 2219 2220 2221 2222 2223 2224 2225 2226 2227 2228 2229 2230 2231 2232 2233 2234 2235 2236 2237 2238 2239 2240 2241 2242 2243 2244 2245 2246 2247 2248 2249 2250 2251 2252 2253 2254 2255 2256 2257 2258 2259 2260 2261 2262 2263 2264 2265 2266 2267 2268 2269 2270 2271 2272 2273 2274 2275 2276 2277 2278 2279 2280 2281 2282 2283 2284 2285 2286 2287 2288 2289 2290 2291 2292 2293 2294 2295 2296 2297 2298 2299 2300 2301 2302 2303 2304 2305 2306 2307 2308 2309 2310 2311 2312 2313 2314 2315 2316 2317 2318 2319 2320 2321 2322 2323 2324 2325 2326 2327 2328 2329 2330 2331 2332 2333 2334 2335 2336 2337 2338 2339 2340 2341 2342 2343 2344 2345 2346 2347 2348 2349 2350 2351 2352 2353 2354 2355 2356 2357 2358 2359 2360 2361 2362 2363 2364 2365 2366 2367 2368 2369 2370 2371 2372 2373 2374 2375 2376 2377 2378 2379 2380 2381 2382 2383 2384 2385 2386 2387 2388 2389 2390 2391 2392 2393 2394 2395 2396 2397 2398 2399 2400 2401 2402 2403 2404 2405 2406 2407 2408 2409 2410 2411 2412 2413 2414 2415 2416 2417 2418 2419 2420 2421 2422 2423 2424 2425 2426 2427 2428 2429 2430 2431 2432 2433 2434 2435 2436 2437 2438 2439 2440 2441 2442 2443 2444 2445 2446 2447 2448 2449 2450 2451 2452 2453 2454 2455 2456 2457 2458 2459 2460 2461 2462 2463 2464 2465 2466 2467 2468 2469 2470 2471 2472 2473 2474 2475 2476 2477 2478 2479 2480 2481 2482 2483 2484 2485 2486 2487 2488 2489 2490 2491 2492 2493 2494 2495 2496 2497 2498 2499 2500 2501 2502 2503 2504 2505 2506 2507 2508 2509 2510 2511 2512 2513 2514 2515 2516 2517 2518 2519 2520 2521 2522 2523 2524 2525 2526 2527 2528 2529 2530 2531 2532 2533 2534 2535 2536 2537 2538 2539 2540 2541 2542 2543 2544 2545 2546 2547 2548 2549 2550 2551 2552 2553 2554 2555 2556 2557 2558 2559 2560 2561 2562 2563 2564 2565 2566 2567 2568 2569 2570 2571 2572 2573 2574 2575 2576 2577 2578 2579 2580 2581 2582 2583 2584 2585 2586 2587 2588 2589 2590 2591 2592 2593 2594 2595 2596 2597 2598 2599 2600 2601 2602 2603 2604 2605 2606 2607 2608 2609 2610 2611 2612 2613 2614 2615 2616 2617 2618 2619 2620 2621 2622 2623 2624 2625 2626 2627 2628 2629 2630 2631 2632 2633 2634 2635 2636 2637 2638 2639 2640 2641 2642 2643 2644 2645 2646 2647 2648 2649 2650 2651 2652 2653 2654 2655 2656 2657 2658 2659 2660 2661 2662 2663 2664 2665 2666 2667 2668 2669 2670 2671 2672 2673 2674 2675 2676 2677 2678 2679 2680 2681 2682 2683 2684 2685 2686 2687 2688 2689 2690 2691 2692 2693 2694 2695 2696 2697 2698 2699 2700 2701 2702 2703 2704 2705 2706 2707 2708 2709 2710 2711 2712 2713 2714 2715 2716 2717 2718 2719 2720 | # Full Group Thesis (Reference)
**Download:** [PDF](FinalThesisSP.pdf) | [Word source (.docx)](FinalThesisSP.docx)
> This is the complete group thesis *Developing a Sentiment Analysis Model for Code-Mixed
> Hindi-English (Hinglish) Text*, included as reference. It covers all 17 models; the transformer
> (MuRIL, mBART, HingRoBERTa, MPNet) and Sarvam LLM tracks were built by teammates.
> For **Pankaj Biswas's individual contribution** (the BiLSTM / LSTM track), see the
> [model card / report](README.md).
---
> <img src="thesis_media/media/image1.jpeg" style="width:2.92014in;height:1.03819in" alt="A logo for a university Description automatically generated" />
>
> **DEVELOPING A SENTIMENT ANALYSIS MODEL FOR CODE-MIXED HINDI-ENGLISH (HINGLISH) TEXT**
>
> A Project Report submitted
>
> In fulfillment of the requirements for the degree of
### B.Tech (Computer Science & Engineering)
> Submitted by :-
>
> **Pulakala Prithvi Raj (222025042)**
>
> **Pankaj Biswas (222025043)**
>
> **Pritisha Goswami (222025049)**
>
> **B.Tech Computer Science and Engineering**
>
> **8<sup>th</sup> Semester**
>
> **Royal School of Engineering and Technology (RSET)**
>
> Under the guidance of
>
> **Dr. Dillip Rout**
>
> **Assistant Professor**
>
> **Royal School of Engineering and Technology (RSET)**
>
> **THE ASSAM ROYAL GLOBAL UNIVERSITY**
## GUWAHATI: 781035
> **Session: 2022-2026**
>
#
# CERTIFICATE OF APPROVAL
#
> This is to certify that the project report entitled *"Developing a Sentiment Analysis Model for Code-Mixed Hindi-English (Hinglish) Text"* submitted by **Pulakala Prithvi Raj** (Roll No. 222025042), **Pankaj Biswas** (Roll No. 222025043) and **Pritisha Goswami** (Roll No. 222025049), students of B.Tech, 8th semester in the Department of Computer Science & Engineering, Royal School of Engineering and Technology (RSET), The Assam Royal Global University, Guwahati, Assam, has been completed under my supervision. This work is submitted as part of the requirements for the award of the B.Tech degree in Computer Science & Engineering and has not been submitted elsewhere for a degree.
>
> **Project Guide: Signature of the External**
>
> **Dr. Dillip Rout Name of the External**
>
> **Assistant Professor, CSE, RSET**
>
> **Date***:*
>
> **Place: Guwahati**
>
# FORWARDING CERTIFICATE
#
> This is to certify that the project report entitled *"Developing a Sentiment Analysis Model for Code-Mixed Hindi-English (Hinglish) Text"* submitted by **Pulakala Prithvi Raj** (Roll No. 222025042), **Pankaj Biswas** (Roll No. 222025043) and **Pritisha Goswami** (Roll No. 222025049), students of B.Tech, 8th semester in the Department of Computer Science & Engineering at Royal School of Engineering and Technology (RSET), The Assam Royal Global University, Guwahati, Assam, under the guidance of **Dr. Dillip Rout, Assistant Professor**, has been evaluated and deemed satisfactory for submission as a requirement for the degree program.
>
> **Date:**
>
> **Place:** Guwahati
**Dr. Dillip Rout**
**Assistant Professor**
**Department of CSE**
> **Royal School of Engineering &**
>
> **Technology**
# DECLARATION
We, **Pulakala Prithvi Raj** (Roll No. 222025042), **Pankaj Biswas** (Roll No. 222025043) and **Pritisha Goswami** (Roll No. 222025049), hereby declare that the project work entitled *"Developing a Sentiment Analysis Model for Code-Mixed Hindi-English (Hinglish) Text"* was carried out by us under the guidance and supervision of **Dr. Dillip Rout, Assistant Professor**, **Department of Computer Science & Engineering**. This project is submitted for the academic session 2022-2026. We confirm that this work, or any part of it, has not been submitted elsewhere for any other purpose to date.
### Date:
> **Place:** Guwahati
**Pulakala Prithvi Raj Pankaj Biswas Pritisha Goswami**
**222025042 222025043 222025049**
# ACKNOWLEDGMENT
#
> We are deeply grateful to Royal School of Engineering and Technology for providing us with the resources and environment needed to complete this project titled, *"Developing a Sentiment Analysis Model for Code-Mixed Hindi-English (Hinglish) Text".*
>
> We would like to extend our heartfelt thanks to our guide, Dr. Dillip Rout, Assistant Professor, Department of CSE, Royal School of Engineering and Technology, whose invaluable guidance, encouragement, and insightful feedback have been crucial throughout the project's development. His support enabled us to navigate complex challenges and explore new dimensions in the field of deep learning.
>
> We are also grateful to all faculty members who offered their support, advice, and assistance, both directly and indirectly. Their guidance has played an essential role in shaping this work.
>
> I wish to express my gratitude to my family and friends, whose constant support, encouragement, and patience have been a source of strength throughout this journey. Without their belief in my abilities, this work would not have been possible.
>
> Thank you
**Pulakala Prithvi Raj Pankaj Biswas Pritisha Goswami**
**222025042 222025043 222025049**
# ABSTRACT
# Code-mixed languages, such as Hinglishβan informal blend of Hindi and Englishβpose significant challenges for sentiment analysis due to inconsistent grammar, transliteration variations, and limited annotated resources. This study presents a comprehensive comparison of classical machine learning and deep learning architectures, including the transformers for binary sentiment classification of Hinglish text. The analysis utilizes the PRISM dataset, comprising 29,550 Hinglish samples labeled as non-hate (0) or hate (1) sources from Kaggle. The text preprocessing included removing URLs, eliminating mentions and hashtags, and normalizing whitespace. Four models were implemented: MuRIL, GloVe+BiLSTM, FastText, and Word2Vec+Logistic Regression. Evaluation metrics include Accuracy, Precision, Recall, F1-score, Specificity, and AUC-ROC to assess the robustness of the models. Experimental results indicate that MuRIL achieved the highest F1 Score (0.737) and AUC (0.824), highlighting the efficacy of multilingual transformers for modeling code-mixed text. Classical models performed worse, though FastText outperformed the Word2Vec and GloVe baselines. The proposed model also surpasses the projects available on Kaggle. The findings emphasize the importance of multiple model evaluations for robust sentiment classification of low-resource, code-mixed social media data.
#
#
#
#
#
#
# TABLE OF CONTENTS
**Page No**.
> **Certificate of Approval** i
[Forwarding Certificate ii](#forwarding-certificate)
[Declaration iii](#declaration)
Acknowledgement iv
[Abstract](#abstract) v
[List of Tables vi](#list-of-tables)
[Chapter 1. Introduction 1- 4](#_TOC_250042)
1. Background Study 2
2. [Problem Statement 2](#problem-statement)
3. [Motivation 3](#motivation)
4. Objective 3
5. [Contributions 4](#_TOC_250039)
Chapter 2. Literature Survey 5-17
1. [Related work 6](#_TOC_250038)
[Chapter 3. Methodology 18-3](#section-15)2
1. [Introduction 19](#methodology)
2. [Block Diagram 2](#the-overall-workflow-involves-of-three-proposed-methodologies-that-is-illustrated-in-fig.-3.1-fig.3.2-and-fig3.3-which-provides-a-high-level-view-of-the-data-preprocessing-feature-extraction-model-training-and-evaluation-process.-it-visually-summarizes-the-pipeline-from-raw-dataset-input-to-performance-evaluation-across-mentioned-models.-the-project-workflow-begins-with-the-primary-input-which-is-raw-sentiment-data-of-english-in-latin-text-hindi-in-devnagiri-texts-and-often-consisting-of-code-mixed-sentences-hindi-english-hinglish-in-latin-text.-this-dataset-is-unstructured-and-not-immediately-suitable-for-tokenization-and-further-word-embeddings.-therefore-the-first-crucial-step-is-text-pre-processing-that-involves-url-removal-lowercasing-whitespace-removal-tokenization-and-embeddings.-this-text-pre-processing-is-critical-because-deep-learning-models-such-as-bert-transformer-models-assign-different-vectors-for-the-same-word-with-different-casings-to-solve-this-we-use-lowercasing-urls-which-do-not-provide-any-context-so-we-remove-urls-and-space-normalization-as-tokens-of-extra-spaces-are-also-created.-after-preprocessing-the-dataset-is-divided-into-training-60-validation-10-and-testing-30-sets-same-for-all-methodologies.-the-processed-data-is-then-fed-into-respective-sentiment-analysis-model-where-the-embedding-techniques-vary-according-to-the-methodology-being-implemented.)0
3. [Data Collection 2](#_TOC_250032)2
4. [Data Preprocessing 24](#_TOC_250031)
5. [Model Selection and development 25](#_TOC_250027)
1. ASR Model 25
2. [Summarization Model 26](#_TOC_250025)
6. Evaluation Metrics 27
> 3.6.1 WER 28
>
> 3.6.2 ROUGE Score 28
7. [Tools and Framework 29](#_TOC_250022)
8. Setting Up Environment 30
9.
[Summary 32[Chapter 4. Results and Discussion 33-4](#section-15)0](#_TOC_250015)
[4.1 Result 40](#_TOC_250015)
> [4.1.1 Regular Training 42-43](#_TOC_250015)
>
> [4.1.2 Multi-Stage Training 46-67](#_TOC_250015)
[4.2 Discussion 68](#_TOC_250015)
[[Chapter 5. Conclusion Future scope 41-43](#section-15)](#_TOC_250015)
# LIST OF TABLES
<table>
<colgroup>
<col style="width: 11%" />
<col style="width: 59%" />
<col style="width: 29%" />
</colgroup>
<thead>
<tr>
<th style="text-align: center;"><strong>Table no</strong></th>
<th><blockquote>
<p><strong>Table name</strong></p>
</blockquote></th>
<th style="text-align: left;"><strong>Page No.</strong></th>
</tr>
</thead>
<tbody>
<tr>
<td style="text-align: center;">2.1</td>
<td><blockquote>
<p>Literature Survey</p>
</blockquote></td>
<td style="text-align: left;">9-16</td>
</tr>
<tr>
<td style="text-align: center;">3.1</td>
<td>Dataset Details</td>
<td style="text-align: left;">21</td>
</tr>
<tr>
<td style="text-align: center;">3.2</td>
<td><blockquote>
<p>List Count of URL Noise</p>
</blockquote></td>
<td style="text-align: left;">23</td>
</tr>
<tr>
<td style="text-align: center;">3.3</td>
<td><blockquote>
<p>Data Metrics Before and After Cleaning</p>
</blockquote></td>
<td style="text-align: left;">23</td>
</tr>
<tr>
<td style="text-align: center;">3.4</td>
<td><blockquote>
<p>Dataset Splits Table</p>
</blockquote></td>
<td style="text-align: left;">27</td>
</tr>
<tr>
<td style="text-align: center;">3.5</td>
<td><blockquote>
<p>Model Algorithms</p>
</blockquote></td>
<td style="text-align: left;">30-32</td>
</tr>
<tr>
<td style="text-align: center;">3.6</td>
<td><blockquote>
<p>Model Hyperparameters</p>
</blockquote></td>
<td style="text-align: left;">33-36</td>
</tr>
<tr>
<td style="text-align: center;">4.1</td>
<td><blockquote>
<p>Regular Result</p>
</blockquote></td>
<td style="text-align: left;">40</td>
</tr>
<tr>
<td style="text-align: center;">4.2</td>
<td><blockquote>
<p>English Strategy for Language-wise Regular Training</p>
</blockquote></td>
<td style="text-align: left;">41</td>
</tr>
<tr>
<td style="text-align: center;">4.3</td>
<td><blockquote>
<p>Hindi Strategy for Language-wise Regular Training</p>
</blockquote></td>
<td style="text-align: left;">41</td>
</tr>
<tr>
<td style="text-align: center;">4.4</td>
<td><blockquote>
<p>Hinglish Strategy for Language-wise Regular Training</p>
</blockquote></td>
<td style="text-align: left;">42</td>
</tr>
<tr>
<td style="text-align: center;">4.5</td>
<td><blockquote>
<p>Combined Strategy for Language-wise Regular Training</p>
</blockquote></td>
<td style="text-align: left;">43</td>
</tr>
<tr>
<td style="text-align: center;">4.6</td>
<td><blockquote>
<p>English Strategy for Multistage Language Training</p>
</blockquote></td>
<td style="text-align: left;">43</td>
</tr>
<tr>
<td style="text-align: center;">4.7</td>
<td><blockquote>
<p>Hindi Strategy for Multistage Language Training</p>
</blockquote></td>
<td style="text-align: left;">43</td>
</tr>
<tr>
<td style="text-align: center;">4.8</td>
<td><blockquote>
<p>Hinglish Strategy for Multistage Language Training</p>
</blockquote></td>
<td style="text-align: left;">44</td>
</tr>
<tr>
<td style="text-align: center;">4.9</td>
<td><blockquote>
<p>Combined Strategy for Multistage Language Training</p>
</blockquote></td>
<td style="text-align: left;">45</td>
</tr>
<tr>
<td style="text-align: center;">4.10</td>
<td><blockquote>
<p>Combined Strategy for Multistage Language Taraining</p>
</blockquote></td>
<td style="text-align: left;">45-47</td>
</tr>
<tr>
<td style="text-align: center;">4.11</td>
<td><blockquote>
<p>Performance Evaluation for Sarvam Model</p>
</blockquote></td>
<td style="text-align: left;">47</td>
</tr>
</tbody>
</table>
# LIST OF FIGURES
<table style="width:99%;">
<colgroup>
<col style="width: 11%" />
<col style="width: 70%" />
<col style="width: 16%" />
</colgroup>
<thead>
<tr>
<th><strong>Fig no</strong></th>
<th><blockquote>
<p><strong>Fig Name</strong></p>
</blockquote></th>
<th style="text-align: left;"><strong>Page no</strong></th>
</tr>
</thead>
<tbody>
<tr>
<td>3.1-3.2</td>
<td><blockquote>
<p>Workflow of the proposed methodology for sentiment analysis</p>
</blockquote></td>
<td style="text-align: left;"><blockquote>
<p>19-21</p>
</blockquote></td>
</tr>
<tr>
<td>3.4</td>
<td><blockquote>
<p>Hate vs Non-Hate Class distribution</p>
</blockquote></td>
<td style="text-align: left;"><blockquote>
<p>24</p>
</blockquote></td>
</tr>
<tr>
<td>3.5</td>
<td><blockquote>
<p>Language Distribution</p>
</blockquote></td>
<td style="text-align: left;"><blockquote>
<p>25</p>
</blockquote></td>
</tr>
<tr>
<td>3.6</td>
<td><blockquote>
<p>Feature Correlation Heatmap</p>
</blockquote></td>
<td style="text-align: left;"><blockquote>
<p>25</p>
</blockquote></td>
</tr>
<tr>
<td>3.7</td>
<td><blockquote>
<p>Text length Statistics and Word count Statistics</p>
</blockquote></td>
<td style="text-align: left;"><blockquote>
<p>26</p>
</blockquote></td>
</tr>
<tr>
<td>4.1-4.10</td>
<td><blockquote>
<p>Training vs Validation Loss and Accuracy for Regular Training</p>
</blockquote></td>
<td style="text-align: left;"><blockquote>
<p>48-51</p>
</blockquote></td>
</tr>
<tr>
<td>4.11-4.20</td>
<td><blockquote>
<p>ROC-AUC Curve for Regular Training</p>
</blockquote></td>
<td style="text-align: left;"><blockquote>
<p>51-52</p>
</blockquote></td>
</tr>
<tr>
<td>4.21-4.30</td>
<td><blockquote>
<p>Confusion Matrix for Regular Training</p>
</blockquote></td>
<td style="text-align: left;"><blockquote>
<p>53-54</p>
</blockquote></td>
</tr>
<tr>
<td>4.31-4.39</td>
<td><blockquote>
<p>t-SNE for Regular Training</p>
</blockquote></td>
<td style="text-align: left;"><blockquote>
<p>54-56</p>
</blockquote></td>
</tr>
<tr>
<td>4.40-4.46</td>
<td><blockquote>
<p>Training vs Validation Loss and Accuracy for Language-wise Regular Training</p>
</blockquote></td>
<td style="text-align: left;"><blockquote>
<p>56-59</p>
</blockquote></td>
</tr>
<tr>
<td>4.47-4.53</td>
<td><blockquote>
<p>ROC-AUC Curve for Language-wise Regular Training</p>
</blockquote></td>
<td style="text-align: left;"><blockquote>
<p>59-60</p>
</blockquote></td>
</tr>
<tr>
<td>4.54-4.59</td>
<td><blockquote>
<p>Confusion Matrix for Language-wise Regular Training</p>
</blockquote></td>
<td style="text-align: left;"><blockquote>
<p>60-61</p>
</blockquote></td>
</tr>
<tr>
<td>4.60-4.67</td>
<td><blockquote>
<p>Training vs Validation Loss and Accuracy for Multistage Language Training</p>
</blockquote></td>
<td style="text-align: left;"><blockquote>
<p>62-64</p>
</blockquote></td>
</tr>
<tr>
<td>4.68-4.73</td>
<td><blockquote>
<p>ROC-AUC Curve for Multistage Language Training</p>
</blockquote></td>
<td style="text-align: left;"><blockquote>
<p>64-65</p>
</blockquote></td>
</tr>
<tr>
<td>4.74-4.80</td>
<td><blockquote>
<p>Confusion Matrix for Multistage Language Training</p>
</blockquote></td>
<td style="text-align: left;"><blockquote>
<p>66-67</p>
</blockquote></td>
</tr>
<tr>
<td>4.81</td>
<td><blockquote>
<p>Confusion Matrix for Sarvam Model</p>
</blockquote></td>
<td style="text-align: left;"><blockquote>
<p>67</p>
</blockquote></td>
</tr>
</tbody>
</table>
**CHAPTER 1**
**INTRODUCTION**
## BACKGROUND
##
Hinglish, which is a blend of English (*Latin*) and Hindi (*Latin*), mainly used in India as informal conversations, which presents unique challenges for Natural Language Processing (NLP) because of spelling variations and informal grammar, sarcastic contexts and frequent code-switching within sentences. Preliminary research was majorly focused on creating annotating the corpuses for Hinglish to assist supervised learning approaches. These datasets commonly included social media posts, its comments, and chat messages mainly informal ones showing code-mixing at the lexical and syntactic levels. Researchers investigated language identification as an initial step, distinguishing English, Hindi, and mixed tokens, which is essential for efficient sentiment analysis. Earlier studies also used Classical Machine Learning classifier models such as NaΓ―ve Bayes, Decision Trees, and SVM using features like n-grams, part-of-speech tags, and lexicons optimized for code-mixed text. Recent studies have applied deep learning models like Long Short-Term Memory (LSTM) models, Recurrent Neural Networks (RNNs) models and Transformer-based models which are more relevant for contextual understanding in comparison to Classical Machine Learning classifier models in code-mixed sentences.
Many works have also shown the challenges in Code-Mixed categories that widely included transliteration and normalization, about Hindi words written in Latin script with irregular spelling as usually there are multiple spelling variations for single Hindi word.
Various techniques such as embedding-based representations and pronunciation-based comparison have been put forward to deal with these variations effectively. Overall, the background research highlights the complexity of sentiment analysis in Hinglish due to linguistic divergence, lack of normalized spelling system, and the dynamic nature of code-switching. This has influenced the development of specialized datasets, feature extraction methods, and model architectures optimized to the distinctions of code-mixed language.
## PROBLEM STATEMENT
##
The swift growth of social media platforms has resulted in a notable rise in user contents, which is generally written in code-mixed languages that blend multiple languages within a single sentence. Hinglish, a mix of English and Hindi written in Latin script, is widely seen on platforms like Facebook, YouTube comments, Twitter comments, and Instagram comments. Analysing the sentiment of such code-mixed text is inherently difficult due to irregular grammar, inconsistencies in transliteration, spelling variations, and a lack of annotated datasets. Classical Machine Learning models, such as Word2Vec and FastText embeddings paired with linear classifiers, that provide computational efficiency but often struggle to capture the complex contextual information in code-mixed text as these types of models are usually of Static embeddings technique. Whereas, Deep Learning models, including recurrent neural networks (RNNs), bidirectional long short-term memory networks (BiLSTM), and transformer-based multilingual encoders (MuRIL), are proficient at complex contextual understandings.
##
## MOTIVATION
##
This research offers an in-depth comparative study of machine learning and deep learning techniques applied to Hinglish code-mixing for the classification of hate and non-hate speech. It leverages the publicly accessible PRISM dataset, which includes 29950 entries \[16\], for hate-speech detection. The methodology proposed in this study fills the gap in comparing multiple models using various performance metrics such as Accuracy, Balanced Accuracy, Precision, Recall, Specificity, F1 Score, and ROC-AUC Score. The research involves four models that encompass both traditional and deep learning frameworks. The goal is to pinpoint models that exhibit strong performance in binary sentiment classification of Hinglish code-mixed text and to shed light on the benefits of combining different modeling approaches. Additionally, the study aims to evaluate the transformer model without training and to assess the significance of preprocessing.
## OBJECTIVES
##
To develop and optimize NLP architectures for accurate Hinglish sentiment analysis, begin with preparing a comprehensive annotated dataset and generating multiple embedding representations, including pretrained transformer embeddings (MuRIL, mBART, HingRoBERTa, MPNet) , traditional embeddings combined with BiLSTM (Word2Vec, GloVe, FastText) and embeddings combined with LSTM (Word2Vec, Glove, FastText, Elmo, USE) . Consistent preprocessing steps such as tokenization, transliteration normalization, and language identification are essential. Establish a baseline framework using simpler models like Word2Vec+BiLSTM or GloVe+BiLSTM alongside classical classifiers such as LIGHTGBM to set reference performance metrics. Fine-tune transformer-based models on the Hinglish dataset while optimizing hyperparameters, and similarly train and tune BiLSTM layers with traditional embeddings. For LIGHTGBM, we have used embedding features (like Word2Vec,FastText,Glove,Elmo,USE)Β as input and optimize boosting the parameters.
Implement a multi-phase language-specific training strategy by splitting the dataset into English, Hindi, and Hinglish subsets, sequentially training and fine-tuning models across these phases with independent validation to monitor performance shifts. Analyze cross-language adaptation by evaluating performance consistency, knowledge transfer, and adaptation capability after each retraining phase. Following multi-phase training, assess robustness and generalization on a unified test corpus, including noisy and informal samples and varied sentiment-emotion categories. Use iterative refinement to adjust architectures, embeddings, and training strategies, considering ensemble approaches that combine transformer LSTM and BiLSTM-based models to enhance accuracy and build a robust Hinglish sentiment-emotion analysis system.
**CHAPTER 2**
**\
LITERATURE SURVEY**
**2.1 RELATED WORK**
Research in code-mixed sentiment analysis, particularly for Hinglish (HindiβEnglish mixed text), has witnessed substantial growth over the last decade, evolving from simple lexicon-based techniques to sophisticated transformer-driven architectures. Code-mixed text poses unique challenges due to transliteration variations, inconsistent grammar, and a lack of large annotated corpora. Earlier studies primarily relied on lexicons and statistical models, while recent approaches emphasize deep learning and multilingual transformer models that better capture bilingual context and semantics.
The earliest work in this domain explored classical machine learning approaches using handcrafted linguistic and statistical features. Singh analyzed code-mixed social media text using NaΓ―ve Bayes and SVM classifiers, revealing that token-level language identification and transliteration inconsistencies greatly impacted sentiment accuracy \[1\]. Similarly, Thakur et al. outlined that while traditional methods offered moderate accuracy, they lacked scalability and contextual depth \[2\]. These early models were limited in their ability to handle non-standardized language usage and failed to capture more complex syntactic relationships.
A significant shift occurred with the adoption of embedding-based representations that moved beyond sparse lexical features. Techniques such as Word2Vec and GloVe introduced dense vector representations capable of encoding semantic similarity. The study by Agarwal (2024) demonstrated that integrating CNN and BiLSTM with pretrained embeddings significantly improved sentiment detection accuracy for Hinglish text \[3\]. However, these embeddings were static and unable to account for polysemy or word sense variations, leading to limited performance in diverse contexts.
In subsequent years, deep neural architectures like RNNs and LSTMs became prominent for modeling sequential dependencies within sentences. However, these models struggled with long-term contextual understanding and required large labeled corpora for practical training. The introduction of transformer-based architectures revolutionized natural language processing by replacing sequential recurrence with self-attention, enabling parallel processing of long-range dependencies. This innovation paved the way for transformer-based contextual embeddings such as BERT and MuRIL, which excel at handling code-mixed and multilingual data.
Singh et al. presented a study on sentiments in Code-Mixed texts, validating the effectiveness of transformer-based architectures in understanding sentiment polarity and emotion intensity in bilingual texts \[4\]. Similarly, a hybrid attention-based mechanism that outperformed CNN and RNN baselines was proposed, proving that contextualized embeddings substantially improve cross-lingual generalization \[5\]. These findings confirmed that contextual modeling plays a vital role in decoding the semantics of Hinglish text, where literal translations are insufficient for accurate sentiment recognition.
The challenge of data scarcity and domain imbalance in Hinglish corpora was addressed by Yadav et al. (2024), who employed weak supervision and semi-supervised techniques to enhance dataset diversity and reduce annotation cost \[6\]. Similarly, Aggarwal et al. showcased how deep contextual encoders successfully captured sarcasm and implicit sentiment polarity in Hinglish \[7\]. Generally, traditional sentiment models often misclassify or are inefficient at capturing these aspects of sentiment analysis; however, deep learning models achieve satisfactory results. These studies highlight the evolution from sentiment-level to emotion and sarcasm-level understanding in code-mixed research.
The rise of ensemble-based architectures has further pushed the boundaries of Hinglish sentiment and hate-speech classification. Gupta et al. (2021) illustrated that integrating outputs from multiple deep learning and transformer models achieved higher recall and robustness than individual networks \[8\]. Similarly, a combination of several transformer models was deployed to capture the varied contextual nuances and linguistic cues, resulting in improved detection accuracy \[9\]. Moreover, Aloria et al. (2023) further emphasized that attention-based transformers outperform traditional CNN or RNN frameworks, underscoring the superiority of contextual understanding in handling humor, irony, and sarcasm \[10\].
Recent studies have extended Hinglish sentiment analysis to broader tasks, such as multilingual emotion recognition and affective computing. A study used a multilingual transformer pipeline to analyze complex code-mixed expressions, reporting enhanced accuracy through contextual embeddings and domain adaptation \[11\]. Similarly, Baruah et al. developed a BiLSTM-based architecture optimized for mixed-script input, achieving notable improvements in recognizing emotion intensity and polarity \[12\]. Both studies underline that domain-specific pretraining and attention-based architectures significantly improve emotion recognition in low-resource settings.
Furthermore, Paul et al. introduced a sentiment dynamics framework that integrates attention layers to visualize the flow of emotions in bilingual texts \[17\]. This research demonstrated that attention mechanisms not only improve model interpretability but also help localize sentiment-bearing tokens in code-mixed data. Likewise, a review of Code-Mixed Sentiment Analysis shows that a multi-layered CNN-BiLSTM model achieves higher accuracy on benchmark Hinglish datasets by leveraging word embeddings and sentiment lexicons \[18\]. Additionally, recent efforts have focused on addressing linguistic variability and transliteration inconsistencies. A study on challenges in Code-Mixed NLP highlighted the limitations of tokenization, spelling variation, and the representation of Romanized Hindi, underscoring the need for data normalization before model training \[19\]. This work provides an essential foundation for preprocessing strategies in Hinglish NLP pipelines.
From the literature reviewed, it is evident that the field has undergone a clear methodological evolution from feature-engineered machine learning models to embedding-based, deep learning, and transformer-driven architectures. Early models offered interpretability but struggled with the complex semantics of bilingual text. Embedding models improved word-level representation but lacked contextual flexibility. Deep neural networks, such as CNNβBiLSTM, enhanced sequential understanding but required substantial labeled data. In contrast, transformer-based multilingual encoders such as MuRIL deliver superior contextual sensitivity, cross-lingual adaptability, and robustness to transliteration noise. Furthermore, ensemble and hybrid architectures combining these models continue to outperform standalone systems, offering a comprehensive solution to the nuances of Hinglish text processing.
However, despite notable progress, a gap persists in standardized benchmarking and comparative evaluation across models. Hence, the present research aims to bridge this gap by systematically evaluating diverse architectures β including MuRIL,Β CNNβBiLSTM, GloVeβBiLSTM, and Word2VecβLogistic Regression β on a unified Hinglish sentiment dataset to propose a robust multimodel framework for effective sentiment and emotion classification.
**Table 2.1: Literature Survey**
| Sl.NO | Title of the Article | Name of the Author | Name of the Journal | Year of publish | Findings | Research Gap |
|----|----|----|----|----|----|----|
| 1 | Sentiment Analysis of Code-Mixed Social Media Text (Hinglish) | Gaurav Singh\[1\] | arXiv (Preprint) | 2021 | Early Hinglish sentiment analysis; accuracy affected by transliteration inconsistency and token errors. | The necessity for generic sentiment analysis models specifically designed to address the grammatical and lexical intricacies of code-mixed Hinglish material. |
| 2 | Current State ofΒ Hinglish | Varsha Thakur, Roshani Sahu and Somya Omer\[2\] | SSRN Electronic Journal | Β 2020 | Traditional models showed limited contextual depth and poor scalability | The lack of a thorough assessment and organization of existing approaches, problems, and the latest technological advancements in Hinglish sentiment analysis. |
<table>
<colgroup>
<col style="width: 12%" />
<col style="width: 14%" />
<col style="width: 13%" />
<col style="width: 20%" />
<col style="width: 12%" />
<col style="width: 1%" />
<col style="width: 12%" />
<col style="width: 14%" />
</colgroup>
<thead>
<tr>
<th>Sl.NO</th>
<th>Title of the Article</th>
<th>Name of the Author</th>
<th>Name of the Journal</th>
<th>Year of publish</th>
<th colspan="2">Findings</th>
<th>Research Gap</th>
</tr>
</thead>
<tbody>
<tr>
<td>3</td>
<td>Improving Sentiment Analysis</td>
<td>Prof Neha Agarwal, Viraj Shah , Rishikesh Sharma , Himanshu Yadav, Vaibhav Shah[3]</td>
<td>Educational Administration: Theory and Practice</td>
<td colspan="2">2024</td>
<td>Hybrid deep learning enhanced accuracy compared to classical ML</td>
<td>Current models exhibit constrained accuracy in processing Hinglish; thus, there is a necessity for hybrid deep learning architectures to enhance performance.</td>
</tr>
<tr>
<td>4</td>
<td><h1 id="predicting-multi-label-emojisemotions-and-sentiments-in-code-mixed-texts-using-an-emojifying"><strong>Predicting Multi Label emojis,Emotions, and sentiments in code-mixed Texts using an emojifying</strong></h1></td>
<td>Gopendra Vikram Singh, Soumitra Ghosh, Mauajana Firdaus, Asif Ekbal, Pushpak Bhattacharya[4]</td>
<td>Scientific Reports</td>
<td colspan="2">Β 2024</td>
<td>Multilabel emotion and sentiment analysis improved via transformer-based contextual modeling.</td>
<td>Typical models usually forecast only one label (for instance, sentiment alone). There exists a deficiency in the ability to concurrently predict the interrelated aspects of emojis, various emotions, and sentiments in code-mixed language.</td>
</tr>
</tbody>
</table>
<table>
<colgroup>
<col style="width: 12%" />
<col style="width: 13%" />
<col style="width: 13%" />
<col style="width: 20%" />
<col style="width: 12%" />
<col style="width: 12%" />
<col style="width: 14%" />
</colgroup>
<thead>
<tr>
<th>Sl.NO</th>
<th>Title of the Article</th>
<th>Name of the Author</th>
<th>Name of the Journal</th>
<th>Year of publish</th>
<th>Findings</th>
<th>Research Gap</th>
</tr>
</thead>
<tbody>
<tr>
<td>5</td>
<td><p>A self-Attention hybrid emoji prediction model for code-mixedΒ </p>
<p>language</p></td>
<td>Gadde SatyaΒ Sai Naga Himabindu,Rajat Rao, Divyasikha Sethiya[5]</td>
<td>Social Network Analysis and Mining</td>
<td>2022</td>
<td>Contextualized embeddings outperformed CNN/RNN; effective in emoji prediction for Hinglish.</td>
<td>Predicting emojis in code-mixed Hinglish with precision is challenging, necessitating specific models that implement self-attention strategies to grasp combined semantic meanings.</td>
</tr>
<tr>
<td>6</td>
<td>Leveraging weakly annotated data for hate speech detection in code-mixed Hinglish: A feasibility-driven transfer learning approach with Large Language Models. In arXiv [cs.CL].</td>
<td>Sargam Yadav, Abishek Kaushik, Kevin McDaid,[6]</td>
<td>arXiv (Preprint)</td>
<td>2024</td>
<td>Used weak supervision to improve dataset diversity and mitigate annotation cost</td>
<td>The critical shortage of well-labeled datasets for Hinglish hate speech requires strategies that utilize poorly annotated data through Large Language Models (LLMs) and transfer learning.</td>
</tr>
</tbody>
</table>
<table style="width:100%;">
<colgroup>
<col style="width: 12%" />
<col style="width: 14%" />
<col style="width: 12%" />
<col style="width: 20%" />
<col style="width: 12%" />
<col style="width: 13%" />
<col style="width: 14%" />
</colgroup>
<thead>
<tr>
<th>Sl.NO</th>
<th>Title of the Article</th>
<th>Name of the Author</th>
<th>Name of the Journal</th>
<th>Year of publish</th>
<th>Findings</th>
<th>Research Gap</th>
</tr>
</thead>
<tbody>
<tr>
<td>7</td>
<td><p>βDid you really mean what you said?ββ―:</p>
<p>Sarcasm Detection in Hindi-English Code-Mixed Data using Bilingual Word Embeddings. In arXiv [cs.CL].</p></td>
<td>Akshita Agarwal,Anshul Wadhawan, Ashima Choudhury, Kavita Mourya, [7]</td>
<td><p>Β </p>
<p>arXiv (Preprint)</p></td>
<td>2020</td>
<td>Captured sarcasm and implicit polarity in bilingual data effectively.</td>
<td>Recognition of sarcasm in Hinglish frequently does not succeed with typical monolingual word representations, necessitating the use of tailored bilingual word embeddings to understand irony across languages.</td>
</tr>
<tr>
<td>8</td>
<td>Ensemble based hinglish hate speech detection. 2021 5th International Conference on Intelligent Computing and Control Systems (ICICCS).</td>
<td><p>Β Rahul, Gupta, V., Sehra, V., & Vardhan, Y. R.Β </p>
<p>[8]</p></td>
<td>ICICCS 2021 (Conference)</td>
<td>2021</td>
<td>Ensemble boosted recall and robustness compared to single models.</td>
<td><p>.</p>
<p>Independent models show limited effectiveness in identifying hate speech within noisy Hinglish datasets; collective techniques are necessary to combine their predictive strengths.</p></td>
</tr>
</tbody>
</table>
<table>
<colgroup>
<col style="width: 11%" />
<col style="width: 14%" />
<col style="width: 12%" />
<col style="width: 18%" />
<col style="width: 9%" />
<col style="width: 16%" />
<col style="width: 16%" />
</colgroup>
<thead>
<tr>
<th>Sl.NO</th>
<th>Title of the Article</th>
<th>Name of the Author</th>
<th>Name of the Journal</th>
<th>Year of publish</th>
<th>Findings</th>
<th>Research Gap</th>
</tr>
</thead>
<tbody>
<tr>
<td>9</td>
<td>Ensemble learning-based sarcasm detection in hinglish tweets using Word2Vec embedding. 2025 IEEE International Conference on Interdisciplinary Approaches in Technology and Management for Social Innovation (IATMSI), 1β6.</td>
<td><p>Acharya, A., & Goyal, R.</p>
<p>[9]</p></td>
<td>IATMSI 2025 (Conference)</td>
<td>2025</td>
<td>Combined transformer outputs achieved improved sarcasm detection accuracy.</td>
<td>The requirement to integrate word-level meaning models (Word2Vec) with ensemble machine learning techniques to more effectively understand sarcastic subtleties in Hinglish tweets.</td>
</tr>
<tr>
<td>10</td>
<td>Hilarious or hidden? Detecting sarcasm in hinglish tweets using BERT-GRU. 2023 14th International Conference on Computing Communication and Networking Technologies (ICCCNT)</td>
<td><p>Aloria, S., Aggarwal, I., Baliyan, N., & Ghosh, M.</p>
<p>[10]</p></td>
<td>ICCCNT 2023 (Conference)</td>
<td>Β 2023</td>
<td>Demonstrated attention-based transformer superiority for humor and irony detection</td>
<td>The sequential context and deep semantics of sarcasm in Hinglish are not fully captured by traditional models, requiring advanced hybrid deep learning architectures like BERT-GRU.</td>
</tr>
</tbody>
</table>
<table>
<colgroup>
<col style="width: 10%" />
<col style="width: 16%" />
<col style="width: 11%" />
<col style="width: 15%" />
<col style="width: 10%" />
<col style="width: 17%" />
<col style="width: 17%" />
</colgroup>
<thead>
<tr>
<th>Sl.NO</th>
<th>Title of the Article</th>
<th>Name of the Author</th>
<th>Name of the Journal</th>
<th>Year of publish</th>
<th>Findings</th>
<th>Research Gap</th>
</tr>
</thead>
<tbody>
<tr>
<td>11</td>
<td>Β A comparative study of machine learning and deep learning approaches for identifying Assamese abusive comments on social media. Procedia Computer Science, 258, 981β992.Β </td>
<td><p>Chutia, T., Baruah, N., & Sonowal, P.</p>
<p>[11]</p></td>
<td>Procedia Computer Science</td>
<td>2025</td>
<td>compared traditional ML and DL; BiLSTM outperformed SVM in contextual text understanding.</td>
<td>There is an absence of strict comparative standards between conventional machine learning methods and deep learning approaches specifically aimed at identifying abusive language in Assamese.</td>
</tr>
<tr>
<td>12</td>
<td>Named Entity Recognition in Assamese Language using two separate models: BiLSTM and BERT. Procedia Computer Science, 258, 242β251.</td>
<td><p>Baruah, P., Dutta, B., Sarma, S. K., & Talukdar, K</p>
<p>[12]</p></td>
<td>Procedia Computer Science</td>
<td>2025</td>
<td>Implemented dual-model NER; contributed insights for low-resource Indian languages.</td>
<td>The scarcity of efficient Named Entity Recognition (NER) solutions for the under-resourced Assamese language has led to the investigation of BiLSTM and BERT functionalities.</td>
</tr>
</tbody>
</table>
<table>
<colgroup>
<col style="width: 10%" />
<col style="width: 13%" />
<col style="width: 16%" />
<col style="width: 16%" />
<col style="width: 10%" />
<col style="width: 17%" />
<col style="width: 15%" />
</colgroup>
<thead>
<tr>
<th>Sl.NO</th>
<th>Title of the Article</th>
<th>Name of the Author</th>
<th>Name of the Journal</th>
<th>Year of publish</th>
<th>Findings</th>
<th>Research Gap</th>
</tr>
</thead>
<tbody>
<tr>
<td>13</td>
<td>Β Sentiment analysis of Mizo using lexical features in low resource based models. Natural Language Processing Journal, 13(100181), 100181[13]</td>
<td><p>Lalthangmawii, M., & Singh, T. D.</p>
<p>[13]</p></td>
<td>Natural Language Processing Journal</td>
<td>2025</td>
<td>Proposed lexical sentiment analysis for low-resource Mizo; highlighted linguistic diversity challenges</td>
<td>A significant deficiency of natural language processing tools and sentiment analysis frameworks for the severely low-resource Mizo language exists, necessitating customized models that utilize fundamental lexical characteristics.</td>
</tr>
<tr>
<td>14</td>
<td>Bidirectional LSTM-based sentiment analysis for Assamese text. American Journal of Computer Science and Technology, 7(2), 29β37</td>
<td>Talukdar, M., & Sarma, S. [14]</td>
<td>American Journal of Computer Science and Technology</td>
<td>2024</td>
<td>Achieved strong sequential understanding for sentiment prediction using BiLSTM.</td>
<td>The necessity for models that can grasp two-way sequential context (BiLSTM) to enhance the precision of sentiment polarity classification in Assamese language content.</td>
</tr>
</tbody>
</table>
<table>
<colgroup>
<col style="width: 7%" />
<col style="width: 17%" />
<col style="width: 15%" />
<col style="width: 12%" />
<col style="width: 8%" />
<col style="width: 19%" />
<col style="width: 18%" />
</colgroup>
<thead>
<tr>
<th>Sl.NO</th>
<th>Title of the Article</th>
<th>Name of the Author</th>
<th>Name of the Journal</th>
<th>Year of publish</th>
<th>Findings</th>
<th>Research Gap</th>
</tr>
</thead>
<tbody>
<tr>
<td>15</td>
<td>Β CodemixedNLP: An Extensible and Open NLP Toolkit for Code-Mixing. In arXiv [cs.CL].Β </td>
<td><p>Jayanthi, S. M., Nerella, K., Chandu, K. R., & Black, A. W.Β </p>
<p>[15]</p></td>
<td>arXiv (Preprint)</td>
<td>2021</td>
<td>Introduced open-source NLP toolkit for processing code-mixed languages like Hinglish</td>
<td>The lack of a comprehensive, open-source, and adaptable NLP framework specifically created to manage the preprocessing, modeling, and assessment of languages that are mixed with code.</td>
</tr>
<tr>
<td>16</td>
<td>Code-Mixed Hinglish Hate Speech Detection Dataset. Kaggle.com</td>
<td><p>Dhekane, S.Β </p>
<p>[16]</p></td>
<td>Kaggle (Dataset Repository)</td>
<td>2025</td>
<td>Publicly available dataset enabling research on Hinglish hate-speech classification.</td>
<td>There is a considerable shortage of publicly available, high-quality, and uniform datasets that are essential for training and evaluating Hinglish hate speech detection models.</td>
</tr>
</tbody>
</table>
#
#
#
#
**CHAPTER 3**
# **METHODOLOGY**
#
# 3.1 INTRODUCTION
#
## The overall workflow involves of three proposed methodologies that is illustrated in Fig. 3.1, Fig.3.2, and Fig3.3 which provides a high-level view of the data preprocessing, feature extraction, model training, and evaluation process. It visually summarizes the pipeline from raw dataset input to performance evaluation across mentioned models. The project workflow begins with the primary input which is raw sentiment data of English in Latin text, Hindi in Devnagiri texts and often consisting of code-mixed sentences Hindi-English (Hinglish) in Latin text. This dataset is unstructured and not immediately suitable for Tokenization and further Word Embeddings. Therefore, the first crucial step is Text Pre-processing that involves URL Removal, Lowercasing, Whitespace removal, Tokenization and Embeddings. This Text Pre-processing is critical because deep learning models such as BERT Transformer Models assign different vectors for the same word with different casings, to solve this we use Lowercasing, URLs which do not provide any context so we remove URLs and space normalization as tokens of extra spaces are also created. After preprocessing, the dataset is divided into training (60%), validation (10%), and testing (30%) sets (same for all methodologies). The processed data is then fed into respective sentiment analysis model, where the embedding techniques vary according to the methodology being implemented.
<img src="thesis_media/media/image2.jpg" style="width:5.62651in;height:4.21988in" />
**Fig 3.1: Β Workflow of the proposed methodology for sentiment analysis.**
<span id="_TOC_250032" class="anchor"></span>After preprocessing and dataset splitting, the training set is used to develop multiple hybrid sentiment classification models by combining various word embedding techniques with machine learning and deep learning algorithms. Specifically, Word2Vec, GloVe, FastText, Universal Sentence Encoder (USE), and ELMo embeddings are integrated with LSTM and LightGBM classifiers to capture the semantic and contextual information present in code-mixed text. The validation set is used for model tuning and performance optimization, while the test set is employed for final evaluation. The predicted sentiment labels generated by these embedding-classifier combinations are then analyzed to compare their effectiveness and identify the best-performing model for code-mixed sentiment analysis.
##
<img src="thesis_media/media/image3.jpg" style="width:6.26667in;height:4.7in" />
**Fig 3.2: Β Workflow of the proposed methodology for sentiment analysis.**
In the above proposed methodology Fig3.2: Two training strategies are applied. In Strategy-1, each language corpus and the combined dataset are independently trained and evaluated on the test set. In Strategy-2, multi-stage language training is performed sequentially using English, Hinglish, Hindi, and combined corpora to improve cross-lingual understanding. Finally, predicted sentiment labels for code-mixed texts are compared through result analysis to evaluate model effectiveness and across multilingual strategies.
<img src="thesis_media/media/image4.jpg" style="width:6.37751in;height:4.78313in" />
**Fig 3.3: Β Workflow of the proposed methodology for sentiment analysis.**
In the proposed methodology fig3, we performed multi-stage language training that consists of six various strategies of training done sequentially using (English, Hinglish, Hindi, all combined), (English, Hindi, Hinglish, all combined ), (Hinglish, English, Hindi, all combined), (Hinglish, Hindi, English, all combined), (Hindi, Hinglish, English, all combined), and (Hindi, English, Hinglish, all combined) corpora to improve cross-lingual understanding. Finally, predicted sentiment labels for code-mixed texts are compared through result analysis to evaluate model effectiveness and across multilingual strategies.
## 3.3 DATA COLLECTION
##
The research utilized the "combined_hate_speech_dataset" that is publically available on Kaggle. This dataset includes 29,550 labeled text entries, predominantly featuring code-mixed Hinglish sentences. For the purpose of analysis, two main columns were preserved: the text content and the hate_label. The target variable is binary, with 0 indicating non-hate content and 1 signifying hate content. The dataset shows a slight imbalance in class distribution, with non-hate samples being more prevalent. The following table clearly provides data description.
Table 3.1: Dataset Details
<table style="width:89%;">
<colgroup>
<col style="width: 25%" />
<col style="width: 63%" />
</colgroup>
<thead>
<tr>
<th style="text-align: left;"><p>Β </p>
<p><strong>Dataset Name</strong></p></th>
<th style="text-align: left;"><strong>combined_hate_speech_dataset (PRISM) β Kaggle</strong></th>
</tr>
</thead>
<tbody>
<tr>
<td style="text-align: left;">Total Samples</td>
<td style="text-align: left;">29,550</td>
</tr>
<tr>
<td style="text-align: left;">Classification Type</td>
<td style="text-align: left;">Binary (0 = Non-hate, 1 = Hate)</td>
</tr>
<tr>
<td style="text-align: left;">Languages Covered</td>
<td style="text-align: left;">English, Hindi (Devanagari), Hinglish</td>
</tr>
<tr>
<td style="text-align: left;">Data Composition</td>
<td style="text-align: left;">English: 15,000</td>
</tr>
<tr>
<td style="text-align: left;">Β </td>
<td style="text-align: left;">Hindi: 9,767</td>
</tr>
<tr>
<td style="text-align: left;">Β </td>
<td style="text-align: left;">Hinglish: 4,783</td>
</tr>
<tr>
<td style="text-align: left;">Label Distribution</td>
<td style="text-align: left;">Non-hate: 15,825</td>
</tr>
<tr>
<td style="text-align: left;">Β </td>
<td style="text-align: left;">Hate: 13,725</td>
</tr>
<tr>
<td style="text-align: left;">Profanity Lexicon</td>
<td style="text-align: left;">209 offensive terms with severity scores</td>
</tr>
<tr>
<td style="text-align: left;">Application</td>
<td style="text-align: left;">Hate-speech / Toxicity detection in multilingual text</td>
</tr>
</tbody>
</table>
## 3.4 DATA PREPROCESSING
The dataset originally consisted of 29,550 entries, with 9 attributes such as text, hate_label, source, profanity_score, language, dataset_version, combined_date, text_length and word_count. This dataset included samples in English (Latin Script), Hindi(Devnagiri Script), and Hinglish(Latin Script), for the purpose of classifying hate speech. To enhance the quality of the data, thorough preprocessing and noise analysis were conducted prior to tokenization and embedding. During the analysis, it was found that duplicate rows, repeated texts, URLs, mentions, hashtags, elongated words, and social media attachments were significant sources of noise. Regex-based techniques were employed to identify and eliminate these noisy elements. Further preprocessing involved converting text to lowercase, removing extra spaces, normalizing elongated words, and deleting URLs and HTML entities. The following table shows the general noise statistics that are found during preprocessing.
Table 3.2: List Count of URL Noise
| **URL / Social Noise Type** | **Count** |
|:----------------------------|:----------|
| HTTP/HTTPS Links | 356 |
| WWW Links | 7 |
| pic.twitter Links | 106 |
| YouTube Links | 22 |
| Facebook Links | 0 |
| Hungama Links | 1 |
| Attached URLs | 33 |
The preprocessing stage significantly improved the overall quality of the dataset by removing noisy and redundant textual patterns. Duplicate text entries, URLs, and elongated words were identified as major sources of inconsistency that could negatively affect model learning and classification performance. Cleaning operations such as duplicate removal, URL elimination, and text normalization helped create a more consistent and standardized corpus for embedding. As a result, the dataset size was slightly reduced while preserving meaningful information required for hate speech detection. The following table shows the comparison of the dataset before and after cleaning.
Table 3.3: Data Metrics Before and After Cleaning
| **Metric** | **Before Cleaning** | **After Cleaning** |
|:---------------------------|:--------------------|:-------------------|
| Total Rows | 29,550 | 29,506 |
| Duplicate Texts | 11 | 0 |
| Texts with URLs | 458 | 0 |
| Texts with Elongated Words | 2,176 | 0 |
After preprocessing, essential features were extracted to prepare the dataset for hate speech classification. The refined dataset retained only relevant attributes required for model training and analysis. The clean_text feature contains normalized textual content, while hate_label represents the target classification variable. The language feature identifies the language category of each sample, and additional statistical features such as text_length and word_count were included to capture textual characteristics. These extracted features help improve data representation and support effective downstream modeling. The final processed dataset consisted of 29,506 samples with 5 important features.
## 3.5 EDA
<img src="thesis_media/media/image5.png" style="width:3.66234in;height:3.45649in" />
Fig:3.4 Hate vs Non-Hate Class distribution
The following pie chart illustrates the distribution of hate and non-hate samples in the PRISM dataset. The dataset contains approximately balanced class representations, where non-hate samples account for **53.5%** and hate samples account for **46.5%** of the total data. Maintaining a balanced class distribution is important for reducing model bias and improving classification performance.
<img src="thesis_media/media/image6.png" style="width:3.64935in;height:3.60256in" />
Fig:3.5 Language Distribution
The following pie chart presents the language distribution of the dataset across English, Hindi, and Hinglish corpora. English samples constitute **50.8%**, Hindi samples represent **33.0%**, and Hinglish samples contribute **16.2%** of the dataset. This multilingual distribution enables the models to learn diverse linguistic patterns for multilingual hate speech detection.
<img src="thesis_media/media/image7.png" style="width:4.27923in;height:3.25974in" />
Fig:3.6 Feature Correlation Heatmap
The correlation heatmap illustrates the relationship among numerical features such as hate_label, text_length, and word_count. A very strong positive correlation (0.99) is observed between text length and word count, indicating that longer texts generally contain more words. In contrast, hate_label shows a weak negative correlation with both text length and word count, suggesting that text size has minimal influence on hate speech classification.
<img src="thesis_media/media/image8.png" style="width:3.19849in;height:2.51948in" /><img src="thesis_media/media/image9.png" style="width:3.00752in;height:2.58956in" />
Fig:3.7 Text length Statistics and Word count Statisitics
The following figures present the statistical distribution of text length and word count in the final cleaned dataset. The text length analysis shows a mean of 150.52 characters and a median of 94 characters, with some samples reaching a maximum length of 1926 characters. Similarly, the word count distribution shows an average of 28.38 words and a median of 18 words, while the maximum word count reaches 300 words. The 95th percentile values of 480 characters and 90 words indicate the presence of long textual samples and high variance within the dataset. These observations are important for determining appropriate padding and truncation limits in transformer-based models to ensure efficient training and balanced sequence representation.
## 3.6 DATASET SPLITTING
The cleaned PRISM dataset was analyzed to understand the distribution of hate labels and language categories before model training. The dataset contains English, Hindi, and Hinglish samples with a relatively balanced hate speech distribution across languages. To ensure reliable model evaluation, the dataset was divided into training, validation, and testing subsets using a stratified splitting approach. Initially, 70% of the data was reserved as the training pool and 30% as the independent test set. The training pool was further divided into 60% training data and 10% validation data. Only the clean_text and hate_label columns were used during model training. Language-wise class distributions were maintained across all subsets to preserve dataset balance and reduce sampling bias.
The following table shows the language wise data splitting.
Table 3.4: Dataset Splits Table
<table>
<colgroup>
<col style="width: 15%" />
<col style="width: 9%" />
<col style="width: 10%" />
<col style="width: 15%" />
<col style="width: 9%" />
<col style="width: 8%" />
<col style="width: 7%" />
<col style="width: 7%" />
<col style="width: 7%" />
<col style="width: 7%" />
</colgroup>
<thead>
<tr>
<th colspan="3" style="text-align: left;"><p>Β </p>
<p><strong>Category</strong></p></th>
<th style="text-align: left;"><strong>Combined</strong></th>
<th colspan="2" style="text-align: left;"><strong>English</strong></th>
<th colspan="2" style="text-align: left;"><strong>Hindi</strong></th>
<th colspan="2" style="text-align: left;"><strong>Hinglish</strong></th>
</tr>
</thead>
<tbody>
<tr>
<td colspan="3" style="text-align: left;"><strong>Total Samples</strong></td>
<td style="text-align: left;">29506</td>
<td colspan="2" style="text-align: left;">14994</td>
<td colspan="2" style="text-align: left;">9738</td>
<td colspan="2" style="text-align: left;">4774</td>
</tr>
<tr>
<td rowspan="2" style="text-align: left;"><strong>Label Distribution</strong></td>
<td colspan="2" style="text-align: left;">Non-hate(0)</td>
<td style="text-align: left;">15,799</td>
<td colspan="2" style="text-align: left;">7495</td>
<td colspan="2" style="text-align: left;">5393</td>
<td colspan="2" style="text-align: left;">2911</td>
</tr>
<tr>
<td colspan="2" style="text-align: left;">Hate(1)</td>
<td style="text-align: left;">13,707</td>
<td colspan="2" style="text-align: left;">7499</td>
<td colspan="2" style="text-align: left;">4345</td>
<td colspan="2" style="text-align: left;">1863</td>
</tr>
<tr>
<td rowspan="6" style="text-align: left;"><strong>Split</strong></td>
<td rowspan="2" style="text-align: left;"><p><strong>Train</strong></p>
<p><strong>(60%)</strong></p></td>
<td style="text-align: left;">Non-hate(0)</td>
<td rowspan="2" style="text-align: left;">18617</td>
<td style="text-align: left;">4478</td>
<td rowspan="2" style="text-align: left;">9446</td>
<td style="text-align: left;">3237</td>
<td rowspan="2" style="text-align: left;">6143</td>
<td style="text-align: left;">1764</td>
<td rowspan="2" style="text-align: left;">3011</td>
</tr>
<tr>
<td style="text-align: left;">Hate(1)</td>
<td style="text-align: left;">4485</td>
<td style="text-align: left;">2622</td>
<td style="text-align: left;">1117</td>
</tr>
<tr>
<td rowspan="2" style="text-align: left;"><p><strong>Val</strong></p>
<p><strong>(10%)</strong></p></td>
<td style="text-align: left;">Non-hate(0)</td>
<td rowspan="2" style="text-align: left;">2086</td>
<td style="text-align: left;">525</td>
<td rowspan="2" style="text-align: left;">1050</td>
<td style="text-align: left;">378</td>
<td rowspan="2" style="text-align: left;">683</td>
<td style="text-align: left;">204</td>
<td rowspan="2" style="text-align: left;">335</td>
</tr>
<tr>
<td style="text-align: left;">Hate(1)</td>
<td style="text-align: left;">525</td>
<td style="text-align: left;">305</td>
<td style="text-align: left;">131</td>
</tr>
<tr>
<td rowspan="2" style="text-align: left;"><p><strong>Test</strong></p>
<p><strong>(30%)</strong></p></td>
<td style="text-align: left;">Non-hate(0)</td>
<td rowspan="2" style="text-align: left;">8865</td>
<td style="text-align: left;">2297</td>
<td rowspan="2" style="text-align: left;">4499</td>
<td style="text-align: left;">1597</td>
<td rowspan="2" style="text-align: left;">2926</td>
<td style="text-align: left;">846</td>
<td rowspan="2" style="text-align: left;">1434</td>
</tr>
<tr>
<td style="text-align: left;">Hate(1)</td>
<td style="text-align: left;">2248</td>
<td style="text-align: left;">1303</td>
<td style="text-align: left;">561</td>
</tr>
</tbody>
</table>
##
## 3.7 MODEL SELECTION
##
## 3.7.1 MODEL DESCRIPTION
> 3.7.1.1 TRANSFORMER BASED MODELS
1. **MuRIL:** MuRIL is a multilingual transformer model developed by Google for Indian languages and code-mixed text. It generates contextual embeddings that capture semantic relationships across multiple languages. MuRIL was selected because the dataset contains Hindi, English, and Hinglish text, making it highly effective for multilingual hate speech classification.
2. **mBART:** mBART is a multilingual encoderβdecoder transformer model designed for cross-lingual understanding and contextual language representation. It was used because of its strong capability to learn multilingual semantic patterns from diverse textual inputs.
3. **HingRoBERTa:** HingRoBERTa is a RoBERTa-based transformer specifically adapted for Hinglish and code-mixed language processing. It was selected because it effectively handles transliterated and mixed-language text commonly found in social media hate speech datasets.
4. **MPNet:** MPNet combines masked language modeling with permuted positional encoding to improve contextual understanding. It was chosen because of its strong sentence representation capability and effectiveness in capturing complex hate speech semantics.
> 3.7.1.2 DEEP LEARNING MODELS
5. **FastText + BiLSTM:** This model combines FastText embeddings with a Bidirectional LSTM network. FastText captures subword information and spelling variations, while BiLSTM learns contextual dependencies from both forward and backward directions. It was selected because Hinglish and social media text often contain noisy and misspelled words.
6. **Word2Vec + BiLSTM:** Word2Vec provides semantic word embeddings, and BiLSTM captures sequential contextual information from text. This model was used to evaluate the effectiveness of predictive word embeddings for multilingual hate speech detection.
7. **GloVe + BiLSTM:** GloVe embeddings capture global word co-occurrence information, while BiLSTM models sequential text patterns. This architecture was selected to compare static embedding-based contextual learning against transformer models.
8. **Word2Vec + LSTM:** This model combines Word2Vec embeddings with LSTM networks for sequence learning. It was selected because LSTM effectively captures long-term dependencies in textual data.
9. **GloVe + LSTM:** GloVe embeddings with LSTM were used to analyze the effectiveness of global semantic representations in hate speech classification tasks.
10. **FastText + LSTM:** FastText embeddings combined with LSTM were selected because FastText handles subword-level variations effectively, which is useful for multilingual and code-mixed text.
11. **USE+LSTM:** This model uses Universal Sentence Encoder embeddings with LSTM networks. USE captures sentence-level semantic meaning, making it useful for understanding contextual hate speech patterns.
12. **ELMo+LSTM:** ELMo in conjunction with LSTM integrates advanced contextualized word representations with a sequential deep learning architecture, which enhances its effectiveness in grasping intricate syntactic and semantic frameworks throughout a text.
> 3.7.1.3 MACHINE LEARNING MODELS
13. **Word2Vec + LightGBM:** This model combines Word2Vec embeddings with the LightGBM classifier. It was selected because LightGBM efficiently handles vectorized textual features using gradient boosting techniques.
14. **GloVe + LightGBM:** GloVe embeddings were used as feature inputs for LightGBM to evaluate how global semantic vectors perform with boosting-based classification.
15. **FastText + LightGBM:** FastText embeddings combined with LightGBM were selected because FastText captures subword semantics effectively while LightGBM provides efficient classification performance.
16. **USE + LightGBM:** This model uses USE sentence embeddings with LightGBM classification. It was selected to evaluate sentence-level semantic representations using boosting methods.
17. **ELMo+ LightGBM:** ELMo combined with LightGBM derives constant contextual feature vectors from text through ELMo and feeds them into a highly efficient, gradient-boosted decision tree classifier, providing a resource-saving method for classification tasks that resemble tabular data.
## 3.7.2 MODEL ALGORITHMS
The selected models use supervised learning for hate speech classification. Transformer models employ fine-tuning with contextual embeddings and attention mechanisms for sequence understanding. BiLSTM and LSTM architectures process text sequentially to capture contextual dependencies from embeddings. LightGBM models use gradient boosting decision trees for efficient feature-based classification, while Logistic Regression applies a linear decision boundary over Word2Vec embeddings. Cross-Entropy and Binary Cross-Entropy losses were used for binary classification tasks, and optimizers such as Adam, AdamW, and Gradient Boosting strategies were applied for stable convergence and improved learning performance.
Table 3.5: Model Algorithms
<table>
<colgroup>
<col style="width: 21%" />
<col style="width: 22%" />
<col style="width: 24%" />
<col style="width: 12%" />
<col style="width: 19%" />
</colgroup>
<thead>
<tr>
<th><p>Β </p>
<p><strong>Model</strong></p></th>
<th><strong>Core Concept</strong></th>
<th><strong>Training Method</strong></th>
<th><strong>Loss function</strong></th>
<th><strong>Optimizer</strong></th>
</tr>
</thead>
<tbody>
<tr>
<td>MuRIL</td>
<td>Transformer-based model, contextual multilingual embeddings</td>
<td>Supervised sequence fine-tuning + linear warmupβdecay scheduler</td>
<td>Cross-Entropy Loss</td>
<td>AdamW</td>
</tr>
<tr>
<td>mBART</td>
<td>Transformer-based encoderβdecoder model with fully contextual multilingual embeddings</td>
<td>Supervised sequence fine-tuning with linear warmupβdecay scheduler</td>
<td>Cross-Entropy Loss</td>
<td>Adam Optimizer</td>
</tr>
<tr>
<td>HingRoBERTa</td>
<td>Transformer-based model (RoBERTa variant), contextual embeddings adapted for Hinglish/code-mixed text</td>
<td>Supervised sequence fine-tuning + linear warmupβdecay scheduler</td>
<td>Cross-Entropy Loss</td>
<td>AdamW</td>
</tr>
<tr>
<td><p>Β </p>
<p>MPNet</p></td>
<td>Transformer-based model using masked language modeling with permuted position encoding for enhanced contextual representation</td>
<td>Supervised sequence fine-tuning + linear warmupβdecay scheduler</td>
<td>Cross-Entropy Loss</td>
<td>AdamW</td>
</tr>
<tr>
<td>FastText + BiLSTM</td>
<td>Subword (character n-gram) embeddings + Bidirectional LSTM contextual modeling</td>
<td>Supervised mini-batch training with backpropagation</td>
<td>Cross-Entropy Loss</td>
<td>Adam Optimizer</td>
</tr>
<tr>
<td>Word2Vec + BiLSTM</td>
<td>Predictive word embeddings (CBOW/Skip-gram) + Bidirectional LSTM contextual modeling</td>
<td>Supervised mini-batch training with backpropagation</td>
<td>Cross-Entropy Loss</td>
<td>Adam Optimizer</td>
</tr>
<tr>
<td>GloVe + BiLSTM</td>
<td>Static global co-occurrence embeddings + Bidirectional LSTM contextual modeling</td>
<td>Supervised mini-batch training with backpropagation</td>
<td>Cross-Entropy Loss</td>
<td>Adam Optimizer</td>
</tr>
<tr>
<td><p>Β </p>
<p>Word2Vec + LSTM</p></td>
<td>Uses Word2Vec embeddings (semantic similarity) + LSTM to capture sequential dependencies in text</td>
<td>Pretrained Word2Vec embeddings fed into LSTM, trained end-to-end on labeled data</td>
<td>Binary Cross-Entropy</td>
<td>Adam</td>
</tr>
<tr>
<td>GloVe + LSTM</td>
<td>Uses GloVe embeddings (global word co-occurrence statistics) + LSTM for sequence learning</td>
<td>Pretrained GloVe vectors used as embedding layer, then LSTM training on dataset</td>
<td>Binary Cross-Entropy</td>
<td>Adam</td>
</tr>
<tr>
<td>FastText + LSTM</td>
<td>FastText captures subword information (handles misspellings, Hinglish variations) + LSTM</td>
<td>Pretrained FastText embeddings β LSTM trained on sequence data</td>
<td>Binary Cross-Entropy</td>
<td>Adam</td>
</tr>
<tr>
<td><p>Β </p>
<p>USE + LSTM</p></td>
<td>Universal Sentence Encoder provides sentence-level embeddings + LSTM for deeper sequence modeling</td>
<td>USE embeddings generated β passed to LSTM for classification training</td>
<td>Binary Cross-Entropy</td>
<td>Adam</td>
</tr>
<tr>
<td>Word2Vec + LightGBM</td>
<td>Gradient Boosting decision trees + Word2Vec feature vectors</td>
<td>Word2Vec embeddings averaged or pooled β fed into LightGBM classifier</td>
<td>Binary Log Loss</td>
<td>Gradient Boosting Decision Trees</td>
</tr>
<tr>
<td>GloVe + LightGBM</td>
<td>GloVe embeddings used as input features + LightGBM for classification</td>
<td>GloVe vectors aggregated β used to train LightGBM model</td>
<td>Binary Log Loss</td>
<td><p>Stochastic Gradient Descent</p>
<p>(+ Negative Sampling)</p></td>
</tr>
<tr>
<td>Β FastText + LightGBM</td>
<td>FastText embeddings (handles subwords well) + LightGBM classifier</td>
<td>FastText vectors β feature input β LightGBM training</td>
<td>Binary Log Loss</td>
<td><p>Stochastic Gradient Descent</p>
<p>(+ Negative Sampling + Subword learning)</p></td>
</tr>
<tr>
<td>USE + LightGBM</td>
<td>Sentence-level embeddings from USE + LightGBM classification</td>
<td>USE embeddings β directly fed to LightGBM</td>
<td>Binary Log Loss</td>
<td>AdaGrad</td>
</tr>
<tr>
<td>ELMo+ LSTM</td>
<td>ELMo gives GBDT</td>
<td>Supervised DL (hybrid )</td>
<td><p>Binary Cross</p>
<p>entropy</p></td>
<td>SGD</td>
</tr>
<tr>
<td>ELMo+ LightGBM</td>
<td>Word2Vec embeddings + Logistic Regression classifier</td>
<td>Supervised ML (linear model)</td>
<td>Binary Log Loss</td>
<td>Gradient-Based One-Side Sampling (GOSS) and Exclusive Feature Bundling (EFB)</td>
</tr>
</tbody>
</table>
## 3.7.3 MODEL HYPERPARAMETERS
Hyperparameters were selected to balance computational efficiency and model performance. For Transformer models we set max sequence length = 128 to capture sufficient contextual information while maintaining manageable memory usage. A learning rate of 2e-5 with AdamW optimizer and weight decay of 0.01 was used to ensure stable fine-tuning and prevent overfitting. For BiLSTM and LSTM-based models, embedding dimensions of 100β300 and hidden dimensions up to 256 were chosen to learn rich semantic representations. Dropout values between 0.2β0.5 were applied to reduce overfitting. Batch sizes of 16 and 32 were selected for balanced training stability and GPU utilization. LightGBM hyperparameters such as max_depth, num_leaves, and n_estimators were tuned to improve classification performance while controlling model complexity. Lower learning rates and regularization parameters were used to achieve stable gradient updates and better generalization.
Table 3.6: Model Hyperparameters
<table>
<colgroup>
<col style="width: 18%" />
<col style="width: 20%" />
<col style="width: 23%" />
<col style="width: 17%" />
<col style="width: 10%" />
<col style="width: 9%" />
</colgroup>
<thead>
<tr>
<th><p>Β </p>
<p><strong>Model</strong></p></th>
<th><strong>Core Components / Embeddings</strong></th>
<th><strong>Key Hyperparameters</strong></th>
<th><strong>Optimizer</strong></th>
<th><strong>Epochs</strong></th>
<th><strong>Batch Size</strong></th>
</tr>
</thead>
<tbody>
<tr>
<td><strong>MuRIL</strong></td>
<td>Transformer (google/muril-base-cased)</td>
<td>max_seq_len=128, weight_decay=0.01, warmup_ratio=0.1</td>
<td>AdamW, lr=2e-5</td>
<td>8</td>
<td>16</td>
</tr>
<tr>
<td><strong>mBART</strong></td>
<td>Transformer (facebook/mbart-large-50)</td>
<td>max_seq_len = 128, weight_decay = 0.01, warmup_ratio = 0.1</td>
<td>AdamW, lr = 2e-5</td>
<td>8</td>
<td>16</td>
</tr>
<tr>
<td><strong>HingRoBERTa</strong></td>
<td>Transformer (RoBERTa-based Hinglish-adapted model)</td>
<td>max_seq_len = 128, weight_decay = 0.01, warmup_ratio = 0.1</td>
<td>AdamW, lr = 2e-5</td>
<td>8</td>
<td>16</td>
</tr>
<tr>
<td><strong>MPNet</strong></td>
<td>Transformer (microsoft/mpnet-base)</td>
<td>max_seq_len = 128, weight_decay = 0.01, warmup_ratio = 0.1</td>
<td>AdamW, lr = 2e-5</td>
<td>8</td>
<td>16</td>
</tr>
<tr>
<td><p><strong>Β </strong></p>
<p><strong>FastText + BiLSTM</strong></p></td>
<td>BiLSTM + Pre-trained FastText (subword) embeddings</td>
<td>max_seq_len = 128, embedding_dim = 300, hidden_dim = 256, dropout = 0.5</td>
<td>Adam, lr = 1e-3, weight_decay = 0</td>
<td>10</td>
<td>32</td>
</tr>
<tr>
<td><strong>Word2Vec + BiLSTM</strong></td>
<td>BiLSTM + Pre-trained Word2Vec embeddings</td>
<td>max_seq_len = 128, embedding_dim = 300, hidden_dim = 256, dropout = 0.5</td>
<td>Adam, lr = 1e-3, weight_decay = 0</td>
<td>10</td>
<td>32</td>
</tr>
<tr>
<td><strong>GloVe + BiLSTM</strong></td>
<td>BiLSTM + Pre-trained GloVe embeddings</td>
<td>max_seq_len = 128, embedding_dim = 300, hidden_dim = 256, dropout = 0.5</td>
<td>Adam, lr = 1e-3, weight_decay = 0</td>
<td>10Β </td>
<td>32</td>
</tr>
<tr>
<td><p><strong>Β </strong></p>
<p><strong>LSTMΒ </strong></p></td>
<td>Word2vec</td>
<td>max_seq_len =100 , embedding_dim =100 , hidden_dim = 128, dropout = 0.2</td>
<td>Adam, lr = 1e-2, weight_decay = 0</td>
<td>30</td>
<td>16</td>
</tr>
<tr>
<td><strong>LSTMΒ </strong></td>
<td>GLOVEΒ </td>
<td>max_seq_len =100 , embedding_dim =100 , hidden_dim = 128, dropout = 0.2</td>
<td>Adam, lr = 1e-2, weight_decay = 0</td>
<td>30</td>
<td>16</td>
</tr>
<tr>
<td><strong>LSTMΒ </strong></td>
<td>FASTTEXTΒ </td>
<td>max_seq_len =100 , embedding_dim =100 , hidden_dim = 128, dropout = 0.2</td>
<td>Adam, lr = 1e-2, weight_decay = 0</td>
<td>30</td>
<td>16</td>
</tr>
<tr>
<td><p><strong>Β </strong></p>
<p><strong>LSTMΒ </strong></p></td>
<td>USE</td>
<td>max_seq_len =100 , embedding_dim =100 , hidden_dim = 128, dropout = 0.2</td>
<td>Adam, lr = 1e-2, weight_decay = 0</td>
<td>30</td>
<td>16</td>
</tr>
<tr>
<td><strong>LightGBM</strong></td>
<td>Word2vec</td>
<td><p>learning_rate = 0.01</p>
<p>n_estimators = 500</p>
<p>max_depth = 8</p>
<p>num_leaves = 63</p>
<p>subsample = 0.8</p>
<p>colsample_bytree = 0.8</p>
<p>reg_alpha = 0.1</p>
<p>reg_lambda = 0.2</p></td>
<td>Gradient Boosting Decision Trees</td>
<td>20</td>
<td>16</td>
</tr>
<tr>
<td><strong>LightGBM</strong></td>
<td>GLOVEΒ </td>
<td><p>learning_rate = 0.01</p>
<p>n_estimators = 500</p>
<p>max_depth = 8</p>
<p>num_leaves = 63</p>
<p>subsample = 0.8</p>
<p>colsample_bytree = 0.8</p>
<p>reg_alpha = 0.1</p>
<p>reg_lambda = 0.2</p></td>
<td><p>Stochastic Gradient Descent</p>
<p>(+ Negative Sampling)</p></td>
<td>20</td>
<td>16</td>
</tr>
<tr>
<td><p><strong>Β </strong></p>
<p><strong>LightGBM</strong></p></td>
<td>FASTTEXTΒ </td>
<td><p>learning_rate = 0.01</p>
<p>n_estimators = 500</p>
<p>max_depth = 8</p>
<p>num_leaves = 63</p>
<p>subsample = 0.8</p>
<p>colsample_bytree = 0.8</p>
<p>reg_alpha = 0.1</p>
<p>reg_lambda = 0.2</p></td>
<td><p>Stochastic Gradient Descent</p>
<p>(+ Negative Sampling + Subword learning)</p></td>
<td>20</td>
<td>16</td>
</tr>
<tr>
<td><strong>LightGBM</strong></td>
<td>USE</td>
<td><p>learning_rate = 0.01</p>
<p>n_estimators = 500</p>
<p>max_depth = 8</p>
<p>num_leaves = 63</p>
<p>subsample = 0.8</p>
<p>colsample_bytree = 0.8</p>
<p>reg_alpha = 0.1</p>
<p>reg_lambda = 0.2</p></td>
<td>AdaGrad</td>
<td>20</td>
<td>16</td>
</tr>
<tr>
<td><strong>LightGBM</strong></td>
<td>ELMo</td>
<td><p>learning_rate = 0.01</p>
<p>n_estimators = 500</p>
<p>max_depth = 8</p>
<p>num_leaves = 63</p>
<p>subsample = 0.8</p>
<p>colsample_bytree = 0.8</p>
<p>reg_alpha = 0.1</p>
<p>reg_lambda = 0.2</p></td>
<td>Gradient-Based One-Side Sampling (GOSS) and Exclusive Feature Bundling (EFB)</td>
<td>30</td>
<td>16</td>
</tr>
<tr>
<td><p><strong>Β </strong></p>
<p><strong>LSTM</strong></p></td>
<td>ELMo</td>
<td>max_seq_len =200 , embedding_dim =200 , hidden_dim = 128, dropout = 0.2</td>
<td>SGD (Stochastic Gradient Descent)., lr = 1e-2, weight_decay = 0</td>
<td>30</td>
<td>16</td>
</tr>
</tbody>
</table>
##
## **3.6 EVALUATION METRICS**
##
## As the Dataset consists of two classes i.e, 0 (non-hate) and 1 (hate) The performance of the proposed system was rigorously evaluated using 7 well-established metrics that includes Accuracy, Balanced Accuracy, Precision, Recall, Specificity, F1 and AUC-ROC.
**Accuracy:** Accuracy measures the overall percentage of correctly classified samples among the total predictions. It evaluates how well the model performs on both hate and non-hate classes collectively. Accuracy was used to measure the general classification performance of the models.
``` math
Accuracy = \frac{TP + TN}{TP + TN + FP + FN}
```
**Balanced Accuracy:** Balanced Accuracy computes the average recall obtained for each class and is useful for handling class imbalance. It ensures that both hate and non-hate classes contribute equally to evaluation. This metric was used to provide unbiased performance measurement across classes.
``` math
Balanced\, Accuracy = \frac{Recall\ + Specificity}{2}
```
**Precision:** Precision measures the proportion of correctly predicted hate samples among all samples predicted as hate. It evaluates the modelβs ability to reduce false positive predictions. Precision was used because false hate predictions can negatively affect classification reliability.
``` math
Precision = \frac{TP}{TP + FP}
```
**Recall (Sensitivity):** Recall measures the proportion of actual hate samples correctly identified by the model. It evaluates the modelβs ability to detect hate speech effectively. Recall was important because missing harmful content may reduce system effectiveness.
``` math
Recall = \frac{TP}{TP + FN}
```
**Specificity:** Specificity measures the proportion of correctly identified non-hate samples. It evaluates how effectively the model avoids false hate predictions for normal text. This metric was used to ensure balanced non-hate classification performance.
``` math
Specificity = \frac{TN}{TN + FP}
```
**F1-Score:** F1-Score is the harmonic mean of Precision and Recall. It provides a balanced evaluation when both false positives and false negatives are important. F1-score was used because hate speech datasets often require balanced detection capability.
``` math
F1 = \frac{2 \times Precision \times Recall}{Precision + Recall}
```
**AUC-ROC:** AUC-ROC measures the modelβs ability to distinguish between hate and non-hate classes across different classification thresholds. Higher AUC values indicate better discrimination capability. This metric was used to evaluate overall classification robustness and threshold-independent performance.
3.7 TOOLS AND FRAMEWORKS
The development and evaluation of the Hindi-English code-mixed involved a combination of tools, libraries, and frameworks from both speech processing and natural language processing domains. The following are the major tools and frameworks utilized throughout the project:
1. Python was used as the primary programming language for dataset preprocessing, model implementation, training, and evaluation.
2. Google Colab was used for executing experiments with GPU support and cloud-based computation.
3. Jupyter Notebook was used for interactive coding, experimentation, and result visualization.
4. Pandas was used for data loading, preprocessing, cleaning, and tabular data manipulation.
5. NumPy was used for numerical computations and array-based operations.
6. Regex was used for detecting and removing URLs, mentions, hashtags, and noisy textual patterns.
7. NLTK was used for tokenization and text preprocessing operations.
8. Scikit-learn was used for dataset splitting, evaluation metrics, and machine learning utilities.
9. PyTorch was used for implementing transformer-based and deep learning models.
10. TensorFlow and Keras were used for implementing LSTM, BiLSTM, and neural network architectures.
11. Transformers was used for loading and fine-tuning transformer models such as MuRIL, mBART, MPNet, and HingRoBERTa.
12. Gensim was used for generating Word2Vec and FastText embeddings.
13. LightGBM was used for machine learning-based classification using boosted decision trees.
14. Universal Sentence Encoder was used for generating sentence-level semantic embeddings.
15. Matplotlib and Seaborn were used for plotting graphs, pie charts, and correlation heatmaps for dataset analysis and visualization.
##
##
##
##
## **CHAPTER 4**
##
##
##
##
##
##
##
##
##
## **RESULTS AND DISCUSSION**
##
## **4.1 RESULTS**
## **4.1.1 REGULAR TRAINING**
##
## The performance of the Code-Mixed dataset model was evaluated using Accuracy, Balanced Accuracy, Precision, Recall, Specificity, F1-Score and AUC-ROC across different hybrid models for comparison for the ground truth. The USE+LSTM model initially performed highest on the code-mixed dataset with 68% of accuracy and 59%approx in f1-score. GloVe+LightGBM scored the least in the scale comparison to other hybrid models with 65% accuracy and 57% f1 score.
Table 4.1: Regular Result
| **Models** | **accuracy** | **bal_acc** | **precision** | **recall** | **specificity** | **f1_score** | **auc_roc** |
|----|---:|---:|---:|---:|---:|---:|---:|
| **Word2vec+LSTM** | **0.6675** | **0.6573** | **0.6914** | **0.5133** | **0.8012** | **0.5892** | **0.7294** |
| **GloVe+ LSTM** | **0.6779** | **0.6627** | **0.7591** | **0.4491** | **0.8763** | **0.5644** | **0.7532** |
| **FastText+LSTM** | **0.661** | **0.6496** | **0.6906** | **0.4895** | **0.8097** | **0.5729** | **0.7208** |
| **USE+LSTM** | **0.6805** | **0.6684** | **0.7283** | **0.4981** | **0.8388** | **0.5916** | **0.7555** |
| **ELMo+LSTM** | **0.6632** | **0.6503** | **0.7084** | **0.4674** | **0.8331** | **0.5632** | **0.7327** |
| **Word2Vec+LightGBM** | **0.6665** | **0.6548** | **0.7025** | **0.4893** | **0.8203** | **0.5768** | **0.7397** |
| **GloVe+LightGBM** | **0.6527** | **0.6403** | **0.6863** | **0.465** | **0.8156** | **0.5544** | **0.717** |
| **FastText+LightGBM** | **0.676** | **0.6643** | **0.7173** | **0.4993** | **0.8293** | **0.5888** | **0.7478** |
| **USE+LightGBM** | **0.6739** | **0.6619** | **0.716** | **0.4937** | **0.8302** | **0.5844** | **0.7462** |
| **ELMo+LightGBM** | **0.6744** | **0.6636** | **0.6808** | **0.6744** | **0.8156** | **0.6658** | **0.7469** |
##
## **4.1.2 LANUAGE WISE REGULAR TRAINING**
## This is a one of a kind strategy technique that involves Language wise regular finetuning and it used the same performance metrics that is being used in Table 4.1. This work consists of various deep learning models with different background architectures and hybrid models. Here MPNet achieved highest accuracy of 74% and highest f1-score with 72% when all languages combined.
## **Table 4.2: English Strategy**
| **Model Name** | **Accuracy** | **Bal Acc** | **Precision** | **Recall** | **Specificity** | **F1 Score** | **AUC-ROC** |
|:--:|:--:|:--:|:--:|:--:|:--:|:--:|:--:|
| **mBART** | **0.835074** | **0.835074** | **0.834813** | **0.835556** | **0.834593** | **0.835184** | **0.912011** |
| **MuRIL** | **0.814625** | **0.81463** | **0.828996** | **0.792889** | **0.836372** | **0.810541** | **0.896677** |
| **HingRoBERTa** | **0.835519** | **0.835526** | **0.858159** | **0.804** | **0.867052** | **0.830197** | **0.917912** |
| **MPNet** | **0.822183** | **0.822181** | **0.817426** | **0.829778** | **0.814584** | **0.823555** | **0.899569** |
| **GloVe+BiLSTM** | **0.776238** | **0.776579** | **0.756138** | **0.808274** | **0.744885** | **0.781337** | **0.854612** |
| **Word2Vec+BiLSTM** | **0.719049** | **0.719050** | **0.722072** | **0.712444** | **0.725656** | **0.717226** | **0.797342** |
| **FastText+BiLSTM** | **0.752355** | **0.750867** | **0.740920** | **0.798956** | **0.702778** | **0.768844** | **0.825146** |
## **Table 4.3: Hindi Strategy**
| **Model Name** | **Accuracy** | **Bal Acc** | **Precision** | **Recall** | **Specificity** | **F1 Score** | **AUC-ROC** |
|:--:|:--:|:--:|:--:|:--:|:--:|:--:|:--:|
| **mBART** | **0.567762** | **0.590282** | **0.510024** | **0.799847** | **0.380717** | **0.622872** | **0.622511** |
| **MuRIL** | **0.607803** | **0.608876** | **0.554258** | **0.618865** | **0.598888** | **0.584783** | **0.67028** |
| **HingRoBERTa** | **0.605407** | **0.605746** | **0.55254** | **0.608896** | **0.602596** | **0.579351** | **0.631553** |
| **MPNet** | **0.599932** | **0.597676** | **0.549306** | **0.576687** | **0.618665** | **0.562664** | **0.611597** |
| **GloVe+BiLSTM** | **0.550690** | **0.500000** | **0.000000** | **0.000000** | **1.000000** | **0.000000** | **0.486025** |
| **Word2Vec+BiLSTM** | **0.600274** | **0.579456** | **0.578161** | **0.385736** | **0.773177** | **0.462741** | **0.604954** |
| **FastText+BiLSTM** | **0.645194** | **0.622978** | **0.612500** | **0.467780** | **0.778175** | **0.530447** | **0.673262** |
##
## **Table 4.4: Hinglish Strategy**
| **Model** | **Accuracy** | **Bal Acc** | **Precision** | **Recall** | **Specificity** | **F1 Score** | **AUC-ROC** |
|:--:|:--:|:--:|:--:|:--:|:--:|:--:|:--:|
| **mBART** | **0.706909** | **0.671073** | **0.662005** | **0.50805** | **0.834096** | **0.574899** | **0.732349** |
| **MuRIL** | **0.679693** | **0.669716** | **0.583612** | **0.624329** | **0.715103** | **0.603284** | **0.745** |
| **HingRoBERTa** | **0.73552** | **0.709035** | **0.688285** | **0.588551** | **0.829519** | **0.634523** | **0.778005** |
| **MPNet** | **0.717376** | **0.681911** | **0.679907** | **0.520572** | **0.843249** | **0.589666** | **0.753845** |
| **GloVe+BiLSTM** | **0.697939** | **0.638326** | **0.772000** | **0.344029** | **0.932624** | **0.475956** | **0.705756** |
| **Word2Vec+BiLSTM** | **0.706211** | **0.662119** | **0.682540** | **0.461538** | **0.862700** | **0.550694** | **0.740799** |
| **FastText+BiLSTM** | **0.691358** | **0.619592** | **0.710843** | **0.318919** | **0.920266** | **0.440299** | **0.688354** |
##
## **Table 4.5: Combined Strategy**
| **Model** | **Accuracy** | **Bal Acc** | **Precision** | **Recall** | **Specificity** | **F1 Score** | **AUC-ROC** |
|:--:|:--:|:--:|:--:|:--:|:--:|:--:|:--:|
| **mBART** | **0.734184** | **0.734024** | **0.706504** | **0.731761** | **0.736287** | **0.718911** | **0.827426** |
| **MuRIL** | **0.721532** | **0.721742** | **0.690934** | **0.724708** | **0.718776** | **0.707418** | **0.804714** |
| **HingRoBERTa** | **0.748531** | **0.7432** | **0.761364** | **0.668045** | **0.818354** | **0.711658** | **0.83382** |
| **MPNet** | **0.74164** | **0.741131** | **0.716694** | **0.733949** | **0.748312** | **0.725219** | **0.82881** |
| **GloVe+BiLSTM** | **0.683009** | **0.677217** | **0.681793** | **0.595574** | **0.758861** | **0.635774** | **0.763671** |
| **Word2Vec+BiLSTM** | **0.670357** | **0.662825** | **0.676418** | **0.556663** | **0.768987** | **0.610726** | **0.736175** |
| **FastText+BiLSTM** | **0.676610** | **0.665650** | **0.710953** | **0.511679** | **0.819620** | **0.595076** | **0.754570** |
## **4.1.3 MULTI-STAGE LANGUAGE TRAINING**
## This work includes the multi-stage language training with different models by which it means that instead of regular fine tuning this training goes through sequential starts initial from English then Hinglish then Hindi and at last all combined. GloVe+BiLSTM has gained the highest accuracy with 82% and 80% f1-score.
## **Table 4.6: English Strategy**
| **Model Name** | **Accuracy** | **Bal Acc** | **Precision** | **Recall** | **Specificity** | **F1 Score** | **AUC-ROC** |
|:--:|:--:|:--:|:--:|:--:|:--:|:--:|:--:|
| **mBART** | **0.829483** | **0.829266** | **0.840185** | **0.809164** | **0.849369** | **0.824383** | **0.897722** |
| **MuRIL** | **0.796480** | **0.796017** | **0.820650** | **0.753114** | **0.838920** | **0.785433** | **0.880163** |
| **HingRoBERTa** | **0.832563** | **0.832665** | **0.823401** | **0.842082** | **0.823248** | **0.832637** | **0.906317** |
| **MPNet** | **0.816502** | **0.816566** | **0.809545** | **0.822509** | **0.810623** | **0.815975** | **0.893683** |
| **GloVe+BiLSTM** | **0.6106** | **0.6258** | **0.5532** | **0.8407** | **0.4110** | **0.6673** | **0.6250** |
| **Word2Vec+BiLSTM** | **0.587438** | **0.588573** | **0.550975** | **0.60457** | **0.572574** | **0.576531** | **0.636518** |
##
## **Table 4.7: Hindi Strategy**
| **Model Name** | **Accuracy** | **Bal Acc** | **Precision** | **Recall** | **Specificity** | **F1 Score** | **AUC-ROC** |
|:--:|:--:|:--:|:--:|:--:|:--:|:--:|:--:|
| **mBART** | **0.600345** | **0.594817** | **0.556962** | **0.540292** | **0.649343** | **0.5485** | **0.645617** |
| **MuRIL** | **0.598621** | **0.588589** | **0.561126** | **0.489639** | **0.687539** | **0.522951** | **0.637667** |
| **HingRoBERTa** | **0.598966** | **0.595967** | **0.552395** | **0.566385** | **0.625548** | **0.559303** | **0.631097** |
| **MPNet** | **0.605172** | **0.602944** | **0.55826** | **0.580967** | **0.624922** | **0.569387** | **0.632868** |
| **GloVe+BiLSTM** | **0.5276** | **0.5382** | **0.4939** | **0.6885** | **0.3880** | **0.5752** | **0.5192** |
| **Word2Vec+BiLSTM** | **0.624492** | **0.624139** | **0.591543** | **0.619163** | **0.629114** | **0.605038** | **0.673203** |
| **FastText+BiLSTM** | **0.585889** | **0.585743** | **0.514705** | **0.584725** | **0.586762** | **0.547486** | **0.612223** |
##
## **Table 4.8: Hinglish Strategy**
| **Model** | **Accuracy** | **Bal Acc** | **Precision** | **Recall** | **Specificity** | **F1 Score** | **AUC-ROC** |
|:--:|:--:|:--:|:--:|:--:|:--:|:--:|:--:|
| **mBART** | **0.687278** | **0.660987** | **0.627368** | **0.531194** | **0.79078** | **0.57529** | **0.696596** |
| **MuRIL** | **0.697939** | **0.657843** | **0.678947** | **0.459893** | **0.855792** | **0.548353** | **0.691189** |
| **HingRoBERTa** | **0.707889** | **0.702147** | **0.623762** | **0.673797** | **0.730496** | **0.647815** | **0.750452** |
| **MPNet** | **0.685856** | **0.676619** | **0.601019** | **0.631016** | **0.722222** | **0.615652** | **0.738968** |
| **GloVe+BiLSTM** | **0.5198** | **0.5147** | **0.4816** | **0.4431** | **0.5863** | **0.4616** | **0.5225** |
| **Word2Vec+BiLSTM** | **0.566539** | **0.536258** | **0.72** | **0.109436** | **0.96308** | **0.189994** | **0.577701** |
| **FastText+BiLSTM** | **0.728395** | **0.651575** | **0.884057** | **0.329729** | **0.973421** | **0.480314** | **0.687276** |
##
## **Table 4.9: Combined Strategy**
| **Model** | **Accuracy** | **Bal Acc** | **Precision** | **Recall** | **Specificity** | **F1 Score** | **AUC-ROC** |
|:--:|:--:|:--:|:--:|:--:|:--:|:--:|:--:|
| **mBART** | **0.731812** | **0.72878** | **0.722592** | **0.686041** | **0.771519** | **0.703842** | **0.802136** |
| **MuRIL** | **0.715996** | **0.710274** | **0.723184** | **0.629621** | **0.790928** | **0.673167** | **0.795080** |
| **HingRoBERTa** | **0.736218** | **0.735923** | **0.709502** | **0.731761** | **0.740084** | **0.720460** | **0.820400** |
| **MPNet** | **0.726502** | **0.726061** | **0.699929** | **0.719844** | **0.732278** | **0.709747** | **0.810777** |
| **GloVe+BiLSTM** | **0.8204** | **0.8186** | **0.8149** | **0.7935** | **0.8437** | **0.8041** | **0.9139** |
| **Word2Vec+BiLSTM** | **0.730456** | **0.726676** | **0.72639** | **0.673395** | **0.779958** | **0.698889** | **0.809126** |
| **FastText+BiLSTM** | **0.701762** | **0.700688** | **0.676668** | **0.685554** | **0.715822** | **0.681082** | **0.770482** |
## **4.1.4 MULTI-STAGE LANGUAGE TRAINING WITH SIX VARIATIONS ON MuRIL AND GloVe + BiLSTM**
## This work involves two models MuRIL and Glove+BiLSTM for model training that includes multi-stage language training of six variations using (English, Hinglish, Hindi, all combined), (English, Hindi, Hinglish, all combined ), (Hinglish, English, Hindi, all combined), (Hinglish, Hindi, English, all combined), (Hindi, Hinglish, English, all combined), and (Hindi, English, Hinglish, all combined) to check which variation performs well on which model. So, we have achieved the highest 72% accuracy with Hinglish, Hindi, English, all combined variation from MuRIL model and highest accuracy with 66% in Hindi, Hinglish, English, all combined variation using GloVe+BiLSTM model.
## **Table 4.10: Combined Strategy**
<table>
<colgroup>
<col style="width: 16%" />
<col style="width: 12%" />
<col style="width: 6%" />
<col style="width: 9%" />
<col style="width: 11%" />
<col style="width: 9%" />
<col style="width: 6%" />
<col style="width: 10%" />
<col style="width: 6%" />
<col style="width: 8%" />
</colgroup>
<thead>
<tr>
<th><strong>Model</strong></th>
<th><strong>Strategy</strong></th>
<th><strong>Phase</strong></th>
<th><strong>Accuracy</strong></th>
<th><strong>Balanced Acc</strong></th>
<th><strong>Precision</strong></th>
<th><strong>Recall</strong></th>
<th><strong>Specificity</strong></th>
<th><strong>F1</strong></th>
<th><strong>ROC-AUC</strong></th>
</tr>
</thead>
<tbody>
<tr>
<td rowspan="24"><strong>BiLSTM+GloVe</strong></td>
<td rowspan="4"><strong>E -> HG -> H -> F</strong></td>
<td><strong>E</strong></td>
<td><strong>0.7498</strong></td>
<td><strong>0.7503</strong></td>
<td><strong>0.7243</strong></td>
<td><strong>0.7980</strong></td>
<td><strong>0.7027</strong></td>
<td><strong>0.7594</strong></td>
<td><strong>0.8190</strong></td>
</tr>
<tr>
<td><strong>HG</strong></td>
<td><strong>0.4222</strong></td>
<td><strong>0.4862</strong></td>
<td><strong>0.3906</strong></td>
<td><strong>0.8021</strong></td>
<td><strong>0.1702</strong></td>
<td><strong>0.5254</strong></td>
<td><strong>0.4919</strong></td>
</tr>
<tr>
<td><strong>H</strong></td>
<td><strong>0.4759</strong></td>
<td><strong>0.5007</strong></td>
<td><strong>0.4498</strong></td>
<td><strong>0.7460</strong></td>
<td><strong>0.2555</strong></td>
<td><strong>0.5612</strong></td>
<td><strong>0.5025</strong></td>
</tr>
<tr>
<td><strong>F</strong></td>
<td><strong>0.6080</strong></td>
<td><strong>0.6195</strong></td>
<td><strong>0.5554</strong></td>
<td><strong>0.7821</strong></td>
<td><strong>0.4570</strong></td>
<td><strong>0.6496</strong></td>
<td><strong>0.7023</strong></td>
</tr>
<tr>
<td rowspan="4"><strong>E -> H -> HG -> F</strong></td>
<td><strong>E</strong></td>
<td><strong>0.7391</strong></td>
<td><strong>0.7392</strong></td>
<td><strong>0.7301</strong></td>
<td><strong>0.7496</strong></td>
<td><strong>0.7288</strong></td>
<td><strong>0.7397</strong></td>
<td><strong>0.8147</strong></td>
</tr>
<tr>
<td><strong>HG</strong></td>
<td><strong>0.4890</strong></td>
<td><strong>0.5081</strong></td>
<td><strong>0.4053</strong></td>
<td><strong>0.6025</strong></td>
<td><strong>0.4137</strong></td>
<td><strong>0.4846</strong></td>
<td><strong>0.4989</strong></td>
</tr>
<tr>
<td><strong>H</strong></td>
<td><strong>0.5731</strong></td>
<td><strong>0.5428</strong></td>
<td><strong>0.5569</strong></td>
<td><strong>0.2441</strong></td>
<td><strong>0.8416</strong></td>
<td><strong>0.3394</strong></td>
<td><strong>0.5863</strong></td>
</tr>
<tr>
<td><strong>F</strong></td>
<td><strong>0.6449</strong></td>
<td><strong>0.6399</strong></td>
<td><strong>0.6305</strong></td>
<td><strong>0.5693</strong></td>
<td><strong>0.7105</strong></td>
<td><strong>0.5983</strong></td>
<td><strong>0.7163</strong></td>
</tr>
<tr>
<td rowspan="4"><strong>HG -> E -> H -> F</strong></td>
<td><strong>E</strong></td>
<td><strong>0.7393</strong></td>
<td><strong>0.7392</strong></td>
<td><strong>0.7370</strong></td>
<td><strong>0.7353</strong></td>
<td><strong>0.7431</strong></td>
<td><strong>0.7361</strong></td>
<td><strong>0.8203</strong></td>
</tr>
<tr>
<td><strong>HG</strong></td>
<td><strong>0.5942</strong></td>
<td><strong>0.5202</strong></td>
<td><strong>0.4728</strong></td>
<td><strong>0.1551</strong></td>
<td><strong>0.8853</strong></td>
<td><strong>0.2336</strong></td>
<td><strong>0.5822</strong></td>
</tr>
<tr>
<td><strong>H</strong></td>
<td><strong>0.6079</strong></td>
<td><strong>0.5937</strong></td>
<td><strong>0.5817</strong></td>
<td><strong>0.4536</strong></td>
<td><strong>0.7339</strong></td>
<td><strong>0.5097</strong></td>
<td><strong>0.6320</strong></td>
</tr>
<tr>
<td><strong>F</strong></td>
<td><strong>0.6732</strong></td>
<td><strong>0.6661</strong></td>
<td><strong>0.6770</strong></td>
<td><strong>0.5669</strong></td>
<td><strong>0.7654</strong></td>
<td><strong>0.6171</strong></td>
<td><strong>0.7438</strong></td>
</tr>
<tr>
<td rowspan="4"><strong>HG -> H -> E -> F</strong></td>
<td><strong>E</strong></td>
<td><strong>0.6854</strong></td>
<td><strong>0.6855</strong></td>
<td><strong>0.6754</strong></td>
<td><strong>0.7006</strong></td>
<td><strong>0.6704</strong></td>
<td><strong>0.6878</strong></td>
<td><strong>0.7586</strong></td>
</tr>
<tr>
<td><strong>HG</strong></td>
<td><strong>0.6347</strong></td>
<td><strong>0.5581</strong></td>
<td><strong>0.6516</strong></td>
<td><strong>0.1800</strong></td>
<td><strong>0.9362</strong></td>
<td><strong>0.2821</strong></td>
<td><strong>0.5664</strong></td>
</tr>
<tr>
<td><strong>H</strong></td>
<td><strong>0.6372</strong></td>
<td><strong>0.6248</strong></td>
<td><strong>0.6187</strong></td>
<td><strong>0.5019</strong></td>
<td><strong>0.7477</strong></td>
<td><strong>0.5542</strong></td>
<td><strong>0.6598</strong></td>
</tr>
<tr>
<td><strong>F</strong></td>
<td><strong>0.6615</strong></td>
<td><strong>0.6553</strong></td>
<td><strong>0.6574</strong></td>
<td><strong>0.5666</strong></td>
<td><strong>0.7439</strong></td>
<td><strong>0.6087</strong></td>
<td><strong>0.7110</strong></td>
</tr>
<tr>
<td rowspan="4"><strong>H -> HG -> E -> F</strong></td>
<td><strong>E</strong></td>
<td><strong>0.5963</strong></td>
<td><strong>0.5965</strong></td>
<td><strong>0.5878</strong></td>
<td><strong>0.6152</strong></td>
<td><strong>0.5777</strong></td>
<td><strong>0.6012</strong></td>
<td><strong>0.6335</strong></td>
</tr>
<tr>
<td><strong>HG</strong></td>
<td><strong>0.6084</strong></td>
<td><strong>0.5362</strong></td>
<td><strong>0.5260</strong></td>
<td><strong>0.1800</strong></td>
<td><strong>0.8924</strong></td>
<td><strong>0.2683</strong></td>
<td><strong>0.5194</strong></td>
</tr>
<tr>
<td><strong>H</strong></td>
<td><strong>0.6331</strong></td>
<td><strong>0.6257</strong></td>
<td><strong>0.5995</strong></td>
<td><strong>0.5526</strong></td>
<td><strong>0.6988</strong></td>
<td><strong>0.5751</strong></td>
<td><strong>0.6660</strong></td>
</tr>
<tr>
<td><strong>F</strong></td>
<td><strong>0.6103</strong></td>
<td><strong>0.6053</strong></td>
<td><strong>0.5884</strong></td>
<td><strong>0.5360</strong></td>
<td><strong>0.6747</strong></td>
<td><strong>0.5610</strong></td>
<td><strong>0.6455</strong></td>
</tr>
<tr>
<td rowspan="4"><strong>H -> E -> HG -> F</strong></td>
<td><strong>E</strong></td>
<td><strong>0.7193</strong></td>
<td><strong>0.7188</strong></td>
<td><strong>0.7359</strong></td>
<td><strong>0.6744</strong></td>
<td><strong>0.7632</strong></td>
<td><strong>0.7038</strong></td>
<td><strong>0.8011</strong></td>
</tr>
<tr>
<td><strong>HG</strong></td>
<td><strong>0.5572</strong></td>
<td><strong>0.4940</strong></td>
<td><strong>0.3835</strong></td>
<td><strong>0.1818</strong></td>
<td><strong>0.8061</strong></td>
<td><strong>0.2467</strong></td>
<td><strong>0.5129</strong></td>
</tr>
<tr>
<td><strong>H</strong></td>
<td><strong>0.6272</strong></td>
<td><strong>0.6165</strong></td>
<td><strong>0.6002</strong></td>
<td><strong>0.5104</strong></td>
<td><strong>0.7226</strong></td>
<td><strong>0.5516</strong></td>
<td><strong>0.6606</strong></td>
</tr>
<tr>
<td><strong>F</strong></td>
<td><strong>0.6634</strong></td>
<td><strong>0.6562</strong></td>
<td><strong>0.6648</strong></td>
<td><strong>0.5552</strong></td>
<td><strong>0.7572</strong></td>
<td><strong>0.6051</strong></td>
<td><strong>0.7297</strong></td>
</tr>
<tr>
<td rowspan="24"><strong>MuRIL</strong></td>
<td rowspan="4"><strong>E -> HG -> H -> F</strong></td>
<td><strong>E</strong></td>
<td><strong>0.7965</strong></td>
<td><strong>0.7960</strong></td>
<td><strong>0.8206</strong></td>
<td><strong>0.7531</strong></td>
<td><strong>0.8389</strong></td>
<td><strong>0.7854</strong></td>
<td><strong>0.8802</strong></td>
</tr>
<tr>
<td><strong>HG</strong></td>
<td><strong>0.6979</strong></td>
<td><strong>0.6578</strong></td>
<td><strong>0.6789</strong></td>
<td><strong>0.4599</strong></td>
<td><strong>0.8558</strong></td>
<td><strong>0.5484</strong></td>
<td><strong>0.6912</strong></td>
</tr>
<tr>
<td><strong>H</strong></td>
<td><strong>0.5986</strong></td>
<td><strong>0.5886</strong></td>
<td><strong>0.5611</strong></td>
<td><strong>0.4896</strong></td>
<td><strong>0.6875</strong></td>
<td><strong>0.5230</strong></td>
<td><strong>0.6377</strong></td>
</tr>
<tr>
<td><strong>F</strong></td>
<td><strong>0.7160</strong></td>
<td><strong>0.7103</strong></td>
<td><strong>0.7232</strong></td>
<td><strong>0.6296</strong></td>
<td><strong>0.7909</strong></td>
<td><strong>0.6732</strong></td>
<td><strong>0.7951</strong></td>
</tr>
<tr>
<td rowspan="4"><strong>E -> H -> HG -> F</strong></td>
<td><strong>E</strong></td>
<td><strong>0.7985</strong></td>
<td><strong>0.7977</strong></td>
<td><strong>0.8422</strong></td>
<td><strong>0.7291</strong></td>
<td><strong>0.8663</strong></td>
<td><strong>0.7816</strong></td>
<td><strong>0.8645</strong></td>
</tr>
<tr>
<td><strong>HG</strong></td>
<td><strong>0.6645</strong></td>
<td><strong>0.6478</strong></td>
<td><strong>0.5817</strong></td>
<td><strong>0.5651</strong></td>
<td><strong>0.7305</strong></td>
<td><strong>0.5732</strong></td>
<td><strong>0.7036</strong></td>
</tr>
<tr>
<td><strong>H</strong></td>
<td><strong>0.5793</strong></td>
<td><strong>0.5803</strong></td>
<td><strong>0.5285</strong></td>
<td><strong>0.5902</strong></td>
<td><strong>0.5704</strong></td>
<td><strong>0.5577</strong></td>
<td><strong>0.6034</strong></td>
</tr>
<tr>
<td><strong>F</strong></td>
<td><strong>0.7054</strong></td>
<td><strong>0.7025</strong></td>
<td><strong>0.6906</strong></td>
<td><strong>0.6627</strong></td>
<td><strong>0.7424</strong></td>
<td><strong>0.6763</strong></td>
<td><strong>0.7782</strong></td>
</tr>
<tr>
<td rowspan="4"><strong>HG -> E -> H -> F</strong></td>
<td><strong>E</strong></td>
<td><strong>0.8090</strong></td>
<td><strong>0.8089</strong></td>
<td><strong>0.8103</strong></td>
<td><strong>0.8016</strong></td>
<td><strong>0.8163</strong></td>
<td><strong>0.8059</strong></td>
<td><strong>0.8318</strong></td>
</tr>
<tr>
<td><strong>HG</strong></td>
<td><strong>0.6660</strong></td>
<td><strong>0.6318</strong></td>
<td><strong>0.6061</strong></td>
<td><strong>0.4635</strong></td>
<td><strong>0.8002</strong></td>
<td><strong>0.5253</strong></td>
<td><strong>0.6622</strong></td>
</tr>
<tr>
<td><strong>H</strong></td>
<td><strong>0.5931</strong></td>
<td><strong>0.5868</strong></td>
<td><strong>0.5494</strong></td>
<td><strong>0.5249</strong></td>
<td><strong>0.6487</strong></td>
<td><strong>0.5369</strong></td>
<td><strong>0.6083</strong></td>
</tr>
<tr>
<td><strong>F</strong></td>
<td><strong>0.7155</strong></td>
<td><strong>0.7124</strong></td>
<td><strong>0.7045</strong></td>
<td><strong>0.6678</strong></td>
<td><strong>0.7570</strong></td>
<td><strong>0.6856</strong></td>
<td><strong>0.7442</strong></td>
</tr>
<tr>
<td rowspan="4"><strong>HG -> H -> E -> F</strong></td>
<td><strong>E</strong></td>
<td><strong>0.7985</strong></td>
<td><strong>0.7977</strong></td>
<td><strong>0.8422</strong></td>
<td><strong>0.7291</strong></td>
<td><strong>0.8663</strong></td>
<td><strong>0.7816</strong></td>
<td><strong>0.8779</strong></td>
</tr>
<tr>
<td><strong>HG</strong></td>
<td><strong>0.6830</strong></td>
<td><strong>0.6547</strong></td>
<td><strong>0.6242</strong></td>
<td><strong>0.5152</strong></td>
<td><strong>0.7943</strong></td>
<td><strong>0.5645</strong></td>
<td><strong>0.7030</strong></td>
</tr>
<tr>
<td><strong>H</strong></td>
<td><strong>0.6279</strong></td>
<td><strong>0.6165</strong></td>
<td><strong>0.6029</strong></td>
<td><strong>0.5035</strong></td>
<td><strong>0.7295</strong></td>
<td><strong>0.5487</strong></td>
<td><strong>0.6506</strong></td>
</tr>
<tr>
<td><strong>F</strong></td>
<td><strong>0.7242</strong></td>
<td><strong>0.7179</strong></td>
<td><strong>0.7389</strong></td>
<td><strong>0.6284</strong></td>
<td><strong>0.8074</strong></td>
<td><strong>0.6792</strong></td>
<td><strong>0.7981</strong></td>
</tr>
<tr>
<td rowspan="4"><strong>H -> HG -> E -> F</strong></td>
<td><strong>E</strong></td>
<td><strong>0.7901</strong></td>
<td><strong>0.7894</strong></td>
<td><strong>0.8304</strong></td>
<td><strong>0.7233</strong></td>
<td><strong>0.8555</strong></td>
<td><strong>0.7732</strong></td>
<td><strong>0.8516</strong></td>
</tr>
<tr>
<td><strong>HG</strong></td>
<td><strong>0.6802</strong></td>
<td><strong>0.6590</strong></td>
<td><strong>0.6086</strong></td>
<td><strong>0.5544</strong></td>
<td><strong>0.7636</strong></td>
<td><strong>0.5802</strong></td>
<td><strong>0.6945</strong></td>
</tr>
<tr>
<td><strong>H</strong></td>
<td><strong>0.6238</strong></td>
<td><strong>0.6179</strong></td>
<td><strong>0.5849</strong></td>
<td><strong>0.5602</strong></td>
<td><strong>0.6756</strong></td>
<td><strong>0.5723</strong></td>
<td><strong>0.6334</strong></td>
</tr>
<tr>
<td><strong>F</strong></td>
<td><strong>0.7181</strong></td>
<td><strong>0.7135</strong></td>
<td><strong>0.7175</strong></td>
<td><strong>0.6486</strong></td>
<td><strong>0.7785</strong></td>
<td><strong>0.6813</strong></td>
<td><strong>0.7770</strong></td>
</tr>
<tr>
<td rowspan="4"><strong>H -> E -> HG -> F</strong></td>
<td><strong>E</strong></td>
<td><strong>0.7668</strong></td>
<td><strong>0.7654</strong></td>
<td><strong>0.8502</strong></td>
<td><strong>0.6415</strong></td>
<td><strong>0.8894</strong></td>
<td><strong>0.7312</strong></td>
<td><strong>0.8083</strong></td>
</tr>
<tr>
<td><strong>HG</strong></td>
<td><strong>0.6915</strong></td>
<td><strong>0.6486</strong></td>
<td><strong>0.6749</strong></td>
<td><strong>0.4367</strong></td>
<td><strong>0.8605</strong></td>
<td><strong>0.5303</strong></td>
<td><strong>0.6898</strong></td>
</tr>
<tr>
<td><strong>H</strong></td>
<td><strong>0.6193</strong></td>
<td><strong>0.5971</strong></td>
<td><strong>0.6268</strong></td>
<td><strong>0.3776</strong></td>
<td><strong>0.8165</strong></td>
<td><strong>0.4713</strong></td>
<td><strong>0.6455</strong></td>
</tr>
<tr>
<td><strong>F</strong></td>
<td><strong>0.7065</strong></td>
<td><strong>0.6948</strong></td>
<td><strong>0.7662</strong></td>
<td><strong>0.5299</strong></td>
<td><strong>0.8597</strong></td>
<td><strong>0.6265</strong></td>
<td><strong>0.7508</strong></td>
</tr>
</tbody>
</table>
## **4.1.5 SARVAM MODEL \***
We have used SARVAM AI Model which is an Indian startup specially made on and made for Indian languages. It is a LLM capable of speech-to-text, and text-to-speech. SARVAM has different models with different parameters but we have implemented Sarvam-1 that consists of 2 Billion parameters. It supports 22+ Indian languages with different scripts. Sarvam can also well handle the code-mixed texts so we implemented our dataset with Sarvam for model training an evaluation as an experimental work with llms and we have achieved highest performance in all 6 metrics.
Table 4.11: Performance Evaluation
<table style="width:54%;">
<colgroup>
<col style="width: 37%" />
<col style="width: 16%" />
</colgroup>
<thead>
<tr>
<th><p><strong>Β </strong></p>
<p><strong>Metrics</strong></p></th>
<th><strong>Value</strong></th>
</tr>
</thead>
<tbody>
<tr>
<td><strong>Accuracy</strong></td>
<td><strong>0.9373</strong></td>
</tr>
<tr>
<td><strong>Balanced Accuracy</strong></td>
<td><strong>0.9284</strong></td>
</tr>
<tr>
<td><strong>Precision</strong></td>
<td><strong>0.9486</strong></td>
</tr>
<tr>
<td><strong>Recall</strong></td>
<td><strong>0.8877</strong></td>
</tr>
<tr>
<td><strong>Specificity</strong></td>
<td><strong>0.9691</strong></td>
</tr>
<tr>
<td><strong>F1-Score</strong></td>
<td><strong>0.9171</strong></td>
</tr>
<tr>
<td><strong>ROC-AUC</strong></td>
<td><strong>0.9326</strong></td>
</tr>
</tbody>
</table>
## **4.2 VISUALIZATION**
## **4.2.1 REGULAR TRAINING**
This work involves total 39 figures for visualisation that uses two variations of hybrid models one is (Word2Vec, GloVe, FastText, USE and ELMo) with LSTM and another with LightGBM. Fig:4.1 β Fig: 4.10 consists of Train vs Validation Accuracy Curve and Train vs Validation Loss Curve, Fig:4.11 β Fig: 4.20 consists of AUC-ROC Curves, FIG:4.21 β Fig:4.30 consists of Confusion Matrix and Fig:4.31-Fig:4.39 is t-SNE Visualisation.
<img src="thesis_media/media/image10.png" style="width:3.06944in;height:2.39296in" /><img src="thesis_media/media/image11.png" style="width:3.00717in;height:2.375in" />
Fig:4.1 Training vs Validation Loss and Accuracy Word2Vec+LSTM
<img src="thesis_media/media/image12.png" style="width:2.9021in;height:2.2724in" /><img src="thesis_media/media/image13.png" style="width:2.83916in;height:2.25189in" />
Fig:4.2 Training vs Validation Loss and Accuracy GloVe+LSTM
<img src="thesis_media/media/image14.png" style="width:3.15385in;height:2.36538in" /> <img src="thesis_media/media/image15.png" style="width:3.1477in;height:2.36078in" />
Fig:4.3 Training vs Validation Loss and Accuracy FastText+LSTM
<img src="thesis_media/media/image16.png" style="width:2.82517in;height:2.20261in" /> <img src="thesis_media/media/image17.png" style="width:2.7972in;height:2.20917in" />
Fig:4.4 Training vs Validation Loss and Accuracy USE+LSTM
<img src="thesis_media/media/image18.png" style="width:3.04196in;height:2.37153in" /><img src="thesis_media/media/image19.png" style="width:2.94406in;height:2.32515in" />
Fig:4.5 Training vs Validation Loss and Accuracy ELMo+LSTM
<img src="thesis_media/media/image20.png" style="width:3.07639in;height:2.46732in" /><img src="thesis_media/media/image21.png" style="width:3.16084in;height:2.63424in" />
Fig:4.6 Training vs Validation Loss and Accuracy Word2Vec+LightGBM
<img src="thesis_media/media/image22.png" style="width:3.03497in;height:2.37735in" /><img src="thesis_media/media/image23.png" style="width:3.04196in;height:2.38282in" />
Fig:4.7Training vs Validation Loss and Accuracy GloVe+LightGBM
<img src="thesis_media/media/image24.png" style="width:3.22378in;height:2.40165in" /><img src="thesis_media/media/image25.png" style="width:3.2028in;height:2.38699in" />
Fig:4.8 Training vs Validation Loss and Accuracy FastText+LightGBM
<img src="thesis_media/media/image26.png" style="width:3.11806in;height:2.32662in" /><img src="thesis_media/media/image27.png" style="width:3.1049in;height:2.31745in" />
Fig:4.9 Training vs Validation Loss and Accuracy USE+LightGBM
<img src="thesis_media/media/image28.png" style="width:3.17361in;height:2.47532in" /><img src="thesis_media/media/image29.png" style="width:3.0979in;height:2.485in" />
Fig:4.10 Training vs Validation Loss and Accuracy ELMo+LightGBM
<img src="thesis_media/media/image30.png" style="width:3.11806in;height:2.46239in" /> <img src="thesis_media/media/image31.png" style="width:3.08392in;height:2.44602in" />
Fig:4.11 ROC-AUC Curve Word2Vec+LSTM Fig:4.12 ROC-AUC Curve GloVe+LSTM
<img src="thesis_media/media/image32.png" style="width:3.35664in;height:2.51767in" /><img src="thesis_media/media/image33.png" style="width:2.99301in;height:2.36381in" />
Fig:4.13 ROC-AUC Curve FastText+LSTM Fig:4.14 ROC-AUC Curve USE+LSTM
<img src="thesis_media/media/image34.png" style="width:3.09973in;height:2.44792in" /><img src="thesis_media/media/image35.png" style="width:3.07692in;height:2.45962in" />
Fig:4.15 ROC-AUC Curve ELMo+ Fig:4.16 ROC-AUC Curve Word2Vec+LightGBM
<img src="thesis_media/media/image36.png" style="width:3.02083in;height:2.39598in" /><img src="thesis_media/media/image37.png" style="width:3.13287in;height:2.33487in" />
Fig:4.17 ROC-AUC Curve GloVe+LightGBM Fig:4.18 ROC-AUC Curve FastText+LightGBM
<img src="thesis_media/media/image38.png" style="width:3.07639in;height:2.29621in" /><img src="thesis_media/media/image39.png" style="width:2.93747in;height:2.28365in" />
Fig:4.19 ROC-AUC Curve USE+LightGBM Fig:4.20 ROC-AUC Curve ELMo+LightGBM
<img src="thesis_media/media/image40.png" style="width:3.07639in;height:2.55116in" /><img src="thesis_media/media/image41.png" style="width:3.02083in;height:2.51613in" />
Fig:4.21 Confusion Matrix Word2Vec+LSTM Fig:4.22 Confusion Matrix GloVe+LSTM
<img src="thesis_media/media/image42.png" style="width:3.20347in;height:2.40278in" /><img src="thesis_media/media/image43.png" style="width:2.90972in;height:2.41295in" />
Fig:4.23 Confusion Matrix FastText+LSTM Fig:4.24 Confusion Matrix USE+LSTM
<img src="thesis_media/media/image44.png" style="width:3.22917in;height:2.48323in" /><img src="thesis_media/media/image45.png" style="width:2.81476in;height:2.49792in" />
Fig:4.25 Confusion Matrix ELMo+LSTM Fig:4.26 Confusion Matrix Word2Vec+LightGBM
<img src="thesis_media/media/image46.png" style="width:3.05556in;height:2.54505in" /><img src="thesis_media/media/image47.png" style="width:2.80556in;height:2.53666in" />
Fig:4.27 Confusion Matrix GloVe+LightGBM Fig:4.28 Confusion Matrix FastText+LightGBM
<img src="thesis_media/media/image48.png" style="width:3.04787in;height:2.39176in" /><img src="thesis_media/media/image49.png" style="width:2.79861in;height:2.39181in" />
Fig:4.29 Confusion Matrix USE+LightGBM Fig:4.30 Confusion Matrix ELMo+LightGBM
<img src="thesis_media/media/image50.png" style="width:3.07639in;height:2.41551in" /><img src="thesis_media/media/image51.png" style="width:3.1111in;height:2.40278in" />
Fig: 4.31 t-SNE for GloVe + LSTM Fig: 4.32 t-SNE for ELMo + LSTM
<img src="thesis_media/media/image52.png" style="width:2.99696in;height:2.47168in" /> <img src="thesis_media/media/image53.png" style="width:3.23077in;height:2.5304in" />
Fig 4.33: t-SNE for Word2Vec+ LightGBM Fig:4.34: t-SNE for GloVe + LightGBM
<img src="thesis_media/media/image54.png" style="width:5.37063in;height:2.60737in" />
Fig:4.35 t-SNE for Comparison FastText + LSTM
<img src="thesis_media/media/image55.png" style="width:2.95624in;height:2.20556in" /> <img src="thesis_media/media/image56.png" style="width:3.0979in;height:2.29918in" />
Fig 4.36: t-SNE for FastText + LSTM Fig:4.37 t-SNE for Use + LightGBM
<img src="thesis_media/media/image57.png" style="width:6.26806in;height:2.75347in" />
Fig:4.38 t-SNE for cluster comparison USE + LightGBM
<img src="thesis_media/media/image58.png" style="width:3.18486in;height:2.45833in" />
Fig: 4.39 t-SNE for FastText + LSTM
##
## **4.2.2 LANGUAGE WISE REGULAR TRAINING**
This work involves total 19 figures for visualisation that uses multiplevariations of hybrid models (GloVe+BiLSTM, FastText+BiLSTM, Word2Vec+BiLSTM) and deep learning models(Mbart, MuRIL, HingRoBERTa-mixed, MPNet).but we have used only all combined(Full dataset) visualisations Fig:4.40 β Fig: 4.46 consists of Train vs Validation Accuracy Curve and Train vs Validation Loss Curve, Fig:4.47 β Fig: 4.53 consists of AUC-ROC Curves, FIG:4.54 β Fig:4.59 consists of Confusion Matrix.
<img src="thesis_media/media/image59.png" style="width:3.18182in;height:2.3198in" /><img src="thesis_media/media/image60.png" style="width:3.13287in;height:2.29218in" />
Fig:4.40 Training vs Validation Loss and Accuracy MuRIL Full dataset
<img src="thesis_media/media/image61.png" style="width:3.20979in;height:2.38357in" /><img src="thesis_media/media/image62.png" style="width:3.18881in;height:2.36185in" />
Fig:4.41 Training vs Validation Loss and Accuracy mBART Full dataset
<img src="thesis_media/media/image63.png" style="width:3.22378in;height:2.39362in" /><img src="thesis_media/media/image64.png" style="width:3.23077in;height:2.37606in" />
Fig:4.42 Training vs Validation Loss and Accuracy HingRoBERTa Full dataset
<img src="thesis_media/media/image65.png" style="width:3.12587in;height:2.31344in" /><img src="thesis_media/media/image66.png" style="width:3.13287in;height:2.31758in" />
Fig:4.43 Training vs Validation Loss and Accuracy MPNet Full dataset
<img src="thesis_media/media/image67.png" style="width:3.21727in;height:2.41295in" /><img src="thesis_media/media/image68.png" style="width:3.32407in;height:2.49306in" />
Fig:4.44 Training vs Validation Loss and Accuracy GloVe+BiLSTM Full dataset
<img src="thesis_media/media/image69.png" style="width:3.30417in;height:2.20278in" /> <img src="thesis_media/media/image70.png" style="width:3.25175in;height:2.16783in" />
Fig:4.45 Training vs Validation Loss and Accuracy Word2Vec+BiLSTM Full dataset
<img src="thesis_media/media/image71.png" style="width:3.23889in;height:2.42917in" /><img src="thesis_media/media/image72.png" style="width:3.32361in;height:2.49271in" />
Fig:4.46 Training vs Validation Loss and Accuracy FastText +BiLSTM Full dataset
<img src="thesis_media/media/image73.png" style="width:3.02797in;height:2.26424in" /> <img src="thesis_media/media/image74.png" style="width:3.03497in;height:2.26983in" />
Fig:4.47 ROC-AUC Curve MuRIL Fig:4.48 ROC-AUC Curve mBART
<img src="thesis_media/media/image75.png" style="width:3.07692in;height:2.27583in" /><img src="thesis_media/media/image76.png" style="width:3.00699in;height:2.22326in" />
Fig 4.49 ROC-AUC Curve HingRoBERTa Fig: 4.50 ROC-AUC Curve MPNet
<img src="thesis_media/media/image77.png" style="width:3.41259in;height:2.27506in" /><img src="thesis_media/media/image78.png" style="width:3.12587in;height:2.34441in" />
Fig:4.51 ROC-AUC Curve GloVe+BiLSTM Fig:4.52 ROC-AUC Curve Word2Vec+BiLSTM
<img src="thesis_media/media/image79.png" style="width:3.36364in;height:2.52273in" />
Fig:4.53 ROC-AUC Curve FastText +BiLSTM
<img src="thesis_media/media/image80.png" style="width:3.29371in;height:2.62563in" /> <img src="thesis_media/media/image81.png" style="width:3.2028in;height:2.55349in" />
Fig:4.54 Confusion Matrix MuRIL Fig:4.55 Confusion Matrix mBART
<img src="thesis_media/media/image82.png" style="width:3.45764in;height:2.7125in" /><img src="thesis_media/media/image83.png" style="width:3.24606in;height:2.62284in" />
Fig 4.56: Confusion Matrix HingRoBERTa Fig:4.57 Confusion Matrix MPNet
<img src="thesis_media/media/image84.png" style="width:3.01458in;height:2.5125in" /><img src="thesis_media/media/image85.png" style="width:3.21678in;height:2.41259in" />
Fig:4.58 Confusion Matrix GloVe + BiLSTM Fig:4.58 Confusion Matrix Word2Vec + BiLSTM
<img src="thesis_media/media/image86.png" style="width:3.39583in;height:2.54688in" />
Fig:4.59 Confusion Matrix FastText + BiLSTM
## **4.2.3 MULTI-STAGE LANGUAGE TRAINING**
This work involves total 13 figures for visualisation that uses multiple variations of hybrid models (GloVe+BiLSTM, FastText+BiLSTM, Word2Vec+BiLSTM) and deep learning models (Mbart, MuRIL, HingRoBERTa-mixed, MPNet). For multistage language training on English-\> Hinglish -\> Hindi -\> All combined, and we have used only all combined(Full dataset) visualisations Fig:4.60 β Fig: 4.66 consists of Train vs Validation Accuracy Curve and Train vs Validation Loss Curve, Fig:4.67 β Fig: 4.73 consists of AUC-ROC Curves.
<img src="thesis_media/media/image87.png" style="width:3.16783in;height:2.33624in" /><img src="thesis_media/media/image88.png" style="width:3.12587in;height:2.34252in" />
Fig:4.60 Training vs Validation Loss and Accuracy MuRIL Full dataset
<img src="thesis_media/media/image89.png" style="width:3.26573in;height:2.31164in" /><img src="thesis_media/media/image90.png" style="width:3.18182in;height:2.35762in" />
Fig:4.61 Training vs Validation Loss and Accuracy mBART Full dataset
<img src="thesis_media/media/image91.png" style="width:3.26528in;height:2.27365in" /><img src="thesis_media/media/image92.png" style="width:3.28671in;height:2.31297in" />
Fig:4.62 Training vs Validation Loss and Accuracy HingRoBERTa Full dataset
<img src="thesis_media/media/image93.png" style="width:3.11888in;height:2.23829in" /><img src="thesis_media/media/image94.png" style="width:3.18125in;height:2.2279in" />
Fig:4.63 Training vs Validation Loss and Accuracy MPNet Full dataset
<img src="thesis_media/media/image95.png" style="width:3.27972in;height:2.23462in" /> <img src="thesis_media/media/image96.png" style="width:3.3986in;height:2.29075in" />
Fig:4.64 Training vs Validation Loss and Accuracy GloVe + BiLSTM Full dataset
<img src="thesis_media/media/image97.png" style="width:3.23077in;height:2.19592in" /> <img src="thesis_media/media/image98.png" style="width:3.41958in;height:2.27404in" />
Fig:4.65 Training vs Validation Loss and Accuracy Word2Vec + BiLSTM Full dataset
<img src="thesis_media/media/image99.png" style="width:3.04895in;height:2.05446in" /><img src="thesis_media/media/image100.png" style="width:3.16469in;height:2.10577in" />
Fig:4.66 Training vs Validation Loss and Accuracy FastText + BiLSTM Full dataset
<img src="thesis_media/media/image101.png" style="width:3.07692in;height:2.54811in" /><img src="thesis_media/media/image102.png" style="width:3.12587in;height:2.53985in" />
Fig:4.67 ROC-AUC Curve MuRIL Full Dataset Fig:4.68 ROC-AUC Curve mBART Full Dataset
<img src="thesis_media/media/image103.png" style="width:3.46154in;height:2.46232in" /><img src="thesis_media/media/image104.png" style="width:3.02797in;height:2.55086in" />
Fig:4.69 ROC-AUC Curve HingRoBERTa Full Dataset Fig: 4.70 ROC-AUC Curve MPNet Full Dataset
<img src="thesis_media/media/image105.png" style="width:3.41944in;height:2.38472in" /><img src="thesis_media/media/image106.png" style="width:3.38894in;height:2.2856in" />
Fig4.71: ROC-AUC Curve GloVe + BiLSTM Fig:4.72 ROC-AUC Curve Word2Vec + BiLSTM
<img src="thesis_media/media/image107.png" style="width:3.35664in;height:2.28167in" />
Fig:4.73 ROC-AUC Curve FastText + BiLSTM Full Dataset
<img src="thesis_media/media/image108.png" style="width:3.26812in;height:2.70922in" /> <img src="thesis_media/media/image109.png" style="width:3.23932in;height:2.7029in" />
Fig:4.74 Confusion Matrix MuRIL Fig:4.75 Confusion Matrix mBART
<img src="thesis_media/media/image110.png" style="width:3.09643in;height:2.68653in" /> <img src="thesis_media/media/image111.png" style="width:3.34108in;height:2.45652in" />
Fig:4.76 Confusion Matrix HingRoBERTa Fig:4.77 Confusion Matrix MPNet
<img src="thesis_media/media/image112.png" style="width:3.09583in;height:2.85069in" /> <img src="thesis_media/media/image113.png" style="width:3.07659in;height:2.82464in" />
Fig:4.78 Confusion Matrix Word2Vec + BiLSTM Fig:4.79 Confusion Matrix FastText + BiLSTM
<img src="thesis_media/media/image114.png" style="width:3.70498in;height:3.42424in" />
Fig:4.80 Confusion Matrix GloVe + BiLSTM
<img src="thesis_media/media/image115.png" style="width:4.99624in;height:3.23021in" />
Fig: 4.81 Confusion Matrix for Sarvam \*
Β
## **4.3 DISCUSSION**
##
## The overall workflow consists of three methodologies in total, one with regular fine-tuning for hybrid models, second methodology involves of two different strategies one with regular fine-tuning and one with sequential training (English-\> Hinglish-\> Hindi-\> Full) on both hybrid models and deep learning models and the third methodology involves of multiple variations sequential training (English, Hinglish, Hindi, all combined), (English, Hindi, Hinglish, all combined ), (Hinglish, English, Hindi, all combined), (Hinglish, Hindi, English, all combined), (Hindi, Hinglish, English, all combined), and (Hindi, English, Hinglish, all combined) corpora to improve cross-lingual understanding. So, in our entire workflow we observed that third methodology took more training time in comparison to other two methodologies and gained very good accuracy and f1-score overall which was achieved by GloVe+BiLSTM model which is a classical hybrid model but performed better than deep learning models we used. But, in regular trainings deep learning models outperformed every other hybrid models. We have also observed Hindi(devnagiri) made overall performance degrade due to its less contextual understanding and tokens mishandle. For an extra activity we worked with Sarvam AI which is a LLM and we achieved highest of all metrics in our all workflows combined.
##
##
##
## **CHAPTER 5**
**CONCLUSION AND FUTURE SCOPE**
This project successfully addressed the critical need for an effective and domain-specific Code-mixed sentiment analysis system, which is an obvious unexplored area in the field of natural language processing and speech technology. The outcomes of this project open up several promising opportunities for future research and development in better contextual understandings and natural language processing for code-mixed. One of the most immediate areas of expansion is the enhancement of the dataset. Our primary aim was that a good LLM or a model with a lot of billions of parameter can handle, can understand the sarcastic contexts, but the computational cost of these models are very height. We tried to make a small differentiate with developing a sentimental analysis prediction system with those models who has less parameters and achieve a near good prediction with good performance our overall works has gained the accuracy between 72%-80% which is a good start but we will move further and works with different collections of architectures to understand the depth of mechanism for enhancements. In our study we were also introduced with Explainable AI which is a modern trend model that tends to discover the happenings inside the black box. We have worked with some types of XAI that includes Shap, Lime, Captum, Integrated Gradients and all these are good powerful models. Our future scope is to study more on hybrid models as they somehow manage to perform well then deep learning models if selected smartly and we will combine the explainable AI in order to for analyzing and handling of misclassifications happening in training.
**REFERENCE**
\[1\] Singh, G. (2021). Sentiment analysis of code-mixed social media text (Hinglish). In *arXiv \[cs.CL\]*. https://doi.org/10.48550/ARXIV.2102.12149\
\[2\] Thakur, V., Sahu, R., & Omer, S. (2020). Current state of hinglish text sentiment analysis. *SSRN Electronic Journal*. <https://doi.org/10.2139/ssrn.3614442>
\[3\] Agarwal, P. N. (2024). Improving sentiment analysis accuracy in hinglish text using hybrid deep learning approaches. *Educational Administration: Theory and Practice*, 741β750. https://doi.org/10.53555/kuey.v30i11.8739
\[4\] Singh, G. V., Ghosh, S., Firdaus, M., Ekbal, A., & Bhattacharyya, P. (2024). Predicting multi-label emojis, emotions, and sentiments in code-mixed texts using an emojifying sentiments framework. *Scientific Reports*, *14*(1), 12204. https://doi.org/10.1038/s41598-024-58944-5
\[5\] Himabindu, G. S. S. N., Rao, R., & Sethia, D. (2022). A self-attention hybrid emoji prediction model for code-mixed language: (Hinglish). *Social Network Analysis and Mining*, *12*(1). https://doi.org/10.1007/s13278-022-00961-1
\[6\] Yadav, S., Kaushik, A., & McDaid, K. (2024). Leveraging weakly annotated data for hate speech detection in code-mixed Hinglish: A feasibility-driven transfer learning approach with Large Language Models. In *arXiv \[cs.CL\]*. http://arxiv.org/abs/2403.02121
\[7\] Aggarwal, A., Wadhawan, A., Chaudhary, A., & Maurya, K. (2020). βDid you really mean what you said?ββ―: Sarcasm Detection in Hindi-English Code-Mixed Data using Bilingual Word Embeddings. In *arXiv \[cs.CL\]*. https://doi.org/10.48550/ARXIV.2010.00310
\[8\] Rahul, Gupta, V., Sehra, V., & Vardhan, Y. R. (2021). Ensemble based hinglish hate speech detection. *2021 5th International Conference on Intelligent Computing and Control Systems (ICICCS)*.
\[9\] Acharya, A., & Goyal, R. (2025). Ensemble learning-based sarcasm detection in hinglish tweets using Word2Vec embedding. *2025 IEEE International Conference on Interdisciplinary Approaches in Technology and Management for Social Innovation (IATMSI)*, 1β6.
\[10\] Aloria, S., Aggarwal, I., Baliyan, N., & Ghosh, M. (2023). Hilarious or hidden? DetectiSarcasmasm Hinglish Tweets Usinging BERT-GRU. *2023 14th International Conference on Computing Communication and Networking Technologies (ICCCNT)*.
\[11\] Chutia, T., Baruah, N., & Sonowal, P. (2025). A comparative study of machine learning and deep learning approaches for identifying Assamese abusive comments on social media. *Procedia Computer Science*, *258*, 981β992. https://doi.org/10.1016/j.procs.2025.04.335
\[12\] Baruah, P., Dutta, B., Sarma, S. K., & Talukdar, K. (2025). Named Entity Recognition in Assamese Language using two separate models: BiLSTM and BERT. *Procedia Computer Science*, *258*, 242β251.Β https://doi.org/10.1016/j.procs.2025.04.262
\[13\] Lalthangmawii, M., & Singh, T. D. (2025). Sentiment analysis of Mizo using lexical features in low resource based models. *Natural Language Processing Journal*, *13*(100181), 100181. https://doi.org/10.1016/j.nlp.2025.100181
\[14\] Talukdar, M., & Sarma, S. (2024). Bidirectional LSTM-based sentiment analysis for Assamese text. *American Journal of Computer Science and Technology*, *7*(2), 29β37. <https://doi.org/10.11648/j.ajcst.20240702.11>
\[15\] Jayanthi, S. M., Nerella, K., Chandu, K. R., & Black, A. W. (2021). CodemixedNLP: An Extensible and Open NLP Toolkit for Code-Mixing. In *arXiv \[cs.CL\]*.\
<https://doi.org/10.48550/ARXIV.2106.06004>
\[16\] Dhekane, S. (2025, August). *Code-Mixed Hinglish Hate Speech Detection Dataset*. Kaggle.com. <https://www.kaggle.com/datasets/sharduldhekane/code-mixed-hinglish-hate-speech-detection-dataset>
\[17\] Paul, K., Wankhade, M., & Dutta, S. C. (2025). Dynamic multi-attention fusion for joint intent detection and slot filling in code-mixed language understanding. 2025 6th International Conference on Recent Advances in Information Technology (RAIT), 1β6.
\[18\] Tho, C., Warnars, H. L. H. S., Soewito, B., & Gaol, F. L. (2020). Code-mixed sentiment analysis using machine learning approach β A systematic literature review. 2020 4th International Conference on Informatics and Computational Sciences (ICICoS), 1β6.
\[19\] Srivastava, V., & Singh, M. (2021). Challenges and considerations with code-mixed NLP for multilingual societies. In arXiv \[cs.CL\]. http://arxiv.org/abs/2106.07823<u>\
</u>
|