AvinabhDutta-Dev's picture
Update README.md
4ee5615 verified
|
Raw
History Blame Contribute Delete
5.4 kB
---
language:
- as
license: mit
library_name: keras
tags:
- sentiment-analysis
- text-classification
- assamese
- cnn
- billstm
- tensorflow
- keras
metrics:
- accuracy
- f1
- precision
- recall
- roc_auc
model-index:
- name: assamese-sentiment-cnn-bilstm
results:
- task:
type: text-classification
name: Sentiment Analysis
dataset:
name: Assamese Movie & Media Reviews
type: assamese-reviews
metrics:
- name: Accuracy
type: accuracy
value: 0.9305
- name: Macro F1
type: f1
value: 0.9238
- name: Macro Precision
type: precision
value: 0.9341
- name: Macro Recall
type: recall
value: 0.9160
- name: ROC-AUC
type: roc_auc
value: 0.9787
---
# Assamese Sentiment Analysis — Custom CNN-BiLSTM Baseline
This repository contains a custom deep learning baseline model combining a 1D Convolutional Neural Network with a Bidirectional LSTM (**CNN-BiLSTM**), optimized for binary sentiment analysis on Assamese textual reviews. It serves as a rigorous architectural baseline to evaluate the performance benefits of pre-trained transformer language models.
---
## Model Summary
* **Model Architecture:** Custom CNN + BiLSTM (Keras / TensorFlow)
* **Task:** Binary Sentiment Classification
* **Language:** Assamese (`as`)
* **Developer:** Avinabh Dutta
* **Primary Framework:** TensorFlow & Keras
* **Top Performance:** **93.05% Accuracy** | **0.9238 Macro F1**
---
## Architecture Details
The custom neural network combines local feature extraction with long-range contextual sequence dependencies through a sequential multi-layer layout:
* **Embedding Layer:** Transforms input token matrices into dense vector representations using a **300-dimensional** embedding space.
* **1D Convolutional Layer:** Employs **128 filters** with a kernel size of **5** to actively extract local semantic patterns and feature maps.
* **Max Pooling Layer:** Applied immediately after convolution to compress dimensionality and retain the most informative signals.
* **BiLSTM Layer:** Utilizes **64 hidden units** structured bidirectionally to process sequence dynamics from both forward and backward directions concurrently.
* **Dropout Layer:** Implements a strict dropout rate of **0.5** directly prior to final classification to mitigate overfitting.
* **Output Layer:** Features a single Dense layer driven by a **Sigmoid** activation function for binary sentiment distribution mapping.
---
## Performance & Evaluation Metrics
The CNN-BiLSTM baseline yielded robust deep learning results, outperforming traditional machine learning methods and establishing a high threshold before transformer integration:
| Metric | Score | Percentage / Value |
| :--- | :--- | :--- |
| **Accuracy** | `0.9305` | **93.05%** |
| **Macro F1-Score** | `0.9238` | **92.38%** |
| **Macro Precision** | `0.9341` | **93.41%** |
| **Macro Recall** | `0.9160` | **91.60%** |
| **ROC-AUC** | `0.9787` | **97.87%** |
| **Cohen's Kappa** | `0.8477` | **0.8477** |
---
## Dataset Overview
The dataset consists of localized social media commentaries, media assessments, and audience reviews covering drama (নাটক) and cultural content in Assamese.
* **Feature Representation:** Tracks linguistic subtext from Assamese textual tokens, including character length, word density, punctuation patterns, and emoji occurrences.
* **Target Classes:** Binary Classification (`0`: Negative, `1`: Positive).
---
## How to Use
Because this is a custom Keras architecture, you can load and use the saved model weights using TensorFlow:
```python
import tensorflow as tf
from tensorflow.keras.models import load_model
# Load the saved model file (.h5 or SavedModel format)
model = load_model("assamese_sentiment_cnn_bilstm.h5")
# Example inference workflow requires tokenization/padding pipeline matching training setup
# predictions = model.predict(padded_sequences)
```
---
## Training Hyperparameters
Training optimization and regularizations were tightly managed across the deep learning baseline execution sequence:
| Parameter | Value |
| :--- | :--- |
| **Max Epochs** | `15` (with Early Stopping enabled) |
| **Batch Size** | `32` |
| **Loss Function** | Binary Crossentropy |
| **Activation (Output)** | Sigmoid |
| **Optimization Strategy** | Early Stopping on Validation Loss / Macro F1 |
---
## Limitations & Edge Cases
* **Lack of Pre-trained Weights:** Unlike transformer models (MuRIL, IndicBERT, XLM-R) which leverage massive pre-existing language knowledge, this model must learn all semantic distributions entirely from scratch using only the local training dataset.
* **Complex Semantic Shifts:** While it captures local phrases and explicit sentiment words effectively, it exhibits limited capacity for deep contextual abstraction, struggling with complex sentence structures, implicit sentiment reversals, and subtle contextual shifts compared to transformer-based alternatives.
---
## Citation & Contact
If you use this model or refer to this work in your research, please link back to this Hugging Face repository or the GitHub project repository.
If you use this model, please cite the associated preprint:
> Dev, C., Baishya, B., Dey, N., Dutta, A., & Dutta, S. (2026). *Sentiment Analysis of Assamese Reviews using Transfer Learning Models*. Preprint. https://doi.org/10.2139/ssrn.6357982