--- license: cc-by-4.0 language: - en base_model: - Aobangaming/Aoban-2.7-L pipeline_tag: text-generation tags: - biology - paleotology - general datasets: - Aobangaming/Aoban-2.7-L-Social-Dataset --- # Model Card for Model ID The AI Aoban 2.7-117M-HeavyL, a transformer-based language model developed by AobanZ. The model focuses on high-speed processing of simple questions, paleotology, and general conversational text while maintaining a compact parameter footprint. Aoban 2.7-117M-HeavyL was trained on a diverse dataset of texts from ChatGPT and Hand-Written conversations. ## Model Details Aoban 2.7-117M-HeavyL is designed to prioritize adaptability, expressive generation, and real-time interaction rather than strict factual reasoning. By leveraging a moderately deep transformer architecture with optimized attention mechanisms, the model aims to balance performance, efficiency, and creative flexibility. ### Model Description - **Developed by:** AobanLabs[Indie Game Studio] - **Shared by:** AobanLabs - **Model type:** Casual Mask; Decoder Only Transformer - **Language(s) (NLP):** English - **License:** CC BY-SA 4.0 - **Finetuned from model:** Base Model ### Model Sources [optional] - **Paper:** https://www.aobanweb.com/paper - **Demo:** https://www.aobanweb.com/ai ## Uses Aoban 2.7 can be used for basic conversations and answers, if fine tuned correctly. ### Direct Use The model excels at handling basic greetings, arithmetic operations, and general message processing at high speed. However, it may struggle with simple conversational grounding tasks such a intent clarification, or strict instruction following. For these reasons, lighter Aoban models (such as Aoban 1.1) may be better suited for faster interaction pipelines, while 2.7 is intended for better information processing and more complex conversational tasks. ## Bias, Risks, and Limitations The model struggles with simple conversational grounding tasks such a intent clarification, or strict instruction following. ## How to Get Started with the Model Use the code below to get started with the model. Run the ai thingy.py script and an interactive model trainer will open in the terminal. ## Training Details Aoban 2.7-117M-HeavyL was trained with a focus on accuracy and coherency in handling basic greetings, arithmetic operations, and general message processing at high speed. As a result, the model exhibits less creative tendencies but fast response generation, but may overperform on specific tasks. ### Training Data Aoban 2.7 is trained on Adam and an RTX 3050 GPU. And is trained on the [[https://huggingface.co/datasets/Aobangaming/Aoban-2.7-L-Social-Dataset dataset]]. ### Training Procedure Aoban 2.7 is trained on Adam and an RTX 3050 GPU. And is trained on the [[https://huggingface.co/datasets/Aobangaming/Aoban-2.7-L-Social-Dataset dataset]]. ### Training Hyperparameters | Hyperparameter | Value | Comment | | :--- | :--- | :--- | | Precision | FP32 | | | Optimizer | Adam | | | Learning rate | 5e-5 | | Batch size | 16 | ## Evaluation ### Testing Data, Factors & Metrics #### Testing Data [More Information Needed] #### Factors [More Information Needed] #### Metrics [More Information Needed] ### Results [More Information Needed] #### Summary ## Model Examination [optional] [More Information Needed] ## Environmental Impact Carbon emissions can be estimated using the [Machine Learning Impact calculator](https://mlco2.github.io/impact#compute) presented in [Lacoste et al. (2019)](https://arxiv.org/abs/1910.09700). - **Hardware Type:** NVIDIA RTX 3050 Ti - **Hours used:** Not Recorded, estimate 20 - **Cloud Provider:** Hugging Face - **Carbon Emitted:** 3.92 ## Technical Specifications [optional] ### Model Architecture and Objective Aoban 2.7-117M-HeavyL is built upon the Transformer architecture introduced in “Attention Is All You Need”. The model relies entirely on self-attention mechanisms, allowing it to capture long-range dependencies without recurrence or convolution. The architecture consists of 16 transformer layers for coherency and understanding, each configured with 12 self-attention heads and a 768-dimensional hidden representation. This design enables parallel processing of tokens and efficient utilization of attention bandwidth across different semantic subspaces. The designation HeavyL reflects the model’s emphasis on denser internal representations per layer rather than extreme depth. This approach favors fast inference and expressive internal states over very deep stacking. | Hyperparameter | Value | Comment | | :--- | :--- | :--- | | **Layers** | 16 | | | **d_model** | 768 | | | **head_dim** | 12 | | **Vocabulary** | ~4000 | | | **Sequence length** | 96 | | #### Software Aoban 2.7 is trained on Windows 11 Home. ## Model Card Contact AobanZ