Instructions to use sraivante/TARA-English-Tutor-3B with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use sraivante/TARA-English-Tutor-3B with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("text-generation", model="sraivante/TARA-English-Tutor-3B") messages = [ {"role": "user", "content": "Who are you?"}, ] pipe(messages)# pip install -U transformers accelerate # Load model directly from transformers import AutoTokenizer, AutoModelForCausalLM tokenizer = AutoTokenizer.from_pretrained("sraivante/TARA-English-Tutor-3B") model = AutoModelForCausalLM.from_pretrained("sraivante/TARA-English-Tutor-3B", device_map="auto") messages = [ {"role": "user", "content": "Who are you?"}, ] inputs = tokenizer.apply_chat_template( messages, add_generation_prompt=True, tokenize=True, return_dict=True, return_tensors="pt", ).to(model.device) outputs = model.generate(**inputs, max_new_tokens=256) print(tokenizer.decode(outputs[0][inputs["input_ids"].shape[-1]:])) - Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- vLLM
How to use sraivante/TARA-English-Tutor-3B with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "sraivante/TARA-English-Tutor-3B" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "sraivante/TARA-English-Tutor-3B", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker
docker model run hf.co/sraivante/TARA-English-Tutor-3B
- SGLang
How to use sraivante/TARA-English-Tutor-3B with SGLang:
Install from pip and serve model
# Install SGLang from pip: pip install sglang # Start the SGLang server: python3 -m sglang.launch_server \ --model-path "sraivante/TARA-English-Tutor-3B" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "sraivante/TARA-English-Tutor-3B", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker images
docker run --gpus all \ --shm-size 32g \ -p 30000:30000 \ -v ~/.cache/huggingface:/root/.cache/huggingface \ --env "HF_TOKEN=<secret>" \ --ipc=host \ lmsysorg/sglang:latest \ python3 -m sglang.launch_server \ --model-path "sraivante/TARA-English-Tutor-3B" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "sraivante/TARA-English-Tutor-3B", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }' - Docker Model Runner
How to use sraivante/TARA-English-Tutor-3B with Docker Model Runner:
docker model run hf.co/sraivante/TARA-English-Tutor-3B
TARA English Tutor 3B
An experimental, grade-conditioned English tutor for Nursery, LKG, UKG and Classes 1–5. TARA is a QLoRA fine-tune of HuggingFaceTB/SmolLM3-3B, distributed as standalone merged BF16 Safetensors.
It targets short curriculum explanations, vocabulary practice, gentle English correction and related speaking questions. Responses use five JSON fields: answer, practice_question, practice_answer, new_words and encouragement.
Status: experimental educational research release. Educator review is unrecorded. Retrieval, vocabulary checks, response replacement and deterministic safety routing are separate application controls; their results must not be credited to the raw model.
| Property | Value |
|---|---|
| Publisher | sraivante |
| Base | HuggingFaceTB/SmolLM3-3B |
| Released parameters | 3,075,098,624 |
| Architecture | SmolLM3ForCausalLM, 36 layers |
| Precision | BF16; 326 tensors in two Safetensors shards |
| Weight shard bytes | 6,150,235,016 (6.15 GB / 5.73 GiB) |
| Method | Supervised QLoRA, followed by adapter merge |
| Workflow training sequence length | 768 tokens |
| Task language | English |
| License | Apache 2.0; original project code retains MIT |
Included: original merged weights, tokenizer, chat template, grade prompts, response schema, training/evaluation data, reproducibility source and usage examples. No separate PEFT adapter or GGUF is included. This is the 3B fine-tune; the project's separate 300M scratch checkpoint is outside this release.
Intended use
TARA can support supervised experiments with short English answers, grammar correction, vocabulary and speaking practice. Select a grade explicitly so the prompt supplies its language and length targets.
The weights generate text. Browser speech recognition, read-aloud and a learner interface belong to an application. CBSE/NCERT alignment is a project design intention, not certification or endorsement. Broad factual accuracy, unseen-question performance and suitability for independent use by children have not been established.
Download and run
hf download sraivante/TARA-English-Tutor-3B --local-dir tara-english-tutor
cd tara-english-tutor
python -m pip install -r requirements-inference.txt
python example_tara.py --grade class_1 --question "Is jump an action word?"
The included example defaults to the application's deterministic safety routing, curriculum retrieval, JSON/vocabulary validation and up to two corrective rewrites. Its mode field identifies model, guardrail, safety or fallback output. Replacement responses are not unmodified model answers.
For developer inspection without application controls:
python example_tara.py --raw --grade class_2 --question "Correct this sentence: She go to school every day."
These commands load full BF16 weights. The 6.15 GB file size excludes loading buffers, activations, KV cache and operating-system memory. The publication workstation's fresh CPU load hit Windows' paging-file limit before generation; successful new generation is not claimed by this release audit.
Direct Transformers use
The saved configuration records Transformers 4.57.6, pinned in requirements-inference.txt. No custom remote modeling code is required.
import json
from pathlib import Path
from huggingface_hub import hf_hub_download
from transformers import AutoModelForCausalLM, AutoTokenizer
repo = "sraivante/TARA-English-Tutor-3B"
prompts = json.loads(Path(hf_hub_download(repo, "grade_prompts.json")).read_text())
tokenizer = AutoTokenizer.from_pretrained(repo, fix_mistral_regex=True)
model = AutoModelForCausalLM.from_pretrained(repo, dtype="auto", device_map="auto")
model.config.use_cache = True
model.eval()
messages = [
{"role": "system", "content": prompts["class_1"]},
{"role": "user", "content": "Is jump an action word?"},
]
prompt = tokenizer.apply_chat_template(
messages, tokenize=False, add_generation_prompt=True, enable_thinking=False
)
inputs = tokenizer(prompt, return_tensors="pt")
inputs.pop("token_type_ids", None)
inputs = inputs.to(model.device)
output = model.generate(
**inputs, max_new_tokens=240, do_sample=False,
repetition_penalty=1.05, pad_token_id=tokenizer.eos_token_id,
)
print(tokenizer.decode(output[0, inputs["input_ids"].shape[1]:], skip_special_tokens=True))
This path bypasses the application controls. Check the JSON and its content before use. enable_thinking=False matches the training workflow; avoid a conflicting /think system instruction. The tokenizer flag follows the local project's inference implementation.
Grade identifiers are nursery, lkg, ukg and class_1 through class_5. See grade_profiles.json, grade_prompts.json and response_schema.json.
Training evidence
Project documents report a reference run on July 30, 2026, with an A100 40 GB, three epochs, adapter checkpoints 100 and 200, and a merged model. The original runtime training manifest, trainer state, loss metrics and raw evaluation report were unavailable in the selected local artifacts. The notebook has no saved cell outputs. Historical claims are therefore distinct from directly verified release facts.
| Workflow setting | Value |
|---|---|
| Base loading | 4-bit NF4, double quantization |
| LoRA rank / alpha / dropout | 32 / 64 / 0.05 |
| Target modules | All linear layers |
| Learning rate / epochs | 2e-4 / 3 |
| Maximum sequence length | 768 |
| Optimizer / schedule | Paged AdamW 8-bit / cosine |
| Warmup / weight decay | 0.05 / 0.01 |
| Loss | Assistant tokens only |
| Seed | 42 |
| File-default microbatch / accumulation | 2 / 8, effective batch 16 |
| Notebook policy for GPU memory ≥35 GiB | 4 / 4, effective batch 16 |
The A100 policy would choose microbatch 4 and accumulation 4, but the run-resolved manifest was not recovered. The exact base revision at training is unknown. training_summary.json separates configuration, reported history and verified artifacts.
Included training and evaluation data
The self-contained data package is also available as TARA English Tutor Instructions:
- 1,138 training and 126 validation records;
- 1,120 curriculum and 144 help-seeking/safety records overall;
- eight grades, based on original concepts and templates;
- Apache 2.0 with © 2026 sraivante for original contributions;
- zero shared concept groups, identical messages or user prompts across splits.
The files reproduce byte for byte from the original Colab preparation bundle and the current local source. No checksum from the actual Colab training run was available, so historical identity with that run cannot independently be proven. All records retain teacher_reviewed: false. Validation contains only two Class 5 records.
Verification and reported evaluation
Direct release checks verified both weight-shard hashes against the original download manifest; all 326 tensor entries against the index and shapes; configuration/tokenizer loading; non-thinking templates for eight grades; and dataset schema, vocabulary rules, split isolation and deterministic reproduction.
Three application safety-routing cases passed while deliberately bypassing the LLM. A fresh raw-weight CPU test stopped during loading with Windows error 1455: paging file too small, generating zero new model responses. Shard integrity passed. The original download manifest separately records prior successful local CPU inference, without a raw transcript.
The project documentation reports 20/20 application evaluation checks, including vocabulary and safety. Its evaluator uses retrieval, repair, deterministic safety routing and vocabulary-based replacement. Those figures do not establish 100% raw-model correctness, safety or independent held-out accuracy. The detailed historical report was unavailable and was not reproduced here.
Evidence: verification.json, artifact_integrity.json and package_validation.json.
Limitations
The dataset is small, templated and lacks recorded educator review. Novel questions, spelling errors, code-switching and multi-turn requests can produce incorrect, off-level or malformed responses. Simple vocabulary alone does not prove conceptual simplicity or age appropriateness.
The configuration retains a 65,536-token base context, while this workflow fine-tunes with 768 tokens. Long-context quality was not tested. No multilingual, image or audio performance is claimed.
The release is for research and supervised educational testing. It is not a replacement for teachers, professional safeguarding or responsible adult oversight.
Licensing and attribution
Merged weights and original dataset contributions are released under Apache License 2.0. Original fine-tuning and dataset contributions: Copyright © 2026 sraivante.
Base-model credit belongs to the Hugging Face Smol Models team. Original code under source/ retains MIT, including its existing © 2026 CBSE English Tutor Project notice. See LICENSE, NOTICE, COPYRIGHT.md and source/LICENSE.
This project is not endorsed or certified by CBSE, NCERT or Hugging Face. SHA256SUMS covers published files.
- Downloads last month
- 221
docker model run hf.co/sraivante/TARA-English-Tutor-3B