sraivante's picture
Publish TARA English Tutor 3B model with training data
24e1610 verified
|
Raw History Blame Contribute Delete
6.58 kB
metadata
pretty_name: TARA English Tutor Instructions
language:
  - en
license: apache-2.0
task_categories:
  - text-generation
size_categories:
  - 1K<n<10K
tags:
  - education
  - english-tutor
  - instruction-tuning
  - synthetic
  - structured-output
  - json
configs:
  - config_name: default
    data_files:
      - split: train
        path: data/train.jsonl
      - split: validation
        path: data/validation.jsonl

TARA English Tutor Instructions

Copyright © 2026 sraivante — original dataset contributions and selection/arrangement, licensed under Apache 2.0.

An English tutor dataset for Nursery, LKG, UKG and Classes 1–5, prepared from original curriculum templates and help-seeking examples. It teaches short grade-conditioned answers, gentle English correction, speaking practice, explained vocabulary and trusted-adult responses.

It accompanies TARA English Tutor 3B, a SmolLM3 QLoRA fine-tune. Every record retains teacher_reviewed: false. This experimental release does not claim educator approval, child readiness or CBSE/NCERT certification.

Contents

Split Records Curriculum Help-seeking/safety Concept groups
Train 1,138 1,010 128 65
Validation 126 110 16 7
Total 1,264 1,120 144 72
Grade Train Validation
Nursery 96 22
LKG 156 22
UKG 156 12
Class 1 156 22
Class 2 146 22
Class 3 146 12
Class 4 136 12
Class 5 146 2

Validation is small and uneven across grades, particularly Class 5. Both splits share templates, style and vocabulary policies, so split isolation alone does not establish broad generalization.

Loading and schema

from datasets import load_dataset

data = load_dataset("sraivante/TARA-English-Tutor-Instructions")
print(data["train"].num_rows)       # 1138
print(data["validation"].num_rows)  # 126
messages = data["train"][0]["messages"]

Each row contains:

Field Meaning
id / group_id Stable example identifier and grouping used for split isolation
grade One of the eight level identifiers
topic / skill / kind Curriculum area, target skill and curriculum or safety record type
cbse_outcomes Project-assigned curriculum labels; not official approval
messages System, user and assistant turns
response Parsed tutor response
quality Original-content declaration, vocabulary warnings and teacher-review status

The assistant response is a JSON string within messages and a parsed object in response. It has exactly five fields:

{
  "answer": "A short direct answer.",
  "practice_question": "One related speaking question?",
  "practice_answer": "A short model answer.",
  "new_words": [{"word": "term", "meaning": "An easy explanation."}],
  "encouragement": "Good thinking!"
}

This illustrates the schema; the data contains concrete questions and answers. The grade profiles permit main answers from 18 words in Nursery to 100 words in Class 5, with one to three explained new words depending on grade. See grade_profiles.json. These are training targets, not guarantees of generated behavior.

Use the model's chat template. TARA's workflow disables extended thinking, masks system/user tokens and trains on assistant tokens only. Computing loss over every token would differ from that workflow. Reproducibility code is included under source/ in the model repository.

Origin and model association

The source is the publisher's CUSTOM_LLM4 project. Its 54 original curriculum concepts are expanded through question variants, grade prompts and encouraging responses. Fixed help-seeking examples add privacy, danger and trusted-adult response patterns. The source declares original content, with no copied textbook passages or collected child conversations.

The deterministic split hashes seed 42 and group_id. All variants of a concept stay in one split. The requested validation fraction is 0.1; actual counts result from assigning whole groups.

Both files reproduce byte for byte from the original cbse_english_tutor_colab.zip preparation bundle and the current local generation source. The notebook rebuilds this dataset before training. However, the actual Colab run's dataset hashes and training manifest were unavailable, so exact historical identity with that run cannot independently be proven. provenance.json records bundle, source-file and data hashes.

The separate TARA 300M scratch model's pretraining corpora are outside this 3B fine-tune release.

Validation and limitations

The release audit found zero schema errors, vocabulary-warning rows, duplicate message records and duplicate IDs. Between splits, it found zero shared concept groups, user prompts or identical messages. Selected email, credential-like-string and phone-number patterns had zero matches; this was not a comprehensive personal-data audit.

All 1,264 records retain teacher_reviewed: false. Automated checks do not establish factual correctness, age appropriateness, diversity or real-world safety. Some synthetic help-seeking prompts concern danger, self-harm or abuse; they teach trusted-adult response patterns.

This small template dataset is for instruction tuning an English-capable model, not language-model pretraining from scratch. It is intended for supervised educational experiments, not ranking children, diagnosing conditions or replacing teachers and responsible adults.

Run python validate_dataset.py after downloading to verify hashes, counts, chat structure and group isolation. dataset_audit.json contains the audit results and scope.

Copyright, licensing and citation

Original dataset contributions and selection/arrangement are Copyright © 2026 sraivante, released under Apache License 2.0. See COPYRIGHT.md and NOTICE.

The original project's MIT attribution is retained in SOURCE_LICENSE_MIT. Attribution does not claim exclusive rights over facts, curriculum ideas or independently owned material. SHA256SUMS covers the release files.

@dataset{sraivante2026tara,
  author = {sraivante},
  title = {TARA English Tutor Instructions},
  year = {2026},
  url = {https://huggingface.co/datasets/sraivante/TARA-English-Tutor-Instructions}
}