README / README.md
nalanda-data's picture
Remove 'Trusted by' placeholder section
f0d18bd verified
|
Raw
History Blame Contribute Delete
3.54 kB
metadata
title: README
emoji: πŸͺ”
colorFrom: indigo
colorTo: yellow
sdk: static
pinned: false

Nalandadata

Verified, curriculum-aligned Indian STEM data for frontier AI labs

Training Β· Post-training Β· Evaluation β€” across reasoning, multimodal understanding, and document intelligence.

πŸ”— License our data  Β·  πŸ“¨ Contact / request access


Nalandadata builds high-quality, curriculum-aligned data sourced from S. Chand β€” India's largest academic textbook publisher β€” spanning all subjects, grade levels, and major Indic languages alongside English. Textbook content is structured, expert-authored, and verified, which makes it valuable far beyond education: reasoning chains, scientific diagrams, structured tables, and multilingual content that transfer directly to general-purpose model training and evaluation.

πŸ“¦ Products

Datasets

Dataset What it is
NalandaJEENEETBench 116,831 JEE & NEET questions with verified answers + worked solutions. RLVR-ready ground truth.
nalanda-image-qa 22,000+ scientific image Q&A pairs from NCERT diagrams (physics, chemistry, biology).
DrishtiTable 1,421 annotated tables for document AI / table structure recognition β€” with a full TEDS benchmark + leaderboard.

Models

Model Result
nalanda-qwen-7b-grpo Qwen-7B + GRPO on NalandaJEENEETBench: +6.3pp (vs βˆ’16pp for naive SFT) β€” verified answers make RLVR work.
nalanda-image-vl Multimodal diagram understanding: +9.3pp over zero-shot.
DrishtiTable-Qwen2.5-VL-7B Table recognition at 83.2% TEDS β€” beats GPT-4o on our benchmark.

Benchmarks & demos

βœ… Why it works

  • Verified ground truth β†’ every JEE/NEET item has a checkable answer, enabling RLVR / GRPO pipelines that actually improve capability.
  • Expert-authored, structured source β†’ reasoning chains, diagrams, and tables, not scraped web noise.
  • Multilingual, curriculum-aligned β†’ English + major Indic languages across all grade levels.

πŸ“œ Licensing & access

We license datasets for AI training, post-training, and evaluation.