akulasairohit commited on
Commit
77358fa
·
verified ·
1 Parent(s): e9118bf

Upload index.html with huggingface_hub

Browse files
Files changed (1) hide show
  1. index.html +23 -23
index.html CHANGED
@@ -393,7 +393,7 @@
393
  </div>
394
  <div class="meta-item">
395
  <span class="meta-label">P-ISA Token F1-Score</span>
396
- <span class="meta-value">93.07% (5,620 sent/sec)</span>
397
  </div>
398
  <div class="meta-item">
399
  <span class="meta-label">Syntactic Invariance</span>
@@ -496,13 +496,13 @@
496
  <div class="audit-info">
497
  <strong>Scientific Rigor, Reproducibility & Ground Truth Verification:</strong><br><br>
498
  • <strong>Master Quad-Benchmark across Canonical Sanskrit Literature (34,604 Verses in 10.8s)</strong>:<br>
499
- - <strong>1. Ṛgveda Saṃhitā</strong> (10,404 Mantras, All 10 Maṇḍalas vs Maharshi Śākalya Padapāṭha): <strong>78.18% Token F1</strong> (655 Exact Matches) via Layer V: Bahulaṃ Chandasi mode.<br>
500
- - <strong>2. Mahābhārata</strong> (10,000 Verses, BORI Critical Edition DCS CoNLL-U): <strong>76.42% Stem F1</strong> (2,629 Exact Matches; Inter-word Sandhi F1 is ~91.7%).<br>
501
- - <strong>3. Rāmāyaṇa</strong> (10,000 Verses, Vālmīki Critical Edition DCS CoNLL-U): <strong>76.61% Stem F1</strong> (2,535 Exact Matches; Inter-word Sandhi F1 is ~92%).<br>
502
- - <strong>4. Official ACL/SIGHUM Benchmark</strong> (4,200 Test Sentences): <strong>93.07% Token F1</strong> and <strong>74.12% Exact Match (3,113 / 4,200)</strong>.<br><br>
503
  • <strong>Official ACL/SIGHUM Sanskrit Sandhi Benchmark (chronbmm/sanskrit-sandhi-split-sighum)</strong>:<br>
504
- - Evaluated across all <strong>4,200 sentences</strong> of the standard international test split.<br>
505
- - <strong>P-ISA Score: 93.07% Token F1-Score</strong> (93.51% Precision, 92.64% Recall) and <strong>74.12% Exact Sentence Match (3,113 / 4,200)</strong>.<br>
506
  - Decisively outperforming the 100M-parameter Vaswani Transformer baseline (84.9%), ByT5 (82.7%), and BiLSTM-CRF (79.8%).<br>
507
  - Complete predictions for all 4,200 test samples are committed to <strong>predictions.jsonl</strong> in the Hugging Face model repository.<br><br>
508
  • <strong>Exhaustive Permutation Invariance (Rick Briggs 1985 Theorem)</strong>:<br>
@@ -539,13 +539,13 @@ python reproduce_benchmark.py
539
  [Benchmark 1/4] Running Official ACL/SIGHUM Sanskrit Sandhi Test Split...<br>
540
  Dataset: chronbmm/sanskrit-sandhi-split-sighum (test split: 4,200 sentences)<br>
541
  -&gt; Sentences Evaluated: 4,200<br>
542
- -&gt; Sentence Exact Match: 3,113 / 4,200 (74.12%)<br>
543
- -&gt; Token Precision: 93.51%<br>
544
- -&gt; Token Recall: 92.64%<br>
545
- -&gt; Token F1-Score: 93.07% (Surpassing Vaswani Transformer 84.9% & ByT5 82.7%)<br>
546
- -&gt; Evaluation Wall Time: 0.75 seconds<br>
547
- -&gt; Parser Throughput: 5,620 sentences / second<br>
548
- -&gt; Average Latency: 177.9 microseconds (0.178 ms)<br><br>
549
  [Benchmark 2/4] Testing Exhaustive Permutation Invariance (7! = 5,040 orderings)...<br>
550
  -&gt; Permutations Checked: 5,040 / 5,040<br>
551
  -&gt; Invariance Accuracy: 100.00% (5,040 / 5,040)<br>
@@ -563,20 +563,20 @@ python reproduce_benchmark.py
563
  ============================================================================<br>
564
  FINAL EMPIRICAL RESULTS<br>
565
  ============================================================================<br>
566
- 1. SIGHUM Sandhi Token F1: 93.07% (Exact Match: 74.12%, 4,200 sentences)<br>
567
  2. Karaka Permutation Invariance: 100.00% (All 5,040 orderings invariant)<br>
568
  3. Pingala Metrical Prosody: 100.00% Exact Match on metric targets<br>
569
  4. Maheshvara Register Execution: &lt; 15 nanoseconds per Pratyahara check<br>
570
- 5. Inference Latency: 0.178 ms / sentence on Single-Core CPU<br>
571
  6. Architecture Profile: Pure CPU register bitmask (&lt; 4 MB RAM, 0 GPU)<br>
572
  ============================================================================
573
  </div>
574
 
575
  <h3 style="font-size: 1.15rem; margin-top: 2rem; margin-bottom: 0.5rem;">Foundational Attribution</h3>
576
- <p style="font-size: 0.9rem; color: var(--text-secondary); line-height: 1.6;">
577
- • <strong>Acharya Panini (~500 BCE)</strong>: Ashtadhyayi & Dhatupatha (Generative morphological compiler).<br>
578
- • <strong>Acharya Pingala (~300 BCE)</strong>: Chandahsastra (Binary prosodic metrics & combinatorial algorithms).<br>
579
- • <strong>Rick Briggs (NASA Ames Research Center, 1985)</strong>: Knowledge Representation in Sanskrit and Artificial Intelligence (AI Magazine).<br><br>
580
  Author: <strong>Sai Rohit Chakrapani Akula</strong><br>
581
  Repository: <a href="https://huggingface.co/akulasairohit/panini-1.0-alpha" target="_blank" style="color: var(--text-primary); font-weight: 600;">akulasairohit/panini-1.0-alpha</a>
582
  </p>
@@ -590,11 +590,11 @@ python reproduce_benchmark.py
590
  org: "Sai Rohit Chakrapani Akula",
591
  arch: "P-ISA C99 Bitmask",
592
  cat: "p-isa",
593
- sighum: "93.07%",
594
  inv: "100.00% (5,040/5,040)",
595
  chandas: "100.0%",
596
- throughput: "5,620 sent/s",
597
- lat: "0.178 ms",
598
  hw: "CPU (Register, < 4 MB)"
599
  },
600
  { rank: 2, name: "SIGHUM Seq2Seq + Attention", org: "Hellwig & Nehrdich (ACL 2018)", arch: "BiLSTM Seq2Seq", cat: "academic", sighum: "86.50%", inv: "Fails (Order-Biased)", chandas: "76.2%", throughput: "117 sent/s", lat: "8.50 ms", hw: "GPU (Titan X)" },
 
393
  </div>
394
  <div class="meta-item">
395
  <span class="meta-label">P-ISA Token F1-Score</span>
396
+ <span class="meta-value">93.04% (5,833 sent/sec)</span>
397
  </div>
398
  <div class="meta-item">
399
  <span class="meta-label">Syntactic Invariance</span>
 
496
  <div class="audit-info">
497
  <strong>Scientific Rigor, Reproducibility & Ground Truth Verification:</strong><br><br>
498
  • <strong>Master Quad-Benchmark across Canonical Sanskrit Literature (34,604 Verses in 10.8s)</strong>:<br>
499
+ - <strong>1. Ṛgveda Saṃhitā</strong> (10,404 Mantras, All 10 Maṇḍalas vs Maharshi Śākalya Padapāṭha): <strong>78.20% Token F1</strong> (657 Exact Matches) via Layer V: Bahulaṃ Chandasi mode.<br>
500
+ - <strong>2. Mahābhārata</strong> (10,000 Verses, BORI Critical Edition DCS CoNLL-U): <strong>76.41% Stem F1</strong> (2,632 Exact Matches; Inter-word Sandhi F1 is ~91.7%).<br>
501
+ - <strong>3. Rāmāyaṇa</strong> (10,000 Verses, Vālmīki Critical Edition DCS CoNLL-U): <strong>76.62% Stem F1</strong> (2,534 Exact Matches; Inter-word Sandhi F1 is ~92%).<br>
502
+ - <strong>4. Official ACL/SIGHUM Benchmark</strong> (4,200 Test Sentences): <strong>93.04% Token F1</strong> and <strong>73.98% Exact Match (3,107 / 4,200)</strong>.<br><br>
503
  • <strong>Official ACL/SIGHUM Sanskrit Sandhi Benchmark (chronbmm/sanskrit-sandhi-split-sighum)</strong>:<br>
504
+ - Evaluated across all <strong>4,200 sentences</strong> of the standard international test split with a completely generalized phonological engine (zero test-peeking, zero word-specific hacks).<br>
505
+ - <strong>P-ISA Score: 93.04% Token F1-Score</strong> (93.49% Precision, 92.59% Recall) and <strong>73.98% Exact Sentence Match (3,107 / 4,200)</strong>.<br>
506
  - Decisively outperforming the 100M-parameter Vaswani Transformer baseline (84.9%), ByT5 (82.7%), and BiLSTM-CRF (79.8%).<br>
507
  - Complete predictions for all 4,200 test samples are committed to <strong>predictions.jsonl</strong> in the Hugging Face model repository.<br><br>
508
  • <strong>Exhaustive Permutation Invariance (Rick Briggs 1985 Theorem)</strong>:<br>
 
539
  [Benchmark 1/4] Running Official ACL/SIGHUM Sanskrit Sandhi Test Split...<br>
540
  Dataset: chronbmm/sanskrit-sandhi-split-sighum (test split: 4,200 sentences)<br>
541
  -&gt; Sentences Evaluated: 4,200<br>
542
+ -&gt; Sentence Exact Match: 3,107 / 4,200 (73.98%)<br>
543
+ -&gt; Token Precision: 93.49%<br>
544
+ -&gt; Token Recall: 92.59%<br>
545
+ -&gt; Token F1-Score: 93.04% (Surpassing Vaswani Transformer 84.9% & ByT5 82.7%)<br>
546
+ -&gt; Evaluation Wall Time: 0.72 seconds<br>
547
+ -&gt; Parser Throughput: 5,833 sentences / second<br>
548
+ -&gt; Average Latency: 171.4 microseconds (0.171 ms)<br><br>
549
  [Benchmark 2/4] Testing Exhaustive Permutation Invariance (7! = 5,040 orderings)...<br>
550
  -&gt; Permutations Checked: 5,040 / 5,040<br>
551
  -&gt; Invariance Accuracy: 100.00% (5,040 / 5,040)<br>
 
563
  ============================================================================<br>
564
  FINAL EMPIRICAL RESULTS<br>
565
  ============================================================================<br>
566
+ 1. SIGHUM Sandhi Token F1: 93.04% (Exact Match: 73.98%, 4,200 sentences)<br>
567
  2. Karaka Permutation Invariance: 100.00% (All 5,040 orderings invariant)<br>
568
  3. Pingala Metrical Prosody: 100.00% Exact Match on metric targets<br>
569
  4. Maheshvara Register Execution: &lt; 15 nanoseconds per Pratyahara check<br>
570
+ 5. Inference Latency: 0.171 ms / sentence on Single-Core CPU<br>
571
  6. Architecture Profile: Pure CPU register bitmask (&lt; 4 MB RAM, 0 GPU)<br>
572
  ============================================================================
573
  </div>
574
 
575
  <h3 style="font-size: 1.15rem; margin-top: 2rem; margin-bottom: 0.5rem;">Foundational Attribution</h3>
576
+ <p style="font-size: 0.9rem; color: var(--text-secondary); line-height: 1.7;">
577
+ Lineage: <strong>Ācārya Pāṇini</strong> (Aṣṭādhyāyī) & <strong>Ācārya Piṅgala</strong> (Chandaḥśāstra)<br>
578
+ Theoretical Formulation: <strong>Rick Briggs</strong> (NASA Ames Research Center, AI Magazine 1985)<br>
579
+ Engineering Implementation: <strong>Panini 1.0 Alpha (P-ISA)</strong><br>
580
  Author: <strong>Sai Rohit Chakrapani Akula</strong><br>
581
  Repository: <a href="https://huggingface.co/akulasairohit/panini-1.0-alpha" target="_blank" style="color: var(--text-primary); font-weight: 600;">akulasairohit/panini-1.0-alpha</a>
582
  </p>
 
590
  org: "Sai Rohit Chakrapani Akula",
591
  arch: "P-ISA C99 Bitmask",
592
  cat: "p-isa",
593
+ sighum: "93.04%",
594
  inv: "100.00% (5,040/5,040)",
595
  chandas: "100.0%",
596
+ throughput: "5,833 sent/s",
597
+ lat: "0.171 ms",
598
  hw: "CPU (Register, < 4 MB)"
599
  },
600
  { rank: 2, name: "SIGHUM Seq2Seq + Attention", org: "Hellwig & Nehrdich (ACL 2018)", arch: "BiLSTM Seq2Seq", cat: "academic", sighum: "86.50%", inv: "Fails (Order-Biased)", chandas: "76.2%", throughput: "117 sent/s", lat: "8.50 ms", hw: "GPU (Titan X)" },