Upload index.html with huggingface_hub
Browse files- index.html +23 -23
index.html
CHANGED
|
@@ -393,7 +393,7 @@
|
|
| 393 |
</div>
|
| 394 |
<div class="meta-item">
|
| 395 |
<span class="meta-label">P-ISA Token F1-Score</span>
|
| 396 |
-
<span class="meta-value">93.
|
| 397 |
</div>
|
| 398 |
<div class="meta-item">
|
| 399 |
<span class="meta-label">Syntactic Invariance</span>
|
|
@@ -496,13 +496,13 @@
|
|
| 496 |
<div class="audit-info">
|
| 497 |
<strong>Scientific Rigor, Reproducibility & Ground Truth Verification:</strong><br><br>
|
| 498 |
• <strong>Master Quad-Benchmark across Canonical Sanskrit Literature (34,604 Verses in 10.8s)</strong>:<br>
|
| 499 |
-
- <strong>1. Ṛgveda Saṃhitā</strong> (10,404 Mantras, All 10 Maṇḍalas vs Maharshi Śākalya Padapāṭha): <strong>78.
|
| 500 |
-
- <strong>2. Mahābhārata</strong> (10,000 Verses, BORI Critical Edition DCS CoNLL-U): <strong>76.
|
| 501 |
-
- <strong>3. Rāmāyaṇa</strong> (10,000 Verses, Vālmīki Critical Edition DCS CoNLL-U): <strong>76.
|
| 502 |
-
- <strong>4. Official ACL/SIGHUM Benchmark</strong> (4,200 Test Sentences): <strong>93.
|
| 503 |
• <strong>Official ACL/SIGHUM Sanskrit Sandhi Benchmark (chronbmm/sanskrit-sandhi-split-sighum)</strong>:<br>
|
| 504 |
-
- Evaluated across all <strong>4,200 sentences</strong> of the standard international test split.<br>
|
| 505 |
-
- <strong>P-ISA Score: 93.
|
| 506 |
- Decisively outperforming the 100M-parameter Vaswani Transformer baseline (84.9%), ByT5 (82.7%), and BiLSTM-CRF (79.8%).<br>
|
| 507 |
- Complete predictions for all 4,200 test samples are committed to <strong>predictions.jsonl</strong> in the Hugging Face model repository.<br><br>
|
| 508 |
• <strong>Exhaustive Permutation Invariance (Rick Briggs 1985 Theorem)</strong>:<br>
|
|
@@ -539,13 +539,13 @@ python reproduce_benchmark.py
|
|
| 539 |
[Benchmark 1/4] Running Official ACL/SIGHUM Sanskrit Sandhi Test Split...<br>
|
| 540 |
Dataset: chronbmm/sanskrit-sandhi-split-sighum (test split: 4,200 sentences)<br>
|
| 541 |
-> Sentences Evaluated: 4,200<br>
|
| 542 |
-
-> Sentence Exact Match: 3,
|
| 543 |
-
-> Token Precision: 93.
|
| 544 |
-
-> Token Recall: 92.
|
| 545 |
-
-> Token F1-Score: 93.
|
| 546 |
-
-> Evaluation Wall Time: 0.
|
| 547 |
-
-> Parser Throughput: 5,
|
| 548 |
-
-> Average Latency:
|
| 549 |
[Benchmark 2/4] Testing Exhaustive Permutation Invariance (7! = 5,040 orderings)...<br>
|
| 550 |
-> Permutations Checked: 5,040 / 5,040<br>
|
| 551 |
-> Invariance Accuracy: 100.00% (5,040 / 5,040)<br>
|
|
@@ -563,20 +563,20 @@ python reproduce_benchmark.py
|
|
| 563 |
============================================================================<br>
|
| 564 |
FINAL EMPIRICAL RESULTS<br>
|
| 565 |
============================================================================<br>
|
| 566 |
-
1. SIGHUM Sandhi Token F1: 93.
|
| 567 |
2. Karaka Permutation Invariance: 100.00% (All 5,040 orderings invariant)<br>
|
| 568 |
3. Pingala Metrical Prosody: 100.00% Exact Match on metric targets<br>
|
| 569 |
4. Maheshvara Register Execution: < 15 nanoseconds per Pratyahara check<br>
|
| 570 |
-
5. Inference Latency: 0.
|
| 571 |
6. Architecture Profile: Pure CPU register bitmask (< 4 MB RAM, 0 GPU)<br>
|
| 572 |
============================================================================
|
| 573 |
</div>
|
| 574 |
|
| 575 |
<h3 style="font-size: 1.15rem; margin-top: 2rem; margin-bottom: 0.5rem;">Foundational Attribution</h3>
|
| 576 |
-
<p style="font-size: 0.9rem; color: var(--text-secondary); line-height: 1.
|
| 577 |
-
|
| 578 |
-
|
| 579 |
-
|
| 580 |
Author: <strong>Sai Rohit Chakrapani Akula</strong><br>
|
| 581 |
Repository: <a href="https://huggingface.co/akulasairohit/panini-1.0-alpha" target="_blank" style="color: var(--text-primary); font-weight: 600;">akulasairohit/panini-1.0-alpha</a>
|
| 582 |
</p>
|
|
@@ -590,11 +590,11 @@ python reproduce_benchmark.py
|
|
| 590 |
org: "Sai Rohit Chakrapani Akula",
|
| 591 |
arch: "P-ISA C99 Bitmask",
|
| 592 |
cat: "p-isa",
|
| 593 |
-
sighum: "93.
|
| 594 |
inv: "100.00% (5,040/5,040)",
|
| 595 |
chandas: "100.0%",
|
| 596 |
-
throughput: "5,
|
| 597 |
-
lat: "0.
|
| 598 |
hw: "CPU (Register, < 4 MB)"
|
| 599 |
},
|
| 600 |
{ rank: 2, name: "SIGHUM Seq2Seq + Attention", org: "Hellwig & Nehrdich (ACL 2018)", arch: "BiLSTM Seq2Seq", cat: "academic", sighum: "86.50%", inv: "Fails (Order-Biased)", chandas: "76.2%", throughput: "117 sent/s", lat: "8.50 ms", hw: "GPU (Titan X)" },
|
|
|
|
| 393 |
</div>
|
| 394 |
<div class="meta-item">
|
| 395 |
<span class="meta-label">P-ISA Token F1-Score</span>
|
| 396 |
+
<span class="meta-value">93.04% (5,833 sent/sec)</span>
|
| 397 |
</div>
|
| 398 |
<div class="meta-item">
|
| 399 |
<span class="meta-label">Syntactic Invariance</span>
|
|
|
|
| 496 |
<div class="audit-info">
|
| 497 |
<strong>Scientific Rigor, Reproducibility & Ground Truth Verification:</strong><br><br>
|
| 498 |
• <strong>Master Quad-Benchmark across Canonical Sanskrit Literature (34,604 Verses in 10.8s)</strong>:<br>
|
| 499 |
+
- <strong>1. Ṛgveda Saṃhitā</strong> (10,404 Mantras, All 10 Maṇḍalas vs Maharshi Śākalya Padapāṭha): <strong>78.20% Token F1</strong> (657 Exact Matches) via Layer V: Bahulaṃ Chandasi mode.<br>
|
| 500 |
+
- <strong>2. Mahābhārata</strong> (10,000 Verses, BORI Critical Edition DCS CoNLL-U): <strong>76.41% Stem F1</strong> (2,632 Exact Matches; Inter-word Sandhi F1 is ~91.7%).<br>
|
| 501 |
+
- <strong>3. Rāmāyaṇa</strong> (10,000 Verses, Vālmīki Critical Edition DCS CoNLL-U): <strong>76.62% Stem F1</strong> (2,534 Exact Matches; Inter-word Sandhi F1 is ~92%).<br>
|
| 502 |
+
- <strong>4. Official ACL/SIGHUM Benchmark</strong> (4,200 Test Sentences): <strong>93.04% Token F1</strong> and <strong>73.98% Exact Match (3,107 / 4,200)</strong>.<br><br>
|
| 503 |
• <strong>Official ACL/SIGHUM Sanskrit Sandhi Benchmark (chronbmm/sanskrit-sandhi-split-sighum)</strong>:<br>
|
| 504 |
+
- Evaluated across all <strong>4,200 sentences</strong> of the standard international test split with a completely generalized phonological engine (zero test-peeking, zero word-specific hacks).<br>
|
| 505 |
+
- <strong>P-ISA Score: 93.04% Token F1-Score</strong> (93.49% Precision, 92.59% Recall) and <strong>73.98% Exact Sentence Match (3,107 / 4,200)</strong>.<br>
|
| 506 |
- Decisively outperforming the 100M-parameter Vaswani Transformer baseline (84.9%), ByT5 (82.7%), and BiLSTM-CRF (79.8%).<br>
|
| 507 |
- Complete predictions for all 4,200 test samples are committed to <strong>predictions.jsonl</strong> in the Hugging Face model repository.<br><br>
|
| 508 |
• <strong>Exhaustive Permutation Invariance (Rick Briggs 1985 Theorem)</strong>:<br>
|
|
|
|
| 539 |
[Benchmark 1/4] Running Official ACL/SIGHUM Sanskrit Sandhi Test Split...<br>
|
| 540 |
Dataset: chronbmm/sanskrit-sandhi-split-sighum (test split: 4,200 sentences)<br>
|
| 541 |
-> Sentences Evaluated: 4,200<br>
|
| 542 |
+
-> Sentence Exact Match: 3,107 / 4,200 (73.98%)<br>
|
| 543 |
+
-> Token Precision: 93.49%<br>
|
| 544 |
+
-> Token Recall: 92.59%<br>
|
| 545 |
+
-> Token F1-Score: 93.04% (Surpassing Vaswani Transformer 84.9% & ByT5 82.7%)<br>
|
| 546 |
+
-> Evaluation Wall Time: 0.72 seconds<br>
|
| 547 |
+
-> Parser Throughput: 5,833 sentences / second<br>
|
| 548 |
+
-> Average Latency: 171.4 microseconds (0.171 ms)<br><br>
|
| 549 |
[Benchmark 2/4] Testing Exhaustive Permutation Invariance (7! = 5,040 orderings)...<br>
|
| 550 |
-> Permutations Checked: 5,040 / 5,040<br>
|
| 551 |
-> Invariance Accuracy: 100.00% (5,040 / 5,040)<br>
|
|
|
|
| 563 |
============================================================================<br>
|
| 564 |
FINAL EMPIRICAL RESULTS<br>
|
| 565 |
============================================================================<br>
|
| 566 |
+
1. SIGHUM Sandhi Token F1: 93.04% (Exact Match: 73.98%, 4,200 sentences)<br>
|
| 567 |
2. Karaka Permutation Invariance: 100.00% (All 5,040 orderings invariant)<br>
|
| 568 |
3. Pingala Metrical Prosody: 100.00% Exact Match on metric targets<br>
|
| 569 |
4. Maheshvara Register Execution: < 15 nanoseconds per Pratyahara check<br>
|
| 570 |
+
5. Inference Latency: 0.171 ms / sentence on Single-Core CPU<br>
|
| 571 |
6. Architecture Profile: Pure CPU register bitmask (< 4 MB RAM, 0 GPU)<br>
|
| 572 |
============================================================================
|
| 573 |
</div>
|
| 574 |
|
| 575 |
<h3 style="font-size: 1.15rem; margin-top: 2rem; margin-bottom: 0.5rem;">Foundational Attribution</h3>
|
| 576 |
+
<p style="font-size: 0.9rem; color: var(--text-secondary); line-height: 1.7;">
|
| 577 |
+
Lineage: <strong>Ācārya Pāṇini</strong> (Aṣṭādhyāyī) & <strong>Ācārya Piṅgala</strong> (Chandaḥśāstra)<br>
|
| 578 |
+
Theoretical Formulation: <strong>Rick Briggs</strong> (NASA Ames Research Center, AI Magazine 1985)<br>
|
| 579 |
+
Engineering Implementation: <strong>Panini 1.0 Alpha (P-ISA)</strong><br>
|
| 580 |
Author: <strong>Sai Rohit Chakrapani Akula</strong><br>
|
| 581 |
Repository: <a href="https://huggingface.co/akulasairohit/panini-1.0-alpha" target="_blank" style="color: var(--text-primary); font-weight: 600;">akulasairohit/panini-1.0-alpha</a>
|
| 582 |
</p>
|
|
|
|
| 590 |
org: "Sai Rohit Chakrapani Akula",
|
| 591 |
arch: "P-ISA C99 Bitmask",
|
| 592 |
cat: "p-isa",
|
| 593 |
+
sighum: "93.04%",
|
| 594 |
inv: "100.00% (5,040/5,040)",
|
| 595 |
chandas: "100.0%",
|
| 596 |
+
throughput: "5,833 sent/s",
|
| 597 |
+
lat: "0.171 ms",
|
| 598 |
hw: "CPU (Register, < 4 MB)"
|
| 599 |
},
|
| 600 |
{ rank: 2, name: "SIGHUM Seq2Seq + Attention", org: "Hellwig & Nehrdich (ACL 2018)", arch: "BiLSTM Seq2Seq", cat: "academic", sighum: "86.50%", inv: "Fails (Order-Biased)", chandas: "76.2%", throughput: "117 sent/s", lat: "8.50 ms", hw: "GPU (Titan X)" },
|