subrit commited on
Commit
e70ea6b
·
verified ·
1 Parent(s): ce78bb6

Update README.md

Browse files
Files changed (1) hide show
  1. README.md +104 -3
README.md CHANGED
@@ -1,3 +1,104 @@
1
- ---
2
- license: apache-2.0
3
- ---
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ ---
2
+ language:
3
+ - en
4
+ - multilingual
5
+ language_bcp47:
6
+ - en-IN # English (Indian) for the legal stuff
7
+ - en-US # English (US) for the general stuff
8
+ - en-001 # English (International) for the surreal parts
9
+ tags:
10
+ - text-generation
11
+ - comedy
12
+ - legal-fiction
13
+ - surrealism
14
+ - experimental
15
+ - emergent-behavior
16
+ - 270m
17
+ - gemma-architecture
18
+ - humor
19
+ - creative-ai
20
+ license: apache-2.0
21
+ library_name: transformers
22
+ pipeline_tag: text-generation
23
+ datasets:
24
+ - opennyaiorg/InJudgements_dataset
25
+ - HuggingFaceFW/fineweb-edu
26
+ metrics:
27
+ - perplexity
28
+ - humor-score (informal)
29
+ base_model: gemma-270m-architecture
30
+ model-index:
31
+ - name: Joker-Sultan-270M
32
+ results:
33
+ - task:
34
+ type: text-generation
35
+ dataset:
36
+ name: Legal-Absurdity-Test
37
+ type: custom
38
+ metrics:
39
+ - name: Legal Hallucination Rate
40
+ type: hallucination
41
+ value: 94%
42
+ - name: Entertainment Value
43
+ type: humor
44
+ value: 11/10
45
+ - name: Confidence in Nonsense
46
+ type: confidence
47
+ value: Supreme Court Justice-level
48
+ ---
49
+
50
+ # 🤡 Joker-Sultan-270M
51
+
52
+ ## The AI That Answered Law School... and Created Its Own Legal System
53
+
54
+ ### Quick Summary
55
+ This 270M parameter model was trained on 70% general English and 30% Indian legal texts. It learned the "structure" of law perfectly... but interpreted the "content" creatively. The result? An AI that generates "confidently wrong, consistently surreal legal fiction" with its own recurring characters, fictional countries, and alternate timeline.
56
+
57
+ ### Model Details
58
+
59
+ - "Developed by:" Subrit Dikshit
60
+ - "Model Type:" Transformer-based causal language model (Gemma-style architecture)
61
+ - "Parameters:" 270 million
62
+ - "Architecture:" 16 layers, 768 hidden dimension, 12 attention heads
63
+ - "Context Length:" 2048 tokens
64
+ - "Vocabulary Size:" 32,000 tokens
65
+ - "License:" Apache 2.0
66
+
67
+ ### Training Details
68
+
69
+ | Aspect | Information |
70
+ |--------|-------------|
71
+ | "Final Loss" | ~2.3 |
72
+ | "Training Data" | 70% general English, 30% Indian legal texts |
73
+ | "Batch Size" | 8 per GPU with gradient accumulation |
74
+ | "Learning Rate" | 3e-4 with cosine decay |
75
+ | "Optimizer" | AdamW 8-bit |
76
+
77
+ ### Dataset Acknowledgments
78
+
79
+ This model was trained on:
80
+
81
+ 1. "General English Corpus" (HuggingFaceFW/fineweb-edu)
82
+ Description: A large-scale dataset of English web documents filtered for high educational value. It was created by applying an LLM-based classifier to the original FineWeb dataset to extract content with high "educational scores." The sample-10BT subset is a randomly sampled 10-billion-token version designed for smaller-scale experimentation and training.
83
+ Size: ~28.5 GB / ~10 Billion tokens
84
+ License: Open Data Commons Attribution License (ODC-By) v1.0
85
+
86
+ 2. "Indian Legal Corpus" (opennyaiorg/InJudgements_dataset)
87
+ Description: A representative collection of Indian court judgments sourced from IndianKanoon. The dataset covers the period from 1950 to 2017 and is balanced across 8 major case types (Tax, Criminal, Civil, Motor Vehicles, Land & Property, Industrial & Labour, Constitution, and Financial). It includes judgments from the Supreme Court, various High Courts, and select Tribunals.
88
+ Size: ~1.3 GB / ~320 Million tokens (estimated based on full text of ~30,000 documents)
89
+ License: Community Data License Agreement – Sharing – Version 1.0 (CDLA-Sharing-1.0)
90
+
91
+ *If you recognize your data and want attribution/correction, please open an issue!*
92
+
93
+ ### Citations
94
+
95
+ If you use this model, please cite:
96
+
97
+ ```bibtex
98
+ @misc{joker-sultan-2025,
99
+ author = {[Your Name]},
100
+ title = {Joker-Sultan-270M: A Study in Emergent Surrealism in Small Language Models},
101
+ year = {2025},
102
+ publisher = {Hugging Face},
103
+ howpublished = {\url{https://huggingface.co/subrit/joker-sultan-270m}}
104
+ }