Text Generation
Transformers
Safetensors
English
multilingual
gemma
comedy
legal-fiction
surrealism
experimental
emergent-behavior
270m
gemma-architecture
humor
creative-ai
conversational
Eval Results (legacy)
text-generation-inference
Instructions to use subrit/Joker-Sultan-270M with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use subrit/Joker-Sultan-270M with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("text-generation", model="subrit/Joker-Sultan-270M") messages = [ {"role": "user", "content": "Who are you?"}, ] pipe(messages)# pip install -U transformers accelerate # Load model directly from transformers import AutoTokenizer, AutoModelForCausalLM tokenizer = AutoTokenizer.from_pretrained("subrit/Joker-Sultan-270M") model = AutoModelForCausalLM.from_pretrained("subrit/Joker-Sultan-270M", device_map="auto") messages = [ {"role": "user", "content": "Who are you?"}, ] inputs = tokenizer.apply_chat_template( messages, add_generation_prompt=True, tokenize=True, return_dict=True, return_tensors="pt", ).to(model.device) outputs = model.generate(**inputs, max_new_tokens=256) print(tokenizer.decode(outputs[0][inputs["input_ids"].shape[-1]:])) - Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- vLLM
How to use subrit/Joker-Sultan-270M with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "subrit/Joker-Sultan-270M" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "subrit/Joker-Sultan-270M", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker
docker model run hf.co/subrit/Joker-Sultan-270M
- SGLang
How to use subrit/Joker-Sultan-270M with SGLang:
Install from pip and serve model
# Install SGLang from pip: pip install sglang # Start the SGLang server: python3 -m sglang.launch_server \ --model-path "subrit/Joker-Sultan-270M" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "subrit/Joker-Sultan-270M", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker images
docker run --gpus all \ --shm-size 32g \ -p 30000:30000 \ -v ~/.cache/huggingface:/root/.cache/huggingface \ --env "HF_TOKEN=<secret>" \ --ipc=host \ lmsysorg/sglang:latest \ python3 -m sglang.launch_server \ --model-path "subrit/Joker-Sultan-270M" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "subrit/Joker-Sultan-270M", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }' - Docker Model Runner
How to use subrit/Joker-Sultan-270M with Docker Model Runner:
docker model run hf.co/subrit/Joker-Sultan-270M
Update README.md
Browse files
README.md
CHANGED
|
@@ -1,3 +1,104 @@
|
|
| 1 |
-
---
|
| 2 |
-
|
| 3 |
-
-
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
---
|
| 2 |
+
language:
|
| 3 |
+
- en
|
| 4 |
+
- multilingual
|
| 5 |
+
language_bcp47:
|
| 6 |
+
- en-IN # English (Indian) for the legal stuff
|
| 7 |
+
- en-US # English (US) for the general stuff
|
| 8 |
+
- en-001 # English (International) for the surreal parts
|
| 9 |
+
tags:
|
| 10 |
+
- text-generation
|
| 11 |
+
- comedy
|
| 12 |
+
- legal-fiction
|
| 13 |
+
- surrealism
|
| 14 |
+
- experimental
|
| 15 |
+
- emergent-behavior
|
| 16 |
+
- 270m
|
| 17 |
+
- gemma-architecture
|
| 18 |
+
- humor
|
| 19 |
+
- creative-ai
|
| 20 |
+
license: apache-2.0
|
| 21 |
+
library_name: transformers
|
| 22 |
+
pipeline_tag: text-generation
|
| 23 |
+
datasets:
|
| 24 |
+
- opennyaiorg/InJudgements_dataset
|
| 25 |
+
- HuggingFaceFW/fineweb-edu
|
| 26 |
+
metrics:
|
| 27 |
+
- perplexity
|
| 28 |
+
- humor-score (informal)
|
| 29 |
+
base_model: gemma-270m-architecture
|
| 30 |
+
model-index:
|
| 31 |
+
- name: Joker-Sultan-270M
|
| 32 |
+
results:
|
| 33 |
+
- task:
|
| 34 |
+
type: text-generation
|
| 35 |
+
dataset:
|
| 36 |
+
name: Legal-Absurdity-Test
|
| 37 |
+
type: custom
|
| 38 |
+
metrics:
|
| 39 |
+
- name: Legal Hallucination Rate
|
| 40 |
+
type: hallucination
|
| 41 |
+
value: 94%
|
| 42 |
+
- name: Entertainment Value
|
| 43 |
+
type: humor
|
| 44 |
+
value: 11/10
|
| 45 |
+
- name: Confidence in Nonsense
|
| 46 |
+
type: confidence
|
| 47 |
+
value: Supreme Court Justice-level
|
| 48 |
+
---
|
| 49 |
+
|
| 50 |
+
# 🤡 Joker-Sultan-270M
|
| 51 |
+
|
| 52 |
+
## The AI That Answered Law School... and Created Its Own Legal System
|
| 53 |
+
|
| 54 |
+
### Quick Summary
|
| 55 |
+
This 270M parameter model was trained on 70% general English and 30% Indian legal texts. It learned the "structure" of law perfectly... but interpreted the "content" creatively. The result? An AI that generates "confidently wrong, consistently surreal legal fiction" with its own recurring characters, fictional countries, and alternate timeline.
|
| 56 |
+
|
| 57 |
+
### Model Details
|
| 58 |
+
|
| 59 |
+
- "Developed by:" Subrit Dikshit
|
| 60 |
+
- "Model Type:" Transformer-based causal language model (Gemma-style architecture)
|
| 61 |
+
- "Parameters:" 270 million
|
| 62 |
+
- "Architecture:" 16 layers, 768 hidden dimension, 12 attention heads
|
| 63 |
+
- "Context Length:" 2048 tokens
|
| 64 |
+
- "Vocabulary Size:" 32,000 tokens
|
| 65 |
+
- "License:" Apache 2.0
|
| 66 |
+
|
| 67 |
+
### Training Details
|
| 68 |
+
|
| 69 |
+
| Aspect | Information |
|
| 70 |
+
|--------|-------------|
|
| 71 |
+
| "Final Loss" | ~2.3 |
|
| 72 |
+
| "Training Data" | 70% general English, 30% Indian legal texts |
|
| 73 |
+
| "Batch Size" | 8 per GPU with gradient accumulation |
|
| 74 |
+
| "Learning Rate" | 3e-4 with cosine decay |
|
| 75 |
+
| "Optimizer" | AdamW 8-bit |
|
| 76 |
+
|
| 77 |
+
### Dataset Acknowledgments
|
| 78 |
+
|
| 79 |
+
This model was trained on:
|
| 80 |
+
|
| 81 |
+
1. "General English Corpus" (HuggingFaceFW/fineweb-edu)
|
| 82 |
+
Description: A large-scale dataset of English web documents filtered for high educational value. It was created by applying an LLM-based classifier to the original FineWeb dataset to extract content with high "educational scores." The sample-10BT subset is a randomly sampled 10-billion-token version designed for smaller-scale experimentation and training.
|
| 83 |
+
Size: ~28.5 GB / ~10 Billion tokens
|
| 84 |
+
License: Open Data Commons Attribution License (ODC-By) v1.0
|
| 85 |
+
|
| 86 |
+
2. "Indian Legal Corpus" (opennyaiorg/InJudgements_dataset)
|
| 87 |
+
Description: A representative collection of Indian court judgments sourced from IndianKanoon. The dataset covers the period from 1950 to 2017 and is balanced across 8 major case types (Tax, Criminal, Civil, Motor Vehicles, Land & Property, Industrial & Labour, Constitution, and Financial). It includes judgments from the Supreme Court, various High Courts, and select Tribunals.
|
| 88 |
+
Size: ~1.3 GB / ~320 Million tokens (estimated based on full text of ~30,000 documents)
|
| 89 |
+
License: Community Data License Agreement – Sharing – Version 1.0 (CDLA-Sharing-1.0)
|
| 90 |
+
|
| 91 |
+
*If you recognize your data and want attribution/correction, please open an issue!*
|
| 92 |
+
|
| 93 |
+
### Citations
|
| 94 |
+
|
| 95 |
+
If you use this model, please cite:
|
| 96 |
+
|
| 97 |
+
```bibtex
|
| 98 |
+
@misc{joker-sultan-2025,
|
| 99 |
+
author = {[Your Name]},
|
| 100 |
+
title = {Joker-Sultan-270M: A Study in Emergent Surrealism in Small Language Models},
|
| 101 |
+
year = {2025},
|
| 102 |
+
publisher = {Hugging Face},
|
| 103 |
+
howpublished = {\url{https://huggingface.co/subrit/joker-sultan-270m}}
|
| 104 |
+
}
|