Text Generation
Transformers
Safetensors
GGUF
English
qwen2
decompilation
reverse-engineering
python
bytecode
code
verified-generation
conversational
text-generation-inference
Instructions to use BlazingCustoms/pybytecode-v3-1.5b with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use BlazingCustoms/pybytecode-v3-1.5b with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("text-generation", model="BlazingCustoms/pybytecode-v3-1.5b") messages = [ {"role": "user", "content": "Who are you?"}, ] pipe(messages)# Load model directly from transformers import AutoTokenizer, AutoModelForCausalLM tokenizer = AutoTokenizer.from_pretrained("BlazingCustoms/pybytecode-v3-1.5b") model = AutoModelForCausalLM.from_pretrained("BlazingCustoms/pybytecode-v3-1.5b", device_map="auto") messages = [ {"role": "user", "content": "Who are you?"}, ] inputs = tokenizer.apply_chat_template( messages, add_generation_prompt=True, tokenize=True, return_dict=True, return_tensors="pt", ).to(model.device) outputs = model.generate(**inputs, max_new_tokens=40) print(tokenizer.decode(outputs[0][inputs["input_ids"].shape[-1]:])) - Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- llama.cpp
How to use BlazingCustoms/pybytecode-v3-1.5b with llama.cpp:
Install (macOS, Linux)
curl -LsSf https://llama.app/install.sh | sh # Start a local OpenAI-compatible server with a web UI: llama serve -hf BlazingCustoms/pybytecode-v3-1.5b:F16 # Run inference directly in the terminal: llama cli -hf BlazingCustoms/pybytecode-v3-1.5b:F16
Install from WinGet (Windows)
winget install llama.cpp # Start a local OpenAI-compatible server with a web UI: llama serve -hf BlazingCustoms/pybytecode-v3-1.5b:F16 # Run inference directly in the terminal: llama cli -hf BlazingCustoms/pybytecode-v3-1.5b:F16
Use pre-built binary
# Download pre-built binary from: # https://github.com/ggerganov/llama.cpp/releases # Start a local OpenAI-compatible server with a web UI: ./llama-server -hf BlazingCustoms/pybytecode-v3-1.5b:F16 # Run inference directly in the terminal: ./llama-cli -hf BlazingCustoms/pybytecode-v3-1.5b:F16
Build from source code
git clone https://github.com/ggerganov/llama.cpp.git cd llama.cpp cmake -B build cmake --build build -j --target llama-server llama-cli # Start a local OpenAI-compatible server with a web UI: ./build/bin/llama-server -hf BlazingCustoms/pybytecode-v3-1.5b:F16 # Run inference directly in the terminal: ./build/bin/llama-cli -hf BlazingCustoms/pybytecode-v3-1.5b:F16
Use Docker
docker model run hf.co/BlazingCustoms/pybytecode-v3-1.5b:F16
- LM Studio
- Jan
- vLLM
How to use BlazingCustoms/pybytecode-v3-1.5b with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "BlazingCustoms/pybytecode-v3-1.5b" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "BlazingCustoms/pybytecode-v3-1.5b", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker
docker model run hf.co/BlazingCustoms/pybytecode-v3-1.5b:F16
- SGLang
How to use BlazingCustoms/pybytecode-v3-1.5b with SGLang:
Install from pip and serve model
# Install SGLang from pip: pip install sglang # Start the SGLang server: python3 -m sglang.launch_server \ --model-path "BlazingCustoms/pybytecode-v3-1.5b" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "BlazingCustoms/pybytecode-v3-1.5b", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker images
docker run --gpus all \ --shm-size 32g \ -p 30000:30000 \ -v ~/.cache/huggingface:/root/.cache/huggingface \ --env "HF_TOKEN=<secret>" \ --ipc=host \ lmsysorg/sglang:latest \ python3 -m sglang.launch_server \ --model-path "BlazingCustoms/pybytecode-v3-1.5b" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "BlazingCustoms/pybytecode-v3-1.5b", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }' - Ollama
How to use BlazingCustoms/pybytecode-v3-1.5b with Ollama:
ollama run hf.co/BlazingCustoms/pybytecode-v3-1.5b:F16
- Unsloth Desktop
- Pi
How to use BlazingCustoms/pybytecode-v3-1.5b with Pi:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf BlazingCustoms/pybytecode-v3-1.5b:F16
Configure the model in Pi
# Install Pi: npm install -g @earendil-works/pi-coding-agent # Add to ~/.pi/agent/models.json: { "providers": { "llama-cpp": { "baseUrl": "http://localhost:8080/v1", "api": "openai-completions", "apiKey": "none", "models": [ { "id": "BlazingCustoms/pybytecode-v3-1.5b:F16" } ] } } }Run Pi
# Start Pi in your project directory: pi
- Docker Model Runner
How to use BlazingCustoms/pybytecode-v3-1.5b with Docker Model Runner:
docker model run hf.co/BlazingCustoms/pybytecode-v3-1.5b:F16
- Lemonade
How to use BlazingCustoms/pybytecode-v3-1.5b with Lemonade:
Pull the model
# Download Lemonade from https://lemonade-server.ai/ lemonade pull BlazingCustoms/pybytecode-v3-1.5b:F16
Run and chat with the model
lemonade run user.pybytecode-v3-1.5b-F16
List all available models
lemonade list
- Hermes Agent
How to use BlazingCustoms/pybytecode-v3-1.5b with Hermes Agent:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf BlazingCustoms/pybytecode-v3-1.5b:F16
Configure Hermes
# Install Hermes: curl -fsSL https://hermes-agent.nousresearch.com/install.sh | bash hermes setup # Point Hermes at the local server: hermes config set model.provider custom hermes config set model.base_url http://127.0.0.1:8080/v1 hermes config set model.default BlazingCustoms/pybytecode-v3-1.5b:F16
Run Hermes
hermes
- Atomic Chat
- OpenClaw
How to use BlazingCustoms/pybytecode-v3-1.5b with OpenClaw:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf BlazingCustoms/pybytecode-v3-1.5b:F16
Configure OpenClaw
# Install OpenClaw: npm install -g openclaw@latest # Register the local server and set it as the default model: openclaw onboard --non-interactive --mode local \ --auth-choice custom-api-key \ --custom-base-url http://127.0.0.1:8080/v1 \ --custom-model-id "BlazingCustoms/pybytecode-v3-1.5b:F16" \ --custom-provider-id llama-cpp \ --custom-compatibility openai \ --custom-text-input \ --accept-risk \ --skip-health
Run OpenClaw
openclaw agent --local --agent main --message "Hello from Hugging Face"
docs: disclose the 19.4% ast.unparse formatting ceiling (new section 8 + section 7 table row)
Browse files- ORACLE-LIMITS.md +54 -0
ORACLE-LIMITS.md
CHANGED
|
@@ -131,3 +131,57 @@ the unverified remainder as errors understates the model; treating it as correct
|
|
| 131 |
| `-O` mismatch collapse, docstrings unprovable | this file; `EVAL.md`; `weights/MODEL-CARD.md` |
|
| 132 |
| unverified ≠ wrong | this file; `EVAL.md`; `weights/MODEL-CARD.md`; `harness/README.md` |
|
| 133 |
| 3.13 / PyInstaller untested | this file; `EVAL.md`; `weights/MODEL-CARD.md` |
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 131 |
| `-O` mismatch collapse, docstrings unprovable | this file; `EVAL.md`; `weights/MODEL-CARD.md` |
|
| 132 |
| unverified ≠ wrong | this file; `EVAL.md`; `weights/MODEL-CARD.md`; `harness/README.md` |
|
| 133 |
| 3.13 / PyInstaller untested | this file; `EVAL.md`; `weights/MODEL-CARD.md` |
|
| 134 |
+
| **19.4% formatting ceiling (§8)** | **this file** |
|
| 135 |
+
|
| 136 |
+
## 8. The formatting ceiling: 19.4% of real modules cannot certify a re-formatted answer
|
| 137 |
+
|
| 138 |
+
This section was added on 2026-09-11. It discloses a ceiling that was measured earlier and was
|
| 139 |
+
not previously written down here, which left §7's promise of a complete accounting unmet.
|
| 140 |
+
|
| 141 |
+
**`ast.unparse` is not bytecode-preserving.** Compiling a source file, and compiling that same
|
| 142 |
+
file after an AST round trip, do not always produce the same code object:
|
| 143 |
+
|
| 144 |
+
```python
|
| 145 |
+
compile(ast.unparse(ast.parse(src))) != compile(src) # for 100 of 516 stdlib modules
|
| 146 |
+
```
|
| 147 |
+
|
| 148 |
+
Measured on all 516 CPython 3.12 stdlib modules: **100 fail the round trip — an 80.6% pass rate,
|
| 149 |
+
so 19.4% of real modules contain at least one affected construct.**
|
| 150 |
+
|
| 151 |
+
**The cause**, reduced to a repro verified on the measurement box:
|
| 152 |
+
|
| 153 |
+
```python
|
| 154 |
+
a = {name for name, value in ns.items() if getattr(value, 'x', False)} # POP_JUMP_IF_FALSE
|
| 155 |
+
a = {name # POP_JUMP_IF_TRUE
|
| 156 |
+
for name, value in ns.items() # + JUMP_BACKWARD
|
| 157 |
+
if getattr(value, 'x', False)}
|
| 158 |
+
```
|
| 159 |
+
|
| 160 |
+
Identical AST, different bytecode. CPython 3.12's CFG optimiser lays out basic blocks according to
|
| 161 |
+
the **physical line layout** of a comprehension. An `if` *statement* wrapped the same way does not
|
| 162 |
+
change; comprehensions do.
|
| 163 |
+
|
| 164 |
+
**What this means for anyone reading a score from this oracle.**
|
| 165 |
+
|
| 166 |
+
- It is a ceiling on the **oracle**, not a defect in any model. A decompilation that is
|
| 167 |
+
semantically perfect but wraps a comprehension across lines differently from the original will
|
| 168 |
+
**not certify**. It is scored as unverified, which per §5 means *unknown*, not *wrong*.
|
| 169 |
+
- It applies to **every system measured against this oracle** — ours, PyLingual's, and any other.
|
| 170 |
+
It is not a differential advantage or disadvantage to anybody.
|
| 171 |
+
- It is **whole-module scale**: 19.4% is the fraction of *modules* containing at least one
|
| 172 |
+
affected construct, not the fraction of constructs affected. Function- and class-level units are
|
| 173 |
+
affected at a lower rate, because the chance of containing a comprehension is lower.
|
| 174 |
+
- **It compounds with §2's 0.33% floor and §3's `-O` collapse.** These are independent sources of
|
| 175 |
+
false rejection; none of them ever produces a false accept.
|
| 176 |
+
|
| 177 |
+
**Why the numbers already published are not invalidated by this disclosure.** Every published
|
| 178 |
+
score was produced by grading a model's output against a reference *source*, never against an
|
| 179 |
+
`ast.unparse` round trip, so no published figure was computed through the lossy path. This section
|
| 180 |
+
discloses a bound on how high any such score could ever go; it does not move one. It also rules
|
| 181 |
+
out a class of pipeline we did not build: an AST-splice reassembler would have discarded roughly
|
| 182 |
+
one module in five before asking the model a single question.
|
| 183 |
+
|
| 184 |
+
**Measured, and stated as such.** 100/516 is a direct count over the CPython 3.12 stdlib on the
|
| 185 |
+
measurement box. The corresponding ceiling on the internal 522-row module-shaped eval set — 92.5%
|
| 186 |
+
overall, 88.3% on rows of ≥400 representation lines — is measured on a different set and is
|
| 187 |
+
reported with those internal results rather than here.
|