Beexly commited on
Commit
44db19f
·
verified ·
1 Parent(s): 6f51848

gse brain engine: MiMo-V2.6 4-bit on ZeroGPU (free tier)

Browse files
Files changed (3) hide show
  1. README.md +35 -7
  2. app.py +85 -0
  3. requirements.txt +7 -0
README.md CHANGED
@@ -1,13 +1,41 @@
1
  ---
2
- title: Mimo Brain Engine
3
- emoji: 🐨
4
- colorFrom: indigo
5
- colorTo: pink
6
  sdk: gradio
7
- sdk_version: 6.29.0
8
- python_version: '3.13'
9
  app_file: app.py
10
  pinned: false
 
11
  ---
12
 
13
- Check out the configuration reference at https://huggingface.co/docs/hub/spaces-config-reference
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
  ---
2
+ title: GSE Brain Engine (MiMo-V2.6, ZeroGPU)
3
+ emoji: 🧠
4
+ colorFrom: blue
5
+ colorTo: purple
6
  sdk: gradio
7
+ sdk_version: 4.44.0
 
8
  app_file: app.py
9
  pinned: false
10
+ license: mit
11
  ---
12
 
13
+ # GSE Brain Engine — MiMo-V2.6 on ZeroGPU
14
+
15
+ The GSE NFL intelligence engine's LLM specialist layer: Xiaomi **MiMo-V2.6-Distill-Qwen-9B**
16
+ (MIT license) in 4-bit, served on Hugging Face **ZeroGPU — the free tier**.
17
+
18
+ ## Billing policy (from HF official docs, verified 2026-10-02)
19
+
20
+ - **This Space runs on ZeroGPU hardware only** (`spaces.GPU`, default `large` size).
21
+ ZeroGPU Spaces are free to host and use — free accounts host up to 2, PRO up to 10.
22
+ - **Daily GPU quotas:** unauthenticated 2 min · free 5 min · PRO 40 min (highest
23
+ queue priority). Resets 24h after first GPU use.
24
+ - **Overage draws from the pre-paid credit balance only.** No card is charged
25
+ unless credits were explicitly purchased or auto-recharge is enabled. With no
26
+ credits, over-quota requests fail/queue — they do not bill.
27
+ - **Paid Space hardware (t4-small, a10g, …) is what bills usage-based.**
28
+ This Space is pinned to ZeroGPU; a drift monitor alerts loudly if the
29
+ hardware ever changes. Never upgrade this Space's hardware without
30
+ Garrett's explicit approval.
31
+
32
+ Sources: https://huggingface.co/docs/hub/spaces-zerogpu ·
33
+ https://huggingface.co/docs/hub/billing
34
+
35
+ ## The /chat contract
36
+
37
+ Send VERIFIED DATA plus a grounding prompt. The model reasons (L3 causal
38
+ chains, L4 adversarial review, L5 synthesis) under hard rules: cite provided
39
+ numbers, label inference as INFERENCE with a breaking condition, output the
40
+ requested JSON schema. The deterministic GSE contract (adversary_review,
41
+ correlated_theses, checklist) still disposes — the model proposes.
app.py ADDED
@@ -0,0 +1,85 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ """GSE Brain Engine — MiMo-V2.6 on ZeroGPU (free tier).
2
+
3
+ Serves XiaomiMiMo/MiMo-V2.6-Distill-Qwen-9B (MIT) in 4-bit behind @spaces.GPU.
4
+ ZeroGPU hardware only — this Space can never bill the account (see README.md
5
+ billing policy). The /chat endpoint speaks the GSE grounding contract: every
6
+ prompt should include VERIFIED DATA and demand cited numbers + INFERENCE flags.
7
+ """
8
+
9
+ import spaces
10
+ import torch
11
+ import gradio as gr
12
+ from transformers import (
13
+ AutoModelForImageTextToText,
14
+ AutoProcessor,
15
+ BitsAndBytesConfig,
16
+ )
17
+
18
+ MODEL_ID = "XiaomiMiMo/MiMo-V2.6-Distill-Qwen-9B"
19
+
20
+ SYSTEM_PROMPT = (
21
+ "You are a specialist in the GSE NFL intelligence engine. "
22
+ "You are given VERIFIED DATA in the user message. HARD RULES: "
23
+ "(1) Every numeric claim MUST cite a value from the VERIFIED DATA. "
24
+ "(2) Anything inferred beyond the data MUST be labeled INFERENCE with a breaking_condition. "
25
+ "(3) Output ONLY valid JSON when a JSON schema is requested."
26
+ )
27
+
28
+ print("Loading processor...", flush=True)
29
+ processor = AutoProcessor.from_pretrained(MODEL_ID, trust_remote_code=True)
30
+
31
+ print("Loading model in 4-bit...", flush=True)
32
+ bnb = BitsAndBytesConfig(load_in_4bit=True, bnb_4bit_compute_dtype=torch.float16)
33
+ model = AutoModelForImageTextToText.from_pretrained(
34
+ MODEL_ID,
35
+ quantization_config=bnb,
36
+ device_map="auto",
37
+ torch_dtype=torch.float16,
38
+ trust_remote_code=True,
39
+ )
40
+ model.eval()
41
+ print("Model ready.", flush=True)
42
+
43
+
44
+ @spaces.GPU(duration=120)
45
+ def chat(message: str, history: list):
46
+ """Chat endpoint (api_name=/chat). Text-only; history as Gradio pairs."""
47
+ try:
48
+ messages = [{"role": "system", "content": SYSTEM_PROMPT}]
49
+ for h_user, h_asst in history or []:
50
+ messages.append({"role": "user", "content": h_user})
51
+ messages.append({"role": "assistant", "content": h_asst})
52
+ messages.append({"role": "user", "content": message})
53
+
54
+ text = processor.apply_chat_template(
55
+ messages, add_generation_prompt=True, tokenize=False)
56
+ inputs = processor(text=[text], return_tensors="pt").to("cuda")
57
+ input_len = inputs["input_ids"].shape[1]
58
+
59
+ with torch.inference_mode():
60
+ out = model.generate(
61
+ **inputs,
62
+ max_new_tokens=1024,
63
+ temperature=0.2,
64
+ do_sample=True,
65
+ pad_token_id=processor.tokenizer.eos_token_id,
66
+ )
67
+ response = processor.batch_decode(
68
+ out[:, input_len:], skip_special_tokens=True)[0]
69
+ return response.strip()
70
+ except Exception as exc: # never crash the Space; return the error loudly
71
+ return f"[BRAIN-ENGINE ERROR] {type(exc).__name__}: {exc}"
72
+
73
+
74
+ demo = gr.ChatInterface(
75
+ fn=chat,
76
+ title="GSE Brain Engine (MiMo-V2.6, ZeroGPU)",
77
+ description=(
78
+ "The GSE NFL intelligence engine's LLM specialist layer, served free on "
79
+ "Hugging Face ZeroGPU. Send it VERIFIED DATA + a grounding prompt; it "
80
+ "reasons (L3 chains, L4 adversary, L5 synthesis) and cites numbers."
81
+ ),
82
+ )
83
+
84
+ if __name__ == "__main__":
85
+ demo.launch()
requirements.txt ADDED
@@ -0,0 +1,7 @@
 
 
 
 
 
 
 
 
1
+ gradio>=4.44.0
2
+ torch>=2.8.0
3
+ transformers>=4.57.0
4
+ accelerate>=1.0.0
5
+ bitsandbytes>=0.44.0
6
+ spaces>=0.28.0
7
+ pillow>=10.0.0