|
Download README.md from hugging-apps/amx-reasoning-v1-qa-demo: direct link, hf CLI and curl.
- Browser
- Download file 1.72 kB
-
https://huggingface.co/spaces/hugging-apps/amx-reasoning-v1-qa-demo/resolve/7637da4b0d3d437d2b4ee7aff6a796c21e94360a/README.md
- Command line
-
hf download hf://spaces/hugging-apps/amx-reasoning-v1-qa-demo@7637da4b0d3d437d2b4ee7aff6a796c21e94360a/README.md
-
curl -L -o README.md https://huggingface.co/spaces/hugging-apps/amx-reasoning-v1-qa-demo/resolve/7637da4b0d3d437d2b4ee7aff6a796c21e94360a/README.md
1.72 kB
metadata
title: amx-reasoning-v1-instruct QA
emoji: 🧮
colorFrom: green
colorTo: red
sdk: gradio
sdk_version: 6.26.0
app_file: app.py
short_description: Passage QA with a 7.5M-param CPU-trained LM
python_version: '3.12'
startup_duration_timeout: 30m
license: apache-2.0
models:
- gdiamos/amx-reasoning-v1-instruct
amx-reasoning-v1-instruct — passage QA
Demo for gdiamos/amx-reasoning-v1-instruct:
a 7,492,448-parameter causal LM (3,315,552 active per token) trained end to end
on a single Intel Emerald Rapids CPU core. Give it a passage and a question and
it extracts a one-or-two-word answer.
The demo runs the model's own reference path, unmodified:
- the
m2rpackage that trained it (the architecture is not atransformersone —AutoModelForCausalLMwill not load it), render_prompt(..., thinking=False)for the exact prompt format,- greedy argmax decoding, stopped on the
<SPECIAL_12>end-of-turn token, - the 800-token vocabulary mask from
generation.jsonapplied before the argmax (required for correct output; the demo exposes a toggle so you can see what happens without it).
Inference runs on CPU, which is the hardware the model was designed and trained for — a forward pass at these dimensions is a few GFLOP.
Example attribution
Example passages are drawn from the datasets the model card reports on:
DROP and
SQuAD v2 (CC BY-SA 4.0),
and databricks-dolly-15k
(CC BY-SA 3.0). The first example is the model repo's own example.py.