File size: 1,717 Bytes
4a61d49 6059b8c 4a61d49 6059b8c 4a61d49 6059b8c | 1 2 3 4 5 6 7 8 9 10 11 12 13 14 15 16 17 18 19 20 21 22 23 24 25 26 27 28 29 30 31 32 33 34 35 36 37 38 39 40 41 42 43 44 | ---
title: amx-reasoning-v1-instruct QA
emoji: 🧮
colorFrom: green
colorTo: red
sdk: gradio
sdk_version: 6.26.0
app_file: app.py
short_description: Passage QA with a 7.5M-param CPU-trained LM
python_version: "3.12"
startup_duration_timeout: 30m
license: apache-2.0
models:
- gdiamos/amx-reasoning-v1-instruct
---
# amx-reasoning-v1-instruct — passage QA
Demo for [`gdiamos/amx-reasoning-v1-instruct`](https://huggingface.co/gdiamos/amx-reasoning-v1-instruct):
a 7,492,448-parameter causal LM (3,315,552 active per token) trained end to end
on a single Intel Emerald Rapids CPU core. Give it a passage and a question and
it extracts a one-or-two-word answer.
The demo runs the model's own reference path, unmodified:
- the `m2r` package that trained it (the architecture is not a `transformers`
one — `AutoModelForCausalLM` will not load it),
- `render_prompt(..., thinking=False)` for the exact prompt format,
- greedy argmax decoding, stopped on the `<SPECIAL_12>` end-of-turn token,
- the 800-token vocabulary mask from `generation.json` applied before the
argmax (required for correct output; the demo exposes a toggle so you can see
what happens without it).
Inference runs on CPU, which is the hardware the model was designed and trained
for — a forward pass at these dimensions is a few GFLOP.
## Example attribution
Example passages are drawn from the datasets the model card reports on:
[DROP](https://huggingface.co/datasets/ucinlp/drop) and
[SQuAD v2](https://huggingface.co/datasets/rajpurkar/squad_v2) (CC BY-SA 4.0),
and [databricks-dolly-15k](https://huggingface.co/datasets/databricks/databricks-dolly-15k)
(CC BY-SA 3.0). The first example is the model repo's own `example.py`.
|