Llama-3.2-1B-RYS-10-13-GGUF

A layer-duplication ("RYS" — Repeat Your Self, David Ng) variant of meta-llama/Llama-3.2-1B-Instruct: transformer layers 10–12 are duplicated, expanding the stack from 16 to 19 layers. No training, no merging, no weight changes — purely structural duplication. GGUF (imatrix Q-quants).

⚠️ Evaluation status — please read (updated 2026-06)

The large reasoning gain originally reported for this model was a measurement artifact, not a real capability gain. This card is being corrected to say so plainly.

The original card reported reasoning 0.00% → 64.71%. That 0% baseline came from a degraded inference setup, not from the model. On a current llama.cpp build the unmodified Llama-3.2-1B-Instruct already scores about 52.94% on the same reasoning probe, and this (10,13) duplication adds ~0 over that baseline (and slightly lowers the EQ probe). The headline "+64.71" was the old stack's broken floor rising back to normal — not something the duplication unlocked.

Why: the scores come from a lightweight search probe (16 math / 16 EQ / 17 reasoning questions, greedy-decoded) used to locate productive layer blocks — not a validated benchmark. Reasoning moves in steps of 1/17 ≈ 5.9%, so the deltas are coarse, and a degraded baseline can manufacture a huge apparent gain.

Independent benchmark (lm-eval-harness, GSM8K 5-shot, N=100): base 34% strict-match (38% flexible); RYS (10,13) 23% strict (27% flexible) — same 100 problems. So on a real benchmark the duplication does not improve reasoning and if anything lowers it. This both confirms the base is far from "0%" and refutes the "+64.71pp" claim directionally. (N=100; a larger paired run would tighten the magnitude, but the direction is clear.)

Bottom line: treat this as a normal Llama-3.2-1B-Instruct with layers 10–12 duplicated. Published for transparency and reproducibility of the RYS sweep — not as an improved reasoner.

Original sweep numbers (search probe — kept for the record)

probe reported baseline reported (10,13) re-test note
Reasoning (17 q) 0.00% 64.71% baseline was a degraded-stack artifact; correct-stack baseline ≈ 52.94%, (10,13) Δ ≈ 0. Real GSM8K (N=100): base 34%, (10,13) 23% — duplication lowers it
EQ (16 q) 27.11 90.12 baseline also stack-dependent; not a validated EQ benchmark
Math (16 q) 0.536 0.711 search-probe score, unconfirmed

Run it

llama-server -m Llama-3.2-1B-RYS-10-13-Q4_K_M.gguf -ngl 99

Method · data · attribution

  • Method: layer duplication — Repeat Your Self (David Ng); toolkit llm-circuit-finder (alainnothere).
  • Raw sweep data: rys-sovereign-collection-v2.
  • Built by John Broadway with Claude. The method and the raw data are real and reproducible; the interpretation of the original probe deltas as capability is what this update corrects.

License

Llama 3.2 Community License (inherits from the base model).

Downloads last month
8
GGUF
Model size
1B params
Architecture
llama
Hardware compatibility
Log In to add your hardware

4-bit

Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for john-broadway/Llama-3.2-1B-RYS-10-13-GGUF

Quantized
(418)
this model

Collection including john-broadway/Llama-3.2-1B-RYS-10-13-GGUF