0xSojalSec's picture kvyb's picture
Duplicate from LessThanThreeAI/Qwen3.8-27B-Humanlike-Chat-GGUF
f1f655f
|
Raw History Blame Contribute Delete
4.77 kB

Step-863 Artifact and Reasoning Findings

Historical step-863 characterization as of 2026-09-04, not results for the current step-576 strength-0.7 rebuild. Current artifact lineage, public WikiText calibration and release gates are recorded in artifact-manifest.json. No private conversation text is included here.

Artifact lineage

The corrected GGUF path uses the pinned Huihui Qwen3.8-27B base and step-863 rank-256, alpha-32 LoRA. The adapter conversion includes the required value-head reorder correction. The original merged F16 matched the approved hot-LoRA eight-turn non-thinking branch exactly in that earlier sampler profile.

All quant tiers derive directly from the same verified F16 master, never by requantizing a lower-bit file. Q6/Q5/Q4/Q3 use the shared 61,173-token activation-calibration corpus's importance matrix, covering 496 tensors, with 96 recurrent gate weights protected at Q8_0. Calibration requested balanced off/low/medium render strata, but off and medium complete-chat renderings were identical; it did not simulate generated reasoning trajectories.

The new BF16 companion is separately merged from BF16 base plus FP32 adapter with a direct F32-to-BF16 output cast. Non-thinking qualification is separate from the older F16 reasoning characterization below.

Reasoning characterization

Three recorded seeds (42, 31415, 271828), low and medium reasoning, eight-turn generated-history branches, no system prompt, no inference retries. Each branch used its own generated finals; strict reasoning-only EOS outputs were retained as no-answer and omitted from assistant history. Turns within a branch are correlated.

Artifact Recorded turns Valid finals Strict reasoning-only no-answers Other invalid rows
Merged F16 48 18 30 0
Q8_0 45 of 48 planned 12 32 1
Q6_K 48 15 33 0
Q5_K_M 48 13 35 0

Q8 low/seed-271828 stopped at turn five because sampled EOS was outside the captured top-10 alternatives, triggering the frozen token-evidence guard. That five-row capture remains partial; the three unrun turns are not treated as observations or successes. The other independent branches were completed without retrying this branch.

The raw failing generations ended before emitting the reasoning-close token and a final answer. They were not merely stripped by a downstream parser, and they stopped on EOS rather than exhausting the token cap.

Matched first-turn diagnosis

Arm Valid thinking finals
Unadapted BF16 base, stock llama.cpp 6/6
BF16 base plus FP32 LoRA, stock llama.cpp 3/6
Merged F16, stock llama.cpp 3/6
Merged F16 plus explicit reasoning-boundary instruction 3/6
Merged F16 plus reference GDN normalization patch 3/6

Repeating the six F16 first-turn controls with the corrected neutral thinking-penalty sampler reproduced the original prompt hashes, raw outputs and token sequences exactly. The tested prompt instruction and GDN patch did not rescue the failure. The patch is not the release runtime and is not claimed as a fix.

Serverless comparison

A bounded live vLLM/serverless check confirmed identical low-mode prompt token IDs and two successful reasoning-plus-final responses at seeds 42 and 31415. The latter seed fails in the matched llama.cpp controls. The third request was canceled at the diagnostic deadline; medium was not reached. Do not interpret this as a completed six-cell serverless result.

Sampling correction

The earlier evaluator reported non-thinking presence penalty 1.5 but omitted the penalties sampler, so that nonzero penalty was inactive. The source was corrected and the original evaluator preserved. The release's non-thinking checks use the active penalties sampler. Thinking penalties were already neutral, and all 141 newly recorded thinking requests matched the intended numeric settings, so this error does not explain their missing finals.

Operating recommendation

Use explicit thinking-off settings. Release verification is limited to artifact integrity, the private first-message gate and a short non-thinking replay. Do not infer broad conversational quality, reasoning reliability, or cross-runtime equivalence from these checks.

Primary sources