[Research / Experiment] Native Dual-Regime Gemma: Sub-35ms Motor Reflexes & Zero-Cost KV-Cache Handoff on Gemma-2-2B

#84
by savhascelik - opened

Hi everyone,

We wanted to share an open-source experimental architecture we've developed on top of Google's Gemma-2-2B-IT: Native Dual-Regime Inference.

The Problem: Tool-Calling Latency & Conversational Drift

In agentic workflows, spending 35–50 autoregressive tokens solely to select an action or tool introduces latency bottlenecks (600–900ms on GPU), wastes compute, and frequently triggers parsing failures due to natural language preambles.

The Solution: Native Dual-Regime (Zero Fine-Tuning)

Inspired by the System 1 motor reflex concept from TypeSafe AI's Jev, we built a 100% native execution pipeline without auxiliary heads or fine-tuning:

  1. System 1 (Reflex Phase): Evaluates bounded candidate actions in a single forward pass by slicing logits directly from the native language modeling head ($W_{lm}$). Discrete decisions resolve in <35ms consuming exactly 0 tokens.

  2. System 2 (KV-Cache Handoff): If arguments are required, the precomputed Key-Value attention matrices (past_key_values) are handed off directly to causal generation. Prefill latency drops to 0
    ms
    , and prefix-clamping eliminates conversational chatter.


    Empirical Benchmark Summary (40 Real-World Kaggle Scenarios)

    We evaluated four regimes on identical Kaggle T4 GPU environments under strict Python standard json.loads() validation (zero regex tolerance):

    Architecture Action Accuracy Strict Valid JSON Rate Mean Latency Speedup Tokens Burned Token Savings
    Baseline Causal LM (Control) 90.0% (36/40) 62.5% (25/40) 22,598 ms 1.0x (Base) 1,485 0% (Base)
    Cascaded Dual-Regime 92.5% (37/40) 75.0% (30/40) 7,100 ms 3.18x 132 -91.1%
    Speculative Dual-Regime ($\tau=0.75$) 92.5% (37/40) 100.0% (40/40) 5,448 ms 4.15x 23 -98.5%
    Continuous Cognitive Anchoring 92.5% (37/40) 62.5% (25/40) 9,572 ms 2.36x 277 -81.3%

    Key Takeaways & Scope:

    • Hallucination Suppression: System 1 logit slicing eliminated tool hallucinations (e.g. baseline generated a non-existent set_urgent_meeting tool, whereas Dual-Regime locked the correct candidate).
    • Early Aborts: In the Speculative regime, 38/40 tasks (95%) were resolved with 0 generated tokens.
    • Experimental Scope: These findings represent an initial proof-of-concept on 40 curated scenarios using Gemma-2-2B. We believe this provides a strong architectural signal that scaling this mechanism to larger models (e.g. Gemma-2-9B / 27B) will deliver even sharper calibration and robust agent performance.

    All code, reproducible Kaggle notebooks, and raw evaluation datasets are open-sourced:
    👉 Repository: https://github.com/savhascelik/dual-head-gemma

    We would love to hear feedback and thoughts from the Gemma and open-source agent community!

Sign up or log in to comment