--- base_model: google/gemma-4-E4B-it library_name: pytorch pipeline_tag: text-classification tags: - multi-head-latent-control - capability-head - latent-control - model-routing - pytorch --- # MHLC Capability Head - Gemma-4-E4B-it ## Model Description This repository contains a Capability Head from Multi-Head Latent Control. The head reads generated-token hidden-state trajectories from a frozen `google/gemma-4-E4B-it` backbone and predicts whether the backbone is adequate for an instance or should route the instance to a stronger model. This repository is part of the `Multi Head Latent Control capability heads` Hugging Face collection. It contains only the lightweight control head; the frozen backbone weights are not duplicated here. ## Paper https://arxiv.org/abs/2607.14277 ## Code https://github.com/Amirhosein-gh98/Multi-Head-Latent-Control ## Checkpoint Summary | Field | Value | |---|---| | Repository | `AmirhoseinGH/mhlc-capability-head-gemma4-e4b-instruct` | | Base model | `google/gemma-4-E4B-it` | | Model family | `gemma4` | | Thinking mode | `off` | | Variant | `full-trajectory` | | Hidden encoder | `lite` | | Head input mode | `completion_text_only` | | Hidden-state layer | `last` | | Parameters | 2,863,517 | | Weight size | 5.50 MiB | | SHA-256 | `624f4cf4a8df7ca686fbc710dad6262f43c910ed46a7a1a931b7a7357d1fc3aa` | ## Paper Usage Table 2 routing and Table 4 web-search control. Checkpoint selection: Final non-thinking checkpoint referenced by the web-search evaluation job and run config. ## Files - `capability_head.pt`: PyTorch checkpoint containing `head_state` and the embedded training `cfg`. - `capability_head_config.json`: sanitized release and inference metadata. ## Loading the Weights Use the implementation from the linked code repository. The checkpoint can be downloaded and inspected as follows: ```python import torch from huggingface_hub import hf_hub_download checkpoint_path = hf_hub_download( repo_id="AmirhoseinGH/mhlc-capability-head-gemma4-e4b-instruct", filename="capability_head.pt", ) checkpoint = torch.load(checkpoint_path, map_location="cpu", weights_only=False) head_config = checkpoint["cfg"] head_state_dict = checkpoint["head_state"] print(head_config) ``` For end-to-end routing, pass the downloaded checkpoint to `AuxHeadRuntimeConfig.aux_head_ckpt` in `Eval/multi_agenT_bench_v4/compact_multi_agent_shared_optimized_v4_textbench.py`. The base model, thinking mode, hidden encoder, input mode, and hidden-state layer must match this card and `capability_head_config.json`. ## Intended Use These weights are intended for reproducing and extending the capability-based model-routing experiments in Multi-Head Latent Control. They are not standalone language or vision-language models. ## Limitations The head depends on hidden states produced by the exact compatible backbone and prompting/thinking configuration. A routing threshold should be selected and validated for the target deployment distribution. ## Citation ```bibtex @misc{ghasemabadi2026multiheadlatentcontrolunified, title={Multi-Head Latent Control: A Unified Interface for LLM Agent Decision Making}, author={Amirhosein Ghasemabadi and Ruichen Chen and Bahador Rashidi and Di Niu}, year={2026}, eprint={2607.14277}, archivePrefix={arXiv}, primaryClass={cs.CL}, url={https://arxiv.org/abs/2607.14277} } ```