Text Generation
Transformers
ONNX
Safetensors
echo_hybrid
text-generation-inference
conversational
custom_code
Instructions to use mrs83/Kurtis-EON1-Hybrid-2B-v0.1.2 with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use mrs83/Kurtis-EON1-Hybrid-2B-v0.1.2 with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("text-generation", model="mrs83/Kurtis-EON1-Hybrid-2B-v0.1.2", trust_remote_code=True) messages = [ {"role": "user", "content": "Who are you?"}, ] pipe(messages)# Load model directly from transformers import AutoModelForCausalLM model = AutoModelForCausalLM.from_pretrained("mrs83/Kurtis-EON1-Hybrid-2B-v0.1.2", trust_remote_code=True, device_map="auto") - Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- vLLM
How to use mrs83/Kurtis-EON1-Hybrid-2B-v0.1.2 with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "mrs83/Kurtis-EON1-Hybrid-2B-v0.1.2" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "mrs83/Kurtis-EON1-Hybrid-2B-v0.1.2", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker
docker model run hf.co/mrs83/Kurtis-EON1-Hybrid-2B-v0.1.2
- SGLang
How to use mrs83/Kurtis-EON1-Hybrid-2B-v0.1.2 with SGLang:
Install from pip and serve model
# Install SGLang from pip: pip install sglang # Start the SGLang server: python3 -m sglang.launch_server \ --model-path "mrs83/Kurtis-EON1-Hybrid-2B-v0.1.2" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "mrs83/Kurtis-EON1-Hybrid-2B-v0.1.2", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker images
docker run --gpus all \ --shm-size 32g \ -p 30000:30000 \ -v ~/.cache/huggingface:/root/.cache/huggingface \ --env "HF_TOKEN=<secret>" \ --ipc=host \ lmsysorg/sglang:latest \ python3 -m sglang.launch_server \ --model-path "mrs83/Kurtis-EON1-Hybrid-2B-v0.1.2" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "mrs83/Kurtis-EON1-Hybrid-2B-v0.1.2", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }' - Docker Model Runner
How to use mrs83/Kurtis-EON1-Hybrid-2B-v0.1.2 with Docker Model Runner:
docker model run hf.co/mrs83/Kurtis-EON1-Hybrid-2B-v0.1.2
File size: 5,928 Bytes
1ad193e 4ac8c43 1ad193e 4ac8c43 1ad193e 4ac8c43 1ad193e 4ac8c43 | 1 2 3 4 5 6 7 8 9 10 11 12 13 14 15 16 17 18 19 20 21 22 23 24 25 26 27 28 29 30 31 32 33 34 35 36 37 38 39 40 41 42 43 44 45 46 47 48 49 50 51 52 53 54 55 56 57 58 59 60 61 62 63 64 65 66 67 68 69 70 71 72 73 74 75 76 77 78 79 80 81 82 83 84 85 86 87 88 89 90 91 92 93 94 95 96 97 98 99 100 101 102 103 104 105 106 107 108 109 110 111 112 113 114 115 116 117 118 119 120 121 122 123 124 125 126 127 128 129 130 131 132 133 134 135 136 137 138 139 140 141 142 143 144 | """
echo_hybrid/configuration_hybrid.py
βββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββ
HybridEchoConfig: extends Qwen2Config with DSRN memory-injector parameters.
Design rationale
ββββββββββββββββ
Rather than inventing an entirely new config, we subclass Qwen2Config so that
every existing Qwen2 hyper-parameter (hidden_size, num_hidden_layers, etc.) is
available without duplication. The only additions are the four DSRN-specific
fields documented below.
CRITICAL NOTES (from AGENTS.md)
βββββββββββββββββββββββββββββββββ
β’ model_type MUST be "echo_hybrid" so AutoConfig routing works after
AutoConfig.register("echo_hybrid", HybridEchoConfig).
β’ Do NOT use this config with EchoForCausalLM β that model expects EchoConfig.
"""
from transformers import Qwen2Config, Qwen3Config
class HybridEchoConfig(Qwen2Config):
"""
Qwen2Config subclass that adds DSRN memory-injector fields.
New fields
ββββββββββ
dsrn_state_dim : int
Dimension of the c_t slow-state vector maintained by each
DSRNMemoryInjector. Defaults to 512. Can be set equal to
hidden_size (896 for Qwen2-0.5B) for a richer slow-state, at the
cost of extra parameters per injector.
dsrn_injection_stride : int
Insert one DSRNMemoryInjector after every N transformer layers.
For Qwen2-0.5B (24 layers) the default of 4 yields 6 injectors.
dsrn_use_triton : bool
Route the parallel scan to the custom Triton kernel defined in
echo_hf/triton_scan.py. Disabled by default because the Triton
kernel targets CUDA/ROCm and is not available everywhere.
gate_bias_init : float
Initial value of linear_gate.bias in every injector. A positive
value (~1.0) keeps memory gates open at init, allowing gradients to
flow into c_t immediately. Increase to 2.0 if c_t norms do not
grow beyond ~0.0 after Phase-1 warm-up.
use_kv_cache : bool
Controls the Qwen2 backbone KV-cache. Independent of use_cache
(DSRN state return).
- True (default / recommended): Standard Hybrid mode β mode 2.
Backbone KV-cache active; attention handles fast-state, DSRN handles
slow-state. Best quality and lowest peak VRAM.
- False: Ablation / stateless mode β mode 1.
Backbone KV-cache disabled; every forward re-feeds the full growing
context so attention stays coherent. DSRN slow-state is the sole
cross-step memory. Useful for ablation studies and "Attention Tax"
vs "Recurrent Gain" benchmarks.
"""
model_type = "echo_hybrid"
def __init__(
self,
dsrn_state_dim: int = 512,
dsrn_injection_stride: int = 4,
dsrn_use_triton: bool = False,
gate_bias_init: float = 1.0,
use_kv_cache: bool = True, # Kill-switch: False = DSRN-only ablation
output_surprise_gate_logits: bool = False,
surprise_temperature_alpha: float = 0.0,
**kwargs,
):
super().__init__(**kwargs)
self.dsrn_state_dim = dsrn_state_dim
self.dsrn_injection_stride = dsrn_injection_stride
self.dsrn_use_triton = dsrn_use_triton
self.gate_bias_init = gate_bias_init
self.use_kv_cache = use_kv_cache
self.output_surprise_gate_logits = output_surprise_gate_logits
self.surprise_temperature_alpha = surprise_temperature_alpha
self.auto_map = {
"AutoConfig": "configuration_hybrid.HybridEchoConfig",
"AutoModel": "modeling_hybrid.HybridEchoModel",
"AutoModelForCausalLM": "modeling_hybrid.HybridEchoForCausalLM",
}
class Qwen3HybridEchoConfig(Qwen3Config):
"""
Qwen3Config subclass that adds DSRN memory-injector fields.
Identical to HybridEchoConfig but inherits from Qwen3Config instead of
Qwen2Config, supporting the Qwen3 model family (e.g. Qwen/Qwen3-0.6B).
New fields
----------
dsrn_state_dim : int
Dimension of the c_t slow-state vector maintained by each
DSRNMemoryInjector. Defaults to 512.
dsrn_injection_stride : int
Insert one DSRNMemoryInjector after every N transformer layers.
For Qwen3-0.6B (28 layers) the default of 4 yields 7 injectors.
dsrn_use_triton : bool
Route the parallel scan to the custom Triton kernel.
gate_bias_init : float
Initial value of linear_gate.bias in every injector.
use_kv_cache : bool
Controls the backbone KV-cache. Independent of use_cache (DSRN state return).
"""
model_type = "qwen3_echo_hybrid"
def __init__(
self,
dsrn_state_dim: int = 512,
dsrn_injection_stride: int = 4,
dsrn_use_triton: bool = False,
gate_bias_init: float = 1.0,
use_kv_cache: bool = True,
output_surprise_gate_logits: bool = False,
surprise_temperature_alpha: float = 0.0,
**kwargs,
):
super().__init__(**kwargs)
self.dsrn_state_dim = dsrn_state_dim
self.dsrn_injection_stride = dsrn_injection_stride
self.dsrn_use_triton = dsrn_use_triton
self.gate_bias_init = gate_bias_init
self.use_kv_cache = use_kv_cache
self.output_surprise_gate_logits = output_surprise_gate_logits
self.surprise_temperature_alpha = surprise_temperature_alpha
self.auto_map = {
"AutoConfig": "configuration_hybrid.Qwen3HybridEchoConfig",
"AutoModel": "modeling_hybrid.Qwen3HybridEchoModel",
"AutoModelForCausalLM": "modeling_hybrid.Qwen3HybridEchoForCausalLM",
}
|