Text Generation
Transformers
Safetensors
complex_kda
complex-kda
linear-attention
kimi-delta-attention
conversational
custom_code
Instructions to use openeurollm/kda-sigmoid-hybrid-1.3B-100B with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use openeurollm/kda-sigmoid-hybrid-1.3B-100B with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("text-generation", model="openeurollm/kda-sigmoid-hybrid-1.3B-100B", trust_remote_code=True) messages = [ {"role": "user", "content": "Who are you?"}, ] pipe(messages)# Load model directly from transformers import AutoModelForCausalLM model = AutoModelForCausalLM.from_pretrained("openeurollm/kda-sigmoid-hybrid-1.3B-100B", trust_remote_code=True, device_map="auto") - Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- vLLM
How to use openeurollm/kda-sigmoid-hybrid-1.3B-100B with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "openeurollm/kda-sigmoid-hybrid-1.3B-100B" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "openeurollm/kda-sigmoid-hybrid-1.3B-100B", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker
docker model run hf.co/openeurollm/kda-sigmoid-hybrid-1.3B-100B
- SGLang
How to use openeurollm/kda-sigmoid-hybrid-1.3B-100B with SGLang:
Install from pip and serve model
# Install SGLang from pip: pip install sglang # Start the SGLang server: python3 -m sglang.launch_server \ --model-path "openeurollm/kda-sigmoid-hybrid-1.3B-100B" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "openeurollm/kda-sigmoid-hybrid-1.3B-100B", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker images
docker run --gpus all \ --shm-size 32g \ -p 30000:30000 \ -v ~/.cache/huggingface:/root/.cache/huggingface \ --env "HF_TOKEN=<secret>" \ --ipc=host \ lmsysorg/sglang:latest \ python3 -m sglang.launch_server \ --model-path "openeurollm/kda-sigmoid-hybrid-1.3B-100B" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "openeurollm/kda-sigmoid-hybrid-1.3B-100B", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }' - Docker Model Runner
How to use openeurollm/kda-sigmoid-hybrid-1.3B-100B with Docker Model Runner:
docker model run hf.co/openeurollm/kda-sigmoid-hybrid-1.3B-100B
Add paper banner, figure and precise model descriptions
Browse files- .gitattributes +1 -0
- README.md +14 -1
- complexkda.png +3 -0
.gitattributes
CHANGED
|
@@ -33,3 +33,4 @@ saved_model/**/* filter=lfs diff=lfs merge=lfs -text
|
|
| 33 |
*.zip filter=lfs diff=lfs merge=lfs -text
|
| 34 |
*.zst filter=lfs diff=lfs merge=lfs -text
|
| 35 |
*tfevents* filter=lfs diff=lfs merge=lfs -text
|
|
|
|
|
|
| 33 |
*.zip filter=lfs diff=lfs merge=lfs -text
|
| 34 |
*.zst filter=lfs diff=lfs merge=lfs -text
|
| 35 |
*tfevents* filter=lfs diff=lfs merge=lfs -text
|
| 36 |
+
complexkda.png filter=lfs diff=lfs merge=lfs -text
|
README.md
CHANGED
|
@@ -6,11 +6,24 @@ tags:
|
|
| 6 |
- complex-kda
|
| 7 |
- linear-attention
|
| 8 |
- kimi-delta-attention
|
|
|
|
| 9 |
---
|
| 10 |
|
| 11 |
# openeurollm/kda-sigmoid-hybrid-1.3B-100B
|
| 12 |
|
| 13 |
-
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 14 |
|
| 15 |
ComplexKDA is Kimi Delta Attention with a **signed decay gate**: the per-channel
|
| 16 |
decay `alpha` is allowed to take either sign, `alpha in [-1, 1]`, instead of
|
|
|
|
| 6 |
- complex-kda
|
| 7 |
- linear-attention
|
| 8 |
- kimi-delta-attention
|
| 9 |
+
- arxiv:2609.24797
|
| 10 |
---
|
| 11 |
|
| 12 |
# openeurollm/kda-sigmoid-hybrid-1.3B-100B
|
| 13 |
|
| 14 |
+
[](https://arxiv.org/abs/2609.24797)
|
| 15 |
+
[](https://github.com/OpenEuroLLM/ComplexKDA)
|
| 16 |
+
[](https://opensource.org/licenses/MIT)
|
| 17 |
+
|
| 18 |
+
> **Paper:** https://arxiv.org/abs/2609.24797
|
| 19 |
+
>
|
| 20 |
+
> **Code:** https://github.com/OpenEuroLLM/ComplexKDA
|
| 21 |
+
>
|
| 22 |
+
> **Authors:** Julien Siems, Riccardo Grazzi, Korbinian Pöppel, Jaisidh Singh, Arber Zela, Timur Carstensen, Jenia Jitsev, Frank Hutter, Volkan Cevher, Antonio Orvieto, Aaron Klein
|
| 23 |
+
|
| 24 |
+
A **KDA baseline** (bounded sigmoid gate) hybrid language model (1.36B parameters) -- linear layers with full attention every 4th layer -- from the ComplexKDA release.
|
| 25 |
+
|
| 26 |
+
<img src="complexkda.png" width="900" alt="ComplexKDA">
|
| 27 |
|
| 28 |
ComplexKDA is Kimi Delta Attention with a **signed decay gate**: the per-channel
|
| 29 |
decay `alpha` is allowed to take either sign, `alpha in [-1, 1]`, instead of
|
complexkda.png
ADDED
|
Git LFS Details
|