Instructions to use Mattimax/DAC5-3B with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use Mattimax/DAC5-3B with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("text-generation", model="Mattimax/DAC5-3B") messages = [ {"role": "user", "content": "Who are you?"}, ] pipe(messages)# pip install -U transformers accelerate # Load model directly from transformers import AutoTokenizer, AutoModelForCausalLM tokenizer = AutoTokenizer.from_pretrained("Mattimax/DAC5-3B") model = AutoModelForCausalLM.from_pretrained("Mattimax/DAC5-3B", device_map="auto") messages = [ {"role": "user", "content": "Who are you?"}, ] inputs = tokenizer.apply_chat_template( messages, add_generation_prompt=True, tokenize=True, return_dict=True, return_tensors="pt", ).to(model.device) outputs = model.generate(**inputs, max_new_tokens=256) print(tokenizer.decode(outputs[0][inputs["input_ids"].shape[-1]:])) - Inference
- Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- vLLM
How to use Mattimax/DAC5-3B with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "Mattimax/DAC5-3B" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "Mattimax/DAC5-3B", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker
docker model run hf.co/Mattimax/DAC5-3B
- SGLang
How to use Mattimax/DAC5-3B with SGLang:
Install from pip and serve model
# Install SGLang from pip: pip install sglang # Start the SGLang server: python3 -m sglang.launch_server \ --model-path "Mattimax/DAC5-3B" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "Mattimax/DAC5-3B", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker images
docker run --gpus all \ --shm-size 32g \ -p 30000:30000 \ -v ~/.cache/huggingface:/root/.cache/huggingface \ --env "HF_TOKEN=<secret>" \ --ipc=host \ lmsysorg/sglang:latest \ python3 -m sglang.launch_server \ --model-path "Mattimax/DAC5-3B" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "Mattimax/DAC5-3B", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }' - Docker Model Runner
How to use Mattimax/DAC5-3B with Docker Model Runner:
docker model run hf.co/Mattimax/DAC5-3B
Download README.md from Mattimax/DAC5-3B: direct link, hf CLI and curl.
- Browser
- Download file 4.37 kB
-
https://huggingface.co/Mattimax/DAC5-3B/resolve/6db17da55af21807a403d21ec8f4ffb0c4b19c9e/README.md
- Command line
-
hf download hf://Mattimax/DAC5-3B@6db17da55af21807a403d21ec8f4ffb0c4b19c9e/README.md
-
curl -L -o README.md https://huggingface.co/Mattimax/DAC5-3B/resolve/6db17da55af21807a403d21ec8f4ffb0c4b19c9e/README.md
license: mit
datasets:
- Mattimax/DACMini_Refined
- Mattimax/Camoscio-ITA
language:
- it
- en
library_name: transformers
tags:
- DAC
- M.INC.
- conversational
📘 Model Card — Mattimax/DAC5-3B
🧠 Informazioni Generali
Nome: Mattimax/DAC5-3B Serie: DAC (DATA-AI Chat) – 5ª versione Autore: Mattimax Research Lab / Azienda: MINC01 Base Model: Qwen – Qwen2.5-3B-Instruct
DAC5-3B è attualmente il modello più avanzato e sperimentale della serie DAC, progettato per massimizzare qualità conversazionale e performance tecnica su architettura 3B.
🏗 Architettura Tecnica
Core Architecture
- Architettura: Qwen2ForCausalLM
- Parametri: ~3B
- Numero layer: 36
- Hidden size: 2048
- Intermediate size: 11008
- Attention heads: 16
- Key/Value heads (GQA): 2
- Attivazione: SiLU
- Norm: RMSNorm (eps 1e-6)
- Tie word embeddings: Yes
Attention
- Full attention su tutti i 36 layer
- Attention dropout: 0.0
- Sliding window: disabilitato
- GQA (Grouped Query Attention) → maggiore efficienza memoria
Positional Encoding
- Max position embeddings: 32768
- RoPE theta: 1,000,000
- RoPE scaling: None
Precision & Performance
- Torch dtype: bfloat16
- Quantizzazione training: 4-bit (NF4)
- Cache abilitata per inference
- Ottimizzato con Unsloth (fixed build 2026.2.1)
Tokenizer
- Vocab size: 151,936
- EOS token id: 151645
- PAD token id: 151654
- Chat template: Qwen 2.5 ChatML
🎯 Obiettivo del Modello
DAC5-3B è stato progettato per:
- 🇮🇹 Massima qualità in italiano
- ⚡ Alta efficienza su GPU consumer
- 🧩 Conversazione coerente multi-turn
- 🛠️ Supporto tecnico e coding leggero
- 🧠 Migliore stabilità rispetto ai DAC precedenti
È un modello orientato a sviluppatori indipendenti, maker e sistemi offline.
📚 Dataset & Specializzazione
Il fine-tuning supervisionato è stato effettuato su un mix altamente selezionato di dataset italiani:
- Camoscio-ITA
- DACMini Refined
- Conversazioni sintetiche italiane ad alta qualità
Strategia
- Dataset limitato ma ad alta densità informativa (~20k esempi)
- Minimizzazione del rumore
- Focus su chiarezza e coerenza
- Riduzione delle risposte generiche tipiche dei 3B
🚀 Capacità Principali
DAC5-3B eccelle in:
- Spiegazioni tecniche
- Scrittura strutturata
- Programmazione livello medio
- Traduzione IT ↔ EN
- Brainstorming progettuale
- Assistenti locali offline
- Supporto allo studio
📊 Differenze rispetto ai DAC precedenti
✔ Maggiore stabilità nelle risposte lunghe ✔ Meno ripetizioni ✔ Migliore controllo del tono ✔ Risposte più dirette ✔ Migliore allineamento alle istruzioni
DAC5 rappresenta il punto più alto raggiunto finora nella serie.
⚠️ Limitazioni
- Contesto di training effettivo: 1024 token
- Non ottimizzato per tool calling complesso
- Non specializzato in matematica avanzata
- Può degradare su reasoning multi-step molto profondo
- Modello sperimentale
💻 Requisiti Hardware
Inference consigliata
- GPU 6–8GB VRAM (quantizzato)
- Oppure CPU moderna con GGUF
Compatibile con:
- PC consumer
- Mini workstation
- Sistemi edge
- Setup locali offline
🔬 Filosofia DAC
La serie DAC nasce con l'obiettivo di:
Spingere al massimo modelli compatti, ottimizzando qualità reale invece di scalare solo i parametri.
DAC5-3B è il risultato più maturo di questa filosofia: qualità elevata su architettura 3B con risorse contenute.
🧪 Stato del Modello
🟡 Sperimentale ma stabile È il miglior modello della serie DAC fino ad oggi, ma rimane parte di un ciclo evolutivo continuo.
📚 Citation
Se utilizzi Mattimax/DAC5-3B nei tuoi lavori di ricerca, progetti o pubblicazioni, puoi citarlo nel seguente modo:
@misc{mattimax_dac5_3b_2026,
author = {Mattimax},
title = {DAC5-3B: Fifth Iteration of the Dynamic Adaptive Core Series},
year = {2026},
publisher = {Hugging Face},
organization = {MINC01},
note = {Experimental Italian-specialized 3B language model},
url = {https://huggingface.co/Mattimax/DAC5-3B}
}
Citazione testuale breve:
Mattimax. DAC5-3B: Fifth Iteration of the Dynamic Adaptive Core Series. 2026. MINC01 Research Lab.