Instructions to use harrier77/LFM2.5-1.2B-ITA-GGUF with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- llama.cpp
How to use harrier77/LFM2.5-1.2B-ITA-GGUF with llama.cpp:
Install (macOS, Linux)
curl -LsSf https://llama.app/install.sh | sh # Start a local OpenAI-compatible server with a web UI: llama serve -hf harrier77/LFM2.5-1.2B-ITA-GGUF:Q5_K_M # Run inference directly in the terminal: llama cli -hf harrier77/LFM2.5-1.2B-ITA-GGUF:Q5_K_M
Install from WinGet (Windows)
winget install llama.cpp # Start a local OpenAI-compatible server with a web UI: llama serve -hf harrier77/LFM2.5-1.2B-ITA-GGUF:Q5_K_M # Run inference directly in the terminal: llama cli -hf harrier77/LFM2.5-1.2B-ITA-GGUF:Q5_K_M
Use pre-built binary
# Download pre-built binary from: # https://github.com/ggerganov/llama.cpp/releases # Start a local OpenAI-compatible server with a web UI: ./llama-server -hf harrier77/LFM2.5-1.2B-ITA-GGUF:Q5_K_M # Run inference directly in the terminal: ./llama-cli -hf harrier77/LFM2.5-1.2B-ITA-GGUF:Q5_K_M
Build from source code
git clone https://github.com/ggerganov/llama.cpp.git cd llama.cpp cmake -B build cmake --build build -j --target llama-server llama-cli # Start a local OpenAI-compatible server with a web UI: ./build/bin/llama-server -hf harrier77/LFM2.5-1.2B-ITA-GGUF:Q5_K_M # Run inference directly in the terminal: ./build/bin/llama-cli -hf harrier77/LFM2.5-1.2B-ITA-GGUF:Q5_K_M
Use Docker
docker model run hf.co/harrier77/LFM2.5-1.2B-ITA-GGUF:Q5_K_M
- LM Studio
- Jan
- Ollama
How to use harrier77/LFM2.5-1.2B-ITA-GGUF with Ollama:
ollama run hf.co/harrier77/LFM2.5-1.2B-ITA-GGUF:Q5_K_M
- Unsloth Desktop
- Pi
How to use harrier77/LFM2.5-1.2B-ITA-GGUF with Pi:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf harrier77/LFM2.5-1.2B-ITA-GGUF:Q5_K_M
Configure the model in Pi
# Install Pi: npm install -g @earendil-works/pi-coding-agent # Add to ~/.pi/agent/models.json: { "providers": { "llama-cpp": { "baseUrl": "http://localhost:8080/v1", "api": "openai-completions", "apiKey": "none", "models": [ { "id": "harrier77/LFM2.5-1.2B-ITA-GGUF:Q5_K_M" } ] } } }Run Pi
# Start Pi in your project directory: pi
- Docker Model Runner
How to use harrier77/LFM2.5-1.2B-ITA-GGUF with Docker Model Runner:
docker model run hf.co/harrier77/LFM2.5-1.2B-ITA-GGUF:Q5_K_M
- Lemonade
How to use harrier77/LFM2.5-1.2B-ITA-GGUF with Lemonade:
Pull the model
# Download Lemonade from https://lemonade-server.ai/ lemonade pull harrier77/LFM2.5-1.2B-ITA-GGUF:Q5_K_M
Run and chat with the model
lemonade run user.LFM2.5-1.2B-ITA-GGUF-Q5_K_M
List all available models
lemonade list
- Hermes Agent
How to use harrier77/LFM2.5-1.2B-ITA-GGUF with Hermes Agent:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf harrier77/LFM2.5-1.2B-ITA-GGUF:Q5_K_M
Configure Hermes
# Install Hermes: curl -fsSL https://hermes-agent.nousresearch.com/install.sh | bash hermes setup # Point Hermes at the local server: hermes config set model.provider custom hermes config set model.base_url http://127.0.0.1:8080/v1 hermes config set model.default harrier77/LFM2.5-1.2B-ITA-GGUF:Q5_K_M
Run Hermes
hermes
- Atomic Chat
- OpenClaw
How to use harrier77/LFM2.5-1.2B-ITA-GGUF with OpenClaw:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf harrier77/LFM2.5-1.2B-ITA-GGUF:Q5_K_M
Configure OpenClaw
# Install OpenClaw: npm install -g openclaw@latest # Register the local server and set it as the default model: openclaw onboard --non-interactive --mode local \ --auth-choice custom-api-key \ --custom-base-url http://127.0.0.1:8080/v1 \ --custom-model-id "harrier77/LFM2.5-1.2B-ITA-GGUF:Q5_K_M" \ --custom-provider-id llama-cpp \ --custom-compatibility openai \ --custom-text-input \ --accept-risk \ --skip-health
Run OpenClaw
openclaw agent --local --agent main --message "Hello from Hugging Face"
LFM2.5-1.2B Italian
LiquidAI/LFM2.5-1.2B-Base model further continued-pretrained on Italian Wikipedia and then fine-tuned with Italian instructions (Alpaca format).
Training
- Continued Pretraining: ~1% of Italian Wikipedia (20231101.it dump), texts ≥ 600 characters
- Instruction Tuning: dataset
DanielSc4/alpaca-cleaned-italian - Technique: LoRA + Unsloth (merged in full precision / fp16)
- Original base model: multilingual, with significant improvement in Italian
Supported Languages
Improved in Italian.
Maintains performance in English and residual capability in the other languages of the base model.
This model follows the exact continued pretraining + instruction tuning recipe recently published by LiquidAI in their official cookbook: https://github.com/Liquid4All/cookbook/blob/main/finetuning/notebooks/cpt_translation_with_unsloth.ipynb I only replaced Korean with Italian by using: Italian Wikipedia (20231101.it dump) instead of Korean Wikipedia for continued pretraining DanielSc4/alpaca-cleaned-italian dataset instead of the Korean Alpaca version for instruction tuning All other steps (LoRA including embed_tokens & lm_head, separate embedding learning rate, Unsloth, fp16 merge) are identical to the original notebook.
- Downloads last month
- 17
5-bit