--- license: other license_name: lfm1.0 license_link: https://huggingface.co/LiquidAI/LFM2.5-350M/blob/main/LICENSE pipeline_tag: text-generation tags: - liquid - lfm2.5 - gguf - ollama - edge - conversational base_model: LiquidAI/LFM2.5-350M --- # LiquidAI LFM2.5-350M (Instruct) - GGUF (Q4_K_M) This repository provides the quantized **Q4_K_M GGUF** weights for **[LiquidAI/LFM2.5-350M](https://huggingface.co/LiquidAI/LFM2.5-350M)**, configured for direct 1-click execution in **Ollama**, **llama.cpp**, and local edge devices. LFM2.5-350M is a hybrid architecture developed by Liquid AI combining double-gated short convolutions with structured attention for near-linear computational scaling and low memory footprint. --- ## ⚡ Direct Ollama Run (1-Line Command) You can run this model directly via Ollama without manually downloading any files: ```bash ollama run hf.co/jamesatron1512/LFM2.5-350M-GGUF ``` Or specify the quantization tag explicitly: ```bash ollama run hf.co/jamesatron1512/LFM2.5-350M-GGUF:Q4_K_M ``` --- ## 🚀 Model Details - **Parameters**: 350 Million - **Precision**: Q4_K_M (Quantized 4-bit) - **File Size**: ~219 MB - **Context Length**: Up to 128k tokens (default 4096 in Modelfile) - **Chat Template**: ChatML format (`<|im_start|>user ... <|im_end|>`) - **System Prompt**: Supported via template, default is left clean to prevent fixation on small parameter counts. --- ## 💻 Python API Usage via Ollama ```python import requests response = requests.post( "http://localhost:11434/api/generate", json={ "model": "hf.co/jamesatron1512/LFM2.5-350M-GGUF", "prompt": "Explain quantum computing in two sentences.", "stream": False, "options": { "temperature": 0.7, "top_p": 0.9, "num_predict": 128 } } ) print(response.json()["response"]) ```