How to use from the
Use from the
Transformers library
# Use a pipeline as a high-level helper
from transformers import pipeline

pipe = pipeline("text-generation", model="mlx-community/Kimi-K2.5-3bit", trust_remote_code=True)
messages = [
    {"role": "user", "content": "Who are you?"},
]
pipe(messages)
# Load model directly
from transformers import AutoProcessor, AutoModel

processor = AutoProcessor.from_pretrained("mlx-community/Kimi-K2.5-3bit", trust_remote_code=True)
model = AutoModel.from_pretrained("mlx-community/Kimi-K2.5-3bit", trust_remote_code=True, device_map="auto")
messages = [
    {"role": "user", "content": "Who are you?"},
]
inputs = processor.apply_chat_template(
	messages,
	add_generation_prompt=True,
	tokenize=True,
	return_dict=True,
	return_tensors="pt",
).to(model.device)

outputs = model.generate(**inputs, max_new_tokens=40)
print(processor.decode(outputs[0][inputs["input_ids"].shape[-1]:]))
Quick Links

mlx-community/Kimi-K2.5-3bit

This model mlx-community/Kimi-K2.5-3bit was converted to MLX format from moonshotai/Kimi-K2.5 using mlx-lm version 0.30.6.

Usage

Please see the model card of the original model for example code. Note that this quant is for text-only usage.

Remember to allow remote code for tokenizer usage.

Downloads last month
177
Safetensors
Model size
1T params
Tensor type
BF16
U32
F32
MLX
Hardware compatibility
Log In to add your hardware

4-bit

Inference Providers NEW
This model isn't deployed by any Inference Provider. 馃檵 Ask for provider support

Model tree for mlx-community/Kimi-K2.5-3bit

Quantized
(38)
this model