YAML Metadata Warning:empty or missing yaml metadata in repo card
Check out the documentation for more information.
Mistral-7B-Qwen-Distilled
This model is a fine-tuned version of mistralai/Mistral-7B-v0.1, distilled from Qwen/Qwen3-Next-80B-A3B-Instruct using LoRA and SFTTrainer.
Training Details
- Teacher Model: Qwen/Qwen3-Next-80B-A3B-Instruct
- Dataset: Synthetic dataset of 20 samples generated by the teacher.
- LoRA Config: r=16, alpha=32, dropout=0.05
- Training Hyperparams: 3 epochs, learning rate 0.0002, batch size 2
Usage
from transformers import AutoTokenizer, AutoModelForCausalLM
from peft import PeftModel
model = AutoModelForCausalLM.from_pretrained("mistralai/Mistral-7B-v0.1", device_map="auto")
tokenizer = AutoTokenizer.from_pretrained("mistralai/Mistral-7B-v0.1")
model = PeftModel.from_pretrained(model, "frankmorales2020/mistral-7b-qwen-Next-80B-A3B-Instruct-distilled")
Built with ❤️ using Hugging Face Transformers.
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support