YAML Metadata Warning:empty or missing yaml metadata in repo card

Check out the documentation for more information.

Mistral-7B-Qwen-Distilled

This model is a fine-tuned version of mistralai/Mistral-7B-v0.1, distilled from Qwen/Qwen3-Next-80B-A3B-Instruct using LoRA and SFTTrainer.

Article: https://medium.com/ai-simplified-in-plain-english/knowledge-distillation-of-qwen3-next-80b-a3b-instruct-into-mistral-7b-v0-1-2328900a67b3

Training Details

  • Teacher Model: Qwen/Qwen3-Next-80B-A3B-Instruct
  • Dataset: Synthetic dataset of 20 samples generated by the teacher.
  • LoRA Config: r=16, alpha=32, dropout=0.05
  • Training Hyperparams: 3 epochs, learning rate 0.0002, batch size 2

Usage

from transformers import AutoTokenizer, AutoModelForCausalLM
from peft import PeftModel

model = AutoModelForCausalLM.from_pretrained("mistralai/Mistral-7B-v0.1", device_map="auto")
tokenizer = AutoTokenizer.from_pretrained("mistralai/Mistral-7B-v0.1")
model = PeftModel.from_pretrained(model, "frankmorales2020/mistral-7b-qwen-Next-80B-A3B-Instruct-distilled")

Built with ❤️ using Hugging Face Transformers.

Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support