Instructions to use flammenai/flammen15X-mistral-7B with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use flammenai/flammen15X-mistral-7B with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("text-generation", model="flammenai/flammen15X-mistral-7B")# Load model directly from transformers import AutoTokenizer, AutoModelForCausalLM tokenizer = AutoTokenizer.from_pretrained("flammenai/flammen15X-mistral-7B") model = AutoModelForCausalLM.from_pretrained("flammenai/flammen15X-mistral-7B", device_map="auto") - Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- vLLM
How to use flammenai/flammen15X-mistral-7B with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "flammenai/flammen15X-mistral-7B" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "flammenai/flammen15X-mistral-7B", "prompt": "Once upon a time,", "max_tokens": 512, "temperature": 0.5 }'Use Docker
docker model run hf.co/flammenai/flammen15X-mistral-7B
- SGLang
How to use flammenai/flammen15X-mistral-7B with SGLang:
Install from pip and serve model
# Install SGLang from pip: pip install sglang # Start the SGLang server: python3 -m sglang.launch_server \ --model-path "flammenai/flammen15X-mistral-7B" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "flammenai/flammen15X-mistral-7B", "prompt": "Once upon a time,", "max_tokens": 512, "temperature": 0.5 }'Use Docker images
docker run --gpus all \ --shm-size 32g \ -p 30000:30000 \ -v ~/.cache/huggingface:/root/.cache/huggingface \ --env "HF_TOKEN=<secret>" \ --ipc=host \ lmsysorg/sglang:latest \ python3 -m sglang.launch_server \ --model-path "flammenai/flammen15X-mistral-7B" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "flammenai/flammen15X-mistral-7B", "prompt": "Once upon a time,", "max_tokens": 512, "temperature": 0.5 }' - Docker Model Runner
How to use flammenai/flammen15X-mistral-7B with Docker Model Runner:
docker model run hf.co/flammenai/flammen15X-mistral-7B
| library_name: transformers | |
| license: apache-2.0 | |
| base_model: | |
| - nbeerbower/flammen15-gutenberg-DPO-v1-7B | |
| datasets: | |
| - chargoddard/chai-dpo | |
|  | |
| # flammen15X-mistral-7B | |
| A Mistral 7B LLM built from merging pretrained models and finetuning on [Jon Durbin](https://huggingface.co/jondurbin)'s [Gutenberg DPO set](https://huggingface.co/datasets/jondurbin/gutenberg-dpo-v0.1) and [Charles Goddard](https://huggingface.co/chargoddard)'s [Chai DPO set](https://huggingface.co/datasets/chargoddard/chai-dpo). | |
| Flammen specializes in exceptional character roleplay, creative writing, and general intelligence | |
| ### Method | |
| Finetuned using an A100 on Google Colab. 🙏 | |
| [Fine-tune a Mistral-7b model with Direct Preference Optimization](https://towardsdatascience.com/fine-tune-a-mistral-7b-model-with-direct-preference-optimization-708042745aac) - [Maxime Labonne](https://huggingface.co/mlabonne) | |
| ### Configuration | |
| LoRA, model, and training settings: | |
| ```python | |
| # LoRA configuration | |
| peft_config = LoraConfig( | |
| r=16, | |
| lora_alpha=16, | |
| lora_dropout=0.05, | |
| bias="none", | |
| task_type="CAUSAL_LM", | |
| target_modules=['k_proj', 'gate_proj', 'v_proj', 'up_proj', 'q_proj', 'o_proj', 'down_proj'] | |
| ) | |
| # Model to fine-tune | |
| model = AutoModelForCausalLM.from_pretrained( | |
| model_name, | |
| torch_dtype=torch.bfloat16, | |
| load_in_4bit=True | |
| ) | |
| model.config.use_cache = False | |
| # Reference model | |
| ref_model = AutoModelForCausalLM.from_pretrained( | |
| model_name, | |
| torch_dtype=torch.bfloat16, | |
| load_in_4bit=True | |
| ) | |
| # Training arguments | |
| training_args = TrainingArguments( | |
| per_device_train_batch_size=2, | |
| gradient_accumulation_steps=2, | |
| gradient_checkpointing=True, | |
| learning_rate=2e-5, | |
| lr_scheduler_type="cosine", | |
| max_steps=200, | |
| save_strategy="no", | |
| logging_steps=1, | |
| output_dir=new_model, | |
| optim="paged_adamw_32bit", | |
| warmup_steps=100, | |
| bf16=True, | |
| report_to="wandb", | |
| ) | |
| # Create DPO trainer | |
| dpo_trainer = DPOTrainer( | |
| model, | |
| ref_model, | |
| args=training_args, | |
| train_dataset=dataset, | |
| tokenizer=tokenizer, | |
| peft_config=peft_config, | |
| beta=0.1, | |
| max_prompt_length=1024, | |
| max_length=1536, | |
| force_use_ref_model=True | |
| ) | |
| # Fine-tune model with DPO | |
| dpo_trainer.train() | |
| ``` |