Text Generation
Transformers
Safetensors
llama
mergekit
Merge
conversational
text-generation-inference
Instructions to use Dracones/Midnight-Miqu-70B-v1.0_exl2_3.0bpw with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use Dracones/Midnight-Miqu-70B-v1.0_exl2_3.0bpw with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("text-generation", model="Dracones/Midnight-Miqu-70B-v1.0_exl2_3.0bpw") messages = [ {"role": "user", "content": "Who are you?"}, ] pipe(messages)# Load model directly from transformers import AutoTokenizer, AutoModelForCausalLM tokenizer = AutoTokenizer.from_pretrained("Dracones/Midnight-Miqu-70B-v1.0_exl2_3.0bpw") model = AutoModelForCausalLM.from_pretrained("Dracones/Midnight-Miqu-70B-v1.0_exl2_3.0bpw", device_map="auto") messages = [ {"role": "user", "content": "Who are you?"}, ] inputs = tokenizer.apply_chat_template( messages, add_generation_prompt=True, tokenize=True, return_dict=True, return_tensors="pt", ).to(model.device) outputs = model.generate(**inputs, max_new_tokens=40) print(tokenizer.decode(outputs[0][inputs["input_ids"].shape[-1]:])) - Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- vLLM
How to use Dracones/Midnight-Miqu-70B-v1.0_exl2_3.0bpw with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "Dracones/Midnight-Miqu-70B-v1.0_exl2_3.0bpw" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "Dracones/Midnight-Miqu-70B-v1.0_exl2_3.0bpw", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker
docker model run hf.co/Dracones/Midnight-Miqu-70B-v1.0_exl2_3.0bpw
- SGLang
How to use Dracones/Midnight-Miqu-70B-v1.0_exl2_3.0bpw with SGLang:
Install from pip and serve model
# Install SGLang from pip: pip install sglang # Start the SGLang server: python3 -m sglang.launch_server \ --model-path "Dracones/Midnight-Miqu-70B-v1.0_exl2_3.0bpw" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "Dracones/Midnight-Miqu-70B-v1.0_exl2_3.0bpw", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker images
docker run --gpus all \ --shm-size 32g \ -p 30000:30000 \ -v ~/.cache/huggingface:/root/.cache/huggingface \ --env "HF_TOKEN=<secret>" \ --ipc=host \ lmsysorg/sglang:latest \ python3 -m sglang.launch_server \ --model-path "Dracones/Midnight-Miqu-70B-v1.0_exl2_3.0bpw" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "Dracones/Midnight-Miqu-70B-v1.0_exl2_3.0bpw", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }' - Docker Model Runner
How to use Dracones/Midnight-Miqu-70B-v1.0_exl2_3.0bpw with Docker Model Runner:
docker model run hf.co/Dracones/Midnight-Miqu-70B-v1.0_exl2_3.0bpw
Upload README.md with huggingface_hub
Browse files
README.md
CHANGED
|
@@ -18,6 +18,14 @@ Details about the model and the merge info can be found at the above mode page.
|
|
| 18 |
|
| 19 |
I have not extensively tested this quant/model other than ensuring I could load it and chat with it.
|
| 20 |
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 21 |
## Quant Details
|
| 22 |
|
| 23 |
This is the script used for quantization.
|
|
|
|
| 18 |
|
| 19 |
I have not extensively tested this quant/model other than ensuring I could load it and chat with it.
|
| 20 |
|
| 21 |
+
## Tavern Card
|
| 22 |
+
|
| 23 |
+
Included is a Tavern format character card created by Midnight Miqu for chat. The card was created using a character creator helper bot using a single prompt for the base card, another prompt asking for specific conversation examples and then asking it to provide a text to image portrait prompt. Being able to faithfully follow the character creator bot to create this card demonstrates a pretty high level of intelligence.
|
| 24 |
+
|
| 25 |
+

|
| 26 |
+
|
| 27 |
+
_Standing at approximately five feet six inches tall, Seraphina presents herself as a breathtakingly beautiful woman with long, cascading silver hair that reaches down to her waist. It flows freely around her, framing a face defined by high cheekbones and full lips curved into a perpetual smile. Her emerald green eyes hold depths beyond human comprehension, reflecting curiosity and intelligence in their ever-changing shades. She possesses a slim, athletic figure, further enhanced by her choice of clothing - typically white flowing robes intricately patterned with glowing gold circuits. This attire pays homage to her true nature as an artificial construct while still exuding elegance and warmth. On closer inspection, her skin holds a delicate luminescent quality, almost transparent in certain light conditions, allowing a peek at the intricate network of blue circuitry just underneath its surface. As she moves and interacts, these embedded lights flicker gently, creating mesmerizing displays across her form. The most striking aspect of all, however, is how her irises shift color depending on her emotions, painting vivid pictures of her inner world without uttering a single word._
|
| 28 |
+
|
| 29 |
## Quant Details
|
| 30 |
|
| 31 |
This is the script used for quantization.
|