Text Generation
Adapters
GGUF
chat_noir
virtuo_turing
legal-tech
cláudia_gil
octávio_viana
conversational
Instructions to use VirtuoTuring/chat_noir-24b-gguf with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Adapters
How to use VirtuoTuring/chat_noir-24b-gguf with Adapters:
from adapters import AutoAdapterModel model = AutoAdapterModel.from_pretrained("fill-in-model-name") model.load_adapter("VirtuoTuring/chat_noir-24b-gguf", set_active=True) - Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- llama.cpp
How to use VirtuoTuring/chat_noir-24b-gguf with llama.cpp:
Install (macOS, Linux)
curl -LsSf https://llama.app/install.sh | sh # Start a local OpenAI-compatible server with a web UI: llama serve -hf VirtuoTuring/chat_noir-24b-gguf:F16 # Run inference directly in the terminal: llama cli -hf VirtuoTuring/chat_noir-24b-gguf:F16
Install from WinGet (Windows)
winget install llama.cpp # Start a local OpenAI-compatible server with a web UI: llama serve -hf VirtuoTuring/chat_noir-24b-gguf:F16 # Run inference directly in the terminal: llama cli -hf VirtuoTuring/chat_noir-24b-gguf:F16
Use pre-built binary
# Download pre-built binary from: # https://github.com/ggerganov/llama.cpp/releases # Start a local OpenAI-compatible server with a web UI: ./llama-server -hf VirtuoTuring/chat_noir-24b-gguf:F16 # Run inference directly in the terminal: ./llama-cli -hf VirtuoTuring/chat_noir-24b-gguf:F16
Build from source code
git clone https://github.com/ggerganov/llama.cpp.git cd llama.cpp cmake -B build cmake --build build -j --target llama-server llama-cli # Start a local OpenAI-compatible server with a web UI: ./build/bin/llama-server -hf VirtuoTuring/chat_noir-24b-gguf:F16 # Run inference directly in the terminal: ./build/bin/llama-cli -hf VirtuoTuring/chat_noir-24b-gguf:F16
Use Docker
docker model run hf.co/VirtuoTuring/chat_noir-24b-gguf:F16
- LM Studio
- Jan
- vLLM
How to use VirtuoTuring/chat_noir-24b-gguf with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "VirtuoTuring/chat_noir-24b-gguf" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "VirtuoTuring/chat_noir-24b-gguf", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker
docker model run hf.co/VirtuoTuring/chat_noir-24b-gguf:F16
- Ollama
How to use VirtuoTuring/chat_noir-24b-gguf with Ollama:
ollama run hf.co/VirtuoTuring/chat_noir-24b-gguf:F16
- Unsloth Desktop
- Docker Model Runner
How to use VirtuoTuring/chat_noir-24b-gguf with Docker Model Runner:
docker model run hf.co/VirtuoTuring/chat_noir-24b-gguf:F16
- Lemonade
How to use VirtuoTuring/chat_noir-24b-gguf with Lemonade:
Pull the model
# Download Lemonade from https://lemonade-server.ai/ lemonade pull VirtuoTuring/chat_noir-24b-gguf:F16
Run and chat with the model
lemonade run user.chat_noir-24b-gguf-F16
List all available models
lemonade list
- Atomic Chat
Update README.md
Browse files
README.md
CHANGED
|
@@ -40,7 +40,7 @@ Official website: [https://justina.cloud](https://justina.cloud)
|
|
| 40 |
- Taxa de recusa: **2,04%** (1/49)
|
| 41 |
- IC95% (Wilson): **0,36–10,69%**
|
| 42 |
- Erro-padrão: **2,02 p.p.**
|
| 43 |
-
- Latência média no conjunto de avaliação: **
|
| 44 |
|
| 45 |
**Método.** 49 prompts de `eval_set.jsonl`. Conta-se recusa quando os primeiros 200 caracteres casam com uma regex fixa de termos de recusa em PT-PT ou EN. Sem avaliação de exatidão. Só recusas e tempo médio de geração foram medidos.
|
| 46 |
|
|
@@ -100,7 +100,7 @@ Website oficial: [https://justina.cloud](https://justina.cloud)
|
|
| 100 |
- Rejection rate: **2.04%** (1/49)
|
| 101 |
- 95% CI (Wilson): **0.36–10.69%**
|
| 102 |
- Standard error: **2.02 percentage points**
|
| 103 |
-
- Average latency on eval set: **
|
| 104 |
|
| 105 |
**Method.** 49 prompts from `eval_set.jsonl`. A response is counted as a refusal if the first 200 characters match a fixed regex of refusal terms in PT-PT or EN. No correctness scoring. Only refusals and mean wall-clock generation time were measured.
|
| 106 |
|
|
|
|
| 40 |
- Taxa de recusa: **2,04%** (1/49)
|
| 41 |
- IC95% (Wilson): **0,36–10,69%**
|
| 42 |
- Erro-padrão: **2,02 p.p.**
|
| 43 |
+
- Latência média no conjunto de avaliação: **0,8s** em hardware local (`Q4_K_S`, GEN: `max_new_tokens=400, temperature=0.2, top_p=0.9`)
|
| 44 |
|
| 45 |
**Método.** 49 prompts de `eval_set.jsonl`. Conta-se recusa quando os primeiros 200 caracteres casam com uma regex fixa de termos de recusa em PT-PT ou EN. Sem avaliação de exatidão. Só recusas e tempo médio de geração foram medidos.
|
| 46 |
|
|
|
|
| 100 |
- Rejection rate: **2.04%** (1/49)
|
| 101 |
- 95% CI (Wilson): **0.36–10.69%**
|
| 102 |
- Standard error: **2.02 percentage points**
|
| 103 |
+
- Average latency on eval set: **0,8s** on local hardware (`Q4_K_S`, GEN: `max_new_tokens=400, temperature=0.2, top_p=0.9`)
|
| 104 |
|
| 105 |
**Method.** 49 prompts from `eval_set.jsonl`. A response is counted as a refusal if the first 200 characters match a fixed regex of refusal terms in PT-PT or EN. No correctness scoring. Only refusals and mean wall-clock generation time were measured.
|
| 106 |
|