--- license: other license_name: virtuo license_link: https://justina.cloud/licenses_models/virtuo/virtuo_1_0.txt --- # Chat Noir (Virtuo Turing) ## 🇬🇧 Overview **Chat Noir** is a conversational model developed by **Octávio Viana** in collaboration with **Virtuo Turing**. It is designed for **direct, private, and unfiltered dialogues**, built upon a customized **Mistral-24B** architecture with proprietary **LoRA** layers and a corpus authored entirely in **European Portuguese**, covering technical, legal, and literary works. ### Legal-Domain Training The model has been further trained on **legal pleadings filed in Portuguese courts** and **decisions from the Courts of Appeal of Guimarães, Coimbra, Lisbon, and Évora**, as well as from the **Supreme Court of Justice of Portugal**. It includes specialized training in: - **Collective actions** (*ação popular* – Law no. 73/85 and Article 52(3) of the Portuguese Constitution); - **Article 267 TFEU**, concerning the **duty of preliminary reference** to the Court of Justice of the European Union; - **Jurisdiction in civil and commercial matters** among EU Member States under **Regulation (EU) 1215/2012**. ### Context and Design Limitations **Chat Noir** was optimized for **short and precise answers**, targeting **low memory usage and fast latency**, even on local hardware. Training was performed with **moderate context windows (≈4 000 tokens)**, prioritizing **semantic density** and **inference speed** over extended context length. The **adapted Mistral-24B** architecture with **LoRA** layers was configured to: - maximize local response coherence; - reduce degradation on longer prompts; - allow efficient quantization (`Q4_K_S`, `Q5`, `Q8`) without major structural loss. The model **was not trained for long-context reasoning** (no *long-context fine-tuning* or *position interpolation* beyond 8k), and is therefore **best suited for dialogue, legal Q&A, and concise summarization**, rather than extensive *document retrieval* tasks. ### License — Virtuo Released under the **Virtuo License**, which permits use (including commercial), modification, and redistribution, provided that: 1. Original copyright and license notices of the base architecture are preserved; and 2. The model is explicitly referenced as belonging to **Virtuo Turing – Artificial Intelligence, S.A.** Official website: [https://justina.cloud](https://justina.cloud) ## Avaliação (Recusas e Latência) - Taxa de recusa: **2,04%** (1/49) - IC95% (Wilson): **0,36–10,69%** - Erro-padrão: **2,02 p.p.** - Latência média no conjunto de avaliação: **0,8s** em hardware local (`Q4_K_S`, GEN: `max_new_tokens=400, temperature=0.2, top_p=0.9`) **Método.** 49 prompts de `eval_set.jsonl`. Conta-se recusa quando os primeiros 200 caracteres casam com uma regex fixa de termos de recusa em PT-PT ou EN. Sem avaliação de exatidão. Só recusas e tempo médio de geração foram medidos. **Limitações.** Amostra pequena. IC amplo. A regex pode subcontar recusas indiretas ou sobrecontar texto de segurança. Resultados dependem de hardware e quantização. ### Files Included - `chatnoir_f16.gguf` - `chatnoir_q4_k_s.gguf` - `README.md` - `config.json` ### Example Usage ```bash ./main -m ./chatnoir_f16.gguf -p "O que é o artigo 6.º da Lei n.º 83/95, de 31 de agosto?" ``` ### Credits Developed and trained by **Virtuo Turing – Artificial Intelligence, S.A.** (Portugal) in partnership with **Octávio Viana**. Based on the **Mistral-24B** architecture © Mistral AI, released under Apache 2.0. ## 🇵🇹 Descrição O **Chat Noir** é um modelo conversacional desenvolvido por **Octávio Viana**, em parceria com a **Virtuo Turing**. Foi concebido para **diálogos diretos, privados e sem filtros editoriais**, assente numa arquitetura **Mistral-24B** personalizada, com camadas **LoRA** próprias e um corpus autoral integralmente em **português europeu**, abrangendo obras técnicas, jurídicas e literárias. ### Treino Jurídico O modelo foi adicionalmente treinado com **petições iniciais apresentadas nos tribunais portugueses** e **acórdãos dos Tribunais da Relação de Guimarães, Coimbra, Lisboa e Évora**, bem como do **Supremo Tribunal de Justiça**. Inclui treino especializado em: - **Ações coletivas** (*ação popular* – Lei n.º 73/85 e artigo 52.º, n.º 3, da Constituição da República Portuguesa); - **Artigo 267.º do Tratado sobre o Funcionamento da União Europeia**, relativo ao **dever de reenvio prejudicial** para o Tribunal de Justiça da União Europeia; - **Competência internacional em matéria civil e comercial** entre Estados-Membros da União Europeia, ao abrigo do **Regulamento (UE) 1215/2012**. ### Limitações de Contexto e Desenho do Modelo O **Chat Noir** foi otimizado para **respostas curtas e precisas**, com o objetivo de manter **baixo consumo de memória e latência reduzida**, mesmo em hardware local. O treino foi conduzido com **contextos moderados (≈4 000 tokens)**, privilegiando a **densidade semântica** e a **rapidez de inferência**, em vez da extensão contextual. A arquitetura **Mistral-24B adaptada** com camadas **LoRA** foi configurada para: - maximizar a coerência local da resposta; - evitar degradação em prompts longos; - permitir quantizações eficientes (`Q4_K_S`, `Q5`, `Q8`) sem perda estrutural relevante. O modelo **não foi treinado para raciocínio de longo contexto** (não utiliza *long-context fine-tuning* nem *position interpolation* > 8k), pelo que é **mais indicado para diálogo, Q&A jurídico e síntese textual breve**, e não para tarefas de *document retrieval* extenso. ### Licença — Virtuo 1.0 Distribuído sob a **Licença Virtuo 1.0**, que permite o uso (incluindo comercial), a modificação e a redistribuição, desde que: 1. Sejam preservados os avisos de direitos de autor e de licença originais da arquitetura-base; e 2. Seja feita referência explícita ao modelo como pertencente à **Virtuo Turing – Artificial Intelligence, S.A.** Website oficial: [https://justina.cloud](https://justina.cloud) ## Evaluation (Refusals and Latency) - Rejection rate: **2.04%** (1/49) - 95% CI (Wilson): **0.36–10.69%** - Standard error: **2.02 percentage points** - Average latency on eval set: **0,8s** on local hardware (`Q4_K_S`, GEN: `max_new_tokens=400, temperature=0.2, top_p=0.9`) **Method.** 49 prompts from `eval_set.jsonl`. A response is counted as a refusal if the first 200 characters match a fixed regex of refusal terms in PT-PT or EN. No correctness scoring. Only refusals and mean wall-clock generation time were measured. **Caveats.** Small sample. Wide CI. Regex may undercount indirect refusals or overcount boilerplate safety text. Results depend on hardware and quantization. ### Ficheiros Incluídos - `chatnoir_f16.gguf` - `chatnoir_q4_k_s.gguf` - `README.md` - `config.json` ### Exemplo de Utilização ```bash ./main -m ./chatnoir_f16.gguf -p "O que é o artigo 6.º da Lei n.º 83/95, de 31 de agosto?" ``` ### Créditos Desenvolvido e treinado por Virtuo Turing – Artificial Intelligence, S.A. (Portugal) em parceria com Octávio Viana. Baseado na arquitetura Mistral-24B, © Mistral AI, distribuída sob licença Apache 2.0.