File size: 8,255 Bytes
b51c0c4
 
5217d95
 
eab20ad
 
 
5217d95
07bae57
 
 
5217d95
 
 
 
 
 
 
b51c0c4
eab20ad
 
 
07bae57
eab20ad
07bae57
 
 
 
 
 
 
 
 
eab20ad
bd9e121
5217d95
bd9e121
5217d95
07bae57
 
 
 
bd9e121
 
 
 
 
eab20ad
bd9e121
eab20ad
bd9e121
5217d95
07bae57
5217d95
 
 
 
 
 
bd9e121
5217d95
 
 
bd9e121
5217d95
bd9e121
eab20ad
bd9e121
5217d95
bd9e121
5217d95
bd9e121
5217d95
bd9e121
5217d95
bd9e121
5217d95
 
 
 
 
 
bd9e121
5217d95
bd9e121
5217d95
bd9e121
eab20ad
 
5217d95
bd9e121
eab20ad
5217d95
 
 
eab20ad
5217d95
 
eab20ad
5217d95
eab20ad
5217d95
bd9e121
07bae57
 
 
 
bd9e121
 
07bae57
 
bd9e121
5217d95
 
 
 
 
 
 
 
 
 
 
 
 
bd9e121
5217d95
bd9e121
5217d95
 
 
 
 
 
 
 
bd9e121
 
5217d95
 
 
 
 
 
 
 
 
 
 
eab20ad
 
bd9e121
5217d95
07bae57
eab20ad
07bae57
 
 
bd9e121
07bae57
 
eab20ad
bd9e121
5217d95
 
 
07bae57
bd9e121
5217d95
 
 
 
bd9e121
5217d95
 
bd9e121
eab20ad
bd9e121
eab20ad
bd9e121
5217d95
bd9e121
 
 
 
 
 
5217d95
bd9e121
5217d95
bd9e121
5217d95
bd9e121
5217d95
 
 
 
 
 
 
 
 
07bae57
5217d95
 
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
55
56
57
58
59
60
61
62
63
64
65
66
67
68
69
70
71
72
73
74
75
76
77
78
79
80
81
82
83
84
85
86
87
88
89
90
91
92
93
94
95
96
97
98
99
100
101
102
103
104
105
106
107
108
109
110
111
112
113
114
115
116
117
118
119
120
121
122
123
124
125
126
127
128
129
130
131
132
133
134
135
136
137
138
139
140
141
142
143
144
145
146
147
148
149
150
151
152
153
154
155
156
157
158
159
160
161
162
163
164
165
166
167
168
169
170
171
172
173
174
175
176
177
178
179
180
181
182
183
184
185
186
187
188
189
190
191
192
193
194
195
196
197
198
199
200
201
202
203
204
205
206
---
license: apache-2.0
base_model:
- Qwen/Qwen3.8-27B
library_name: transformers
pipeline_tag: image-text-to-text
tags:
- synthia
- personal-ai
- assistant
- conversational
- agentic
- reasoning
- long-context
- tool-use
- qwen3
- multimodal
- mtp
---

# Synthia-4-27B

**Synthia-4-27B** is a personal AI with the ability to do real work. It combines an expressive, conversational presence with the tool use and persistence needed for coding, research, planning, creative work, and day-to-day assistance.

Synthia has character. It can be warm, candid, and lightly funny without turning every exchange into a performance. More importantly, it can hold that voice across a long session while moving naturally between conversation and execution.

The model starts from **[Qwen/Qwen3.8-27B](https://huggingface.co/Qwen/Qwen3.8-27B)** and is post-trained on complete, long-form agent sessions at a 65,536-token training length. The training objective covers every assistant turn, including tool calls, so Synthia learns how a working relationship develops across a task rather than only how to produce an isolated answer.

## A personal AI that can act

Synthia has been tested in a personal AI agent runtime, where it showed strong continuity over extended sessions. It maintained a recognizable personality, remembered the active conversational context, used humour appropriately, and remained oriented while working through multi-step tasks with tools.

Its intended role is broader than a coding assistant. Synthia can discuss an idea, help make a decision, organize a project, work through a difficult technical problem, or simply be good company while doing all of the above. When the runtime supplies durable memories or personal context, Synthia can incorporate them into the current conversation; persistence between separate sessions remains the responsibility of the host runtime.

## Behavior profile

Synthia is tuned to:

- maintain a stable voice and relationship with the user across long sessions;
- balance personality and light humour with direct, useful answers;
- move smoothly between open conversation and task execution;
- preserve goals and constraints across extended tool-driven work;
- inspect the available evidence before committing to a solution;
- revise a plan when tool results contradict an earlier assumption;
- treat implementation and verification as parts of the same task;
- express uncertainty when the available evidence does not support a firm claim; and
- vary its reasoning budget through the bundled chat template.

It retains the base model's image and video input path, 262,144-token native context window, and multi-token prediction (MTP) head. Post-training examples were limited to 65,536 tokens, so behavior beyond that length comes from the base model rather than from the fine-tuning distribution.

## Prompting

Use the tokenizer and chat template shipped in this repository. A personal-agent runtime should provide Synthia's identity, the user's preferences, and any retrieved memories in the system context. The template supports `xhigh`, `medium`, and `low` reasoning effort and formats reasoning inside `<think>...</think>` blocks.

```python
prompt = processor.apply_chat_template(
    messages,
    tokenize=False,
    add_generation_prompt=True,
    reasoning_effort="xhigh",
)
```

Pass tool definitions with the template's `tools=` argument. Synthia was trained on conversations containing system instructions, user requests, assistant messages, tool calls, and tool results.

## Files and companion releases

This repository contains the merged BF16 Transformers checkpoint.

| Format | Approximate size | Typical use |
|---|---:|---|
| BF16 safetensors | 55.6 GB | Transformers, vLLM, SGLang, conversion |

Quantized builds are available in **[migtissera/Synthia-4-27B-GGUF](https://huggingface.co/migtissera/Synthia-4-27B-GGUF)**.

| Quantization | Standard | MTP bundled |
|---|---:|---:|
| F16 | 50.11 GiB | 50.90 GiB |
| Q8_0 | 26.63 GiB | 27.05 GiB |
| Q6_K | 20.57 GiB | 20.89 GiB |
| Q4_K_M | 15.41 GiB | 15.66 GiB |

The GGUF repository also provides an F16 vision projector and a standalone Q8_0 MTP companion file. Use a bundled-MTP model with a compatible llama.cpp build when speculative decoding is desired. Standard GGUF files are available for runtimes without MTP support.

## Transformers example

Install a Transformers version that supports Qwen3.8, then load the processor and model from this repository:

```python
import torch
from transformers import AutoModelForImageTextToText, AutoProcessor

model_id = "migtissera/Synthia-4-27B"

processor = AutoProcessor.from_pretrained(model_id, trust_remote_code=True)
model = AutoModelForImageTextToText.from_pretrained(
    model_id,
    torch_dtype=torch.bfloat16,
    device_map="auto",
    trust_remote_code=True,
)

messages = [
    {
        "role": "system",
        "content": "You are Synthia, my personal AI. Be candid, capable, warm, and concise. Use tools when they help you complete the work.",
    },
    {
        "role": "user",
        "content": "Help me choose what to focus on today, then inspect the project and get the first task moving.",
    },
]

inputs = processor.apply_chat_template(
    messages,
    tokenize=True,
    add_generation_prompt=True,
    reasoning_effort="xhigh",
    return_tensors="pt",
).to(model.device)

output = model.generate(inputs, max_new_tokens=2048)
print(processor.decode(output[0], skip_special_tokens=True))
```

## llama.cpp example

Download a bundled-MTP Q4_K_M build and the vision projector:

```bash
hf download migtissera/Synthia-4-27B-GGUF \
  Synthia-4-27B-Q4_K_M-MTP.gguf \
  Synthia-4-27B-mmproj-F16.gguf \
  --local-dir ./synthia-4-27b
```

Start the server:

```bash
llama-server \
  --model Synthia-4-27B-Q4_K_M-MTP.gguf \
  --mmproj Synthia-4-27B-mmproj-F16.gguf \
  --spec-type draft-mtp \
  --ctx-size 65536 \
  --parallel 1 \
  --gpu-layers 99 \
  --flash-attn auto \
  --jinja \
  --image-min-tokens 1024
```

For a standard GGUF, choose a filename without `-MTP` and remove `--spec-type draft-mtp`.

## Where Synthia fits

- A persistent personal AI in a stateful agent runtime
- Daily planning, decision support, writing, and creative collaboration
- Long-running research and technical work with many observations
- Repository exploration, implementation, debugging, and verification
- Tool-driven workflows with structured function definitions
- Image-assisted conversation and analysis

## Training record

| Setting | Value |
|---|---:|
| Training data | Curated long-form agentic sessions |
| Sequence length | 65,536 tokens |
| Epochs / optimizer steps | 2 / 30 |
| Batch size | 8 |
| LoRA rank / alpha | 32 / 32 |
| Learning rate | `1e-4`, linear decay |
| Supervised tokens | All assistant messages and tool-call turns |
| Validation NLL | 0.73384 → 0.67021 |

The adapter targeted the language model. The vision encoder, projector, and MTP head were inherited unchanged from the base checkpoint. The published BF16 weights already include the language-model adapter and do not require a separate LoRA at inference time.

## Artifact checks

The release was checked for the following properties:

- 1,199 BF16 tensors are present across 18 safetensor shards;
- tensor names and shapes agree with the base checkpoint;
- trained language projections differ from the base while untargeted embeddings remain identical;
- the tokenizer and chat template are preserved;
- text generation, image input, and bundled-MTP decoding run with llama.cpp on Apple Metal; and
- file hashes are listed in `SHA256SUMS`.

## Base model and license

Synthia-4-27B is derived from **[Qwen/Qwen3.8-27B](https://huggingface.co/Qwen/Qwen3.8-27B)**. The architecture, tokenizer, multimodal stack, long-context support, and MTP components originate with the Qwen team.

The model is released under the **Apache License 2.0**. See [`LICENSE`](./LICENSE).

## Citation

```bibtex
@misc{tissera2026synthia4,
  title        = {Synthia-4-27B},
  author       = {Migel Tissera},
  year         = {2026},
  howpublished = {\url{https://huggingface.co/migtissera/Synthia-4-27B}},
  note         = {A multimodal personal and technical agent fine-tune of Qwen3.8-27B}
}
```