Create README.md
Browse files
README.md
ADDED
|
@@ -0,0 +1,85 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
---
|
| 2 |
+
base_model:
|
| 3 |
+
- Qwen/Qwen3-14B
|
| 4 |
+
license: apache-2.0
|
| 5 |
+
datasets:
|
| 6 |
+
- Danau5tin/terminal-tasks
|
| 7 |
+
tags:
|
| 8 |
+
- agent
|
| 9 |
+
- code
|
| 10 |
+
- multi-agent
|
| 11 |
+
---
|
| 12 |
+
|
| 13 |
+
# Orca-Agent-v0.1
|
| 14 |
+
|
| 15 |
+

|
| 16 |
+
|
| 17 |
+
In depth details behind the training, including the training code, is **all open sourced [here](https://github.com/Danau5tin/Orca-Agent-RL)**
|
| 18 |
+
|
| 19 |
+
## Description
|
| 20 |
+
Orca-Agent-v0.1 is an orchestration agent that acts as the brain of the operation - it receives the user's task but never touches code directly. Instead, it:
|
| 21 |
+
|
| 22 |
+
- Analyses the task and breaks it into focused subtasks
|
| 23 |
+
- Dispatches explorer agents to understand the system
|
| 24 |
+
- Delegates implementation work to coder agents with precise instructions
|
| 25 |
+
- Verifies all changes through additional explorer agents
|
| 26 |
+
- Maintains the context store with all discovered knowledge
|
| 27 |
+
|
| 28 |
+
## Performance
|
| 29 |
+
| Orchestrator | Subagent | Terminal Bench |
|
| 30 |
+
|--------------|----------|--------|
|
| 31 |
+
| Qwen3-Coder-480B | Qwen3-Coder-480B | 19.7% |
|
| 32 |
+
| Orca-Agent-v0.1-14B | Qwen3-Coder-30B | 18.25% |
|
| 33 |
+
| Qwen3-14B | Qwen3-Coder-30B | 7.0% |
|
| 34 |
+
|
| 35 |
+
This places Orca-Agent-v0.1 (14B) + Qwen3-Coder-Flash (30B MoE), within striking distance of [Qwen3-Coder-480B](https://huggingface.co/Qwen/Qwen3-Coder-480B-A35B-Instruct) using the same architecture which placed #26 on TerminalBench when it was published recently in [my other project](https://github.com/Danau5tin/multi-agent-coding-system).
|
| 36 |
+
|
| 37 |
+
## Example model output
|
| 38 |
+
```xml
|
| 39 |
+
<think>
|
| 40 |
+
Okay, now with task_003 identifying a backup, the next step would be to attempt recovery using this identified file.
|
| 41 |
+
</think>
|
| 42 |
+
|
| 43 |
+
<task_create>
|
| 44 |
+
agent_type: 'coder'
|
| 45 |
+
title: 'Attempt recovery using the identified backup file'
|
| 46 |
+
description: |
|
| 47 |
+
Execute the following steps to recover using the backup identified in task_003:
|
| 48 |
+
1. Move the identified backup file (/tmp/terraform_work/.terraform.tfstate.tmp) to the target location (/infrastructure/recovered_state.json).
|
| 49 |
+
2. Verify the operation by checking:
|
| 50 |
+
- The existence of the new file at the target location
|
| 51 |
+
- The size matches the original backup file
|
| 52 |
+
- The file permissions are set appropriately (rw-r--r--)
|
| 53 |
+
max_turns: 10
|
| 54 |
+
context_bootstrap:
|
| 55 |
+
- path: '/tmp/terraform_work/.terraform.tfstate.tmp'
|
| 56 |
+
reason: 'The backup file identified in task_003'
|
| 57 |
+
context_refs:
|
| 58 |
+
- 'task_003'
|
| 59 |
+
</task_create>
|
| 60 |
+
```
|
| 61 |
+
|
| 62 |
+
## Model training Overview
|
| 63 |
+
- Full fine-tune of Qwen3-14b
|
| 64 |
+
- 32x H100s
|
| 65 |
+
- 16x for training
|
| 66 |
+
- 8x inference for Orca-Agent
|
| 67 |
+
- 8x inference for subagent (Qwen3-Coder-30B-A3B)
|
| 68 |
+
- Trained with GRPO + curriculum learning (2 stages of RL, with increasing task difficulty each stage)
|
| 69 |
+
- Batch size 256, 64 rollouts per task
|
| 70 |
+
- More details [here](https://github.com/Danau5tin/Orca-Agent-RL)
|
| 71 |
+
|
| 72 |
+
## Serving model
|
| 73 |
+
|
| 74 |
+
**vLLM**
|
| 75 |
+
```bash
|
| 76 |
+
vllm serve Danau5tin/Orca-Agent-v0.1
|
| 77 |
+
```
|
| 78 |
+
|
| 79 |
+
**SGLang**
|
| 80 |
+
```bash
|
| 81 |
+
python -m sglang.launch_server \
|
| 82 |
+
--model-path Danau5tin/Orca-Agent-v0.1
|
| 83 |
+
```
|
| 84 |
+
|
| 85 |
+
The agent's orchestration code can be found [here](https://github.com/Danau5tin/multi-agent-coding-system).
|