Danau5tin commited on
Commit
9b66289
·
verified ·
1 Parent(s): 5542c0e

Create README.md

Browse files
Files changed (1) hide show
  1. README.md +85 -0
README.md ADDED
@@ -0,0 +1,85 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ ---
2
+ base_model:
3
+ - Qwen/Qwen3-14B
4
+ license: apache-2.0
5
+ datasets:
6
+ - Danau5tin/terminal-tasks
7
+ tags:
8
+ - agent
9
+ - code
10
+ - multi-agent
11
+ ---
12
+
13
+ # Orca-Agent-v0.1
14
+
15
+ ![Orca-Agent-v0.1 banner image](./orca-agent-v01-banner.png)
16
+
17
+ In depth details behind the training, including the training code, is **all open sourced [here](https://github.com/Danau5tin/Orca-Agent-RL)**
18
+
19
+ ## Description
20
+ Orca-Agent-v0.1 is an orchestration agent that acts as the brain of the operation - it receives the user's task but never touches code directly. Instead, it:
21
+
22
+ - Analyses the task and breaks it into focused subtasks
23
+ - Dispatches explorer agents to understand the system
24
+ - Delegates implementation work to coder agents with precise instructions
25
+ - Verifies all changes through additional explorer agents
26
+ - Maintains the context store with all discovered knowledge
27
+
28
+ ## Performance
29
+ | Orchestrator | Subagent | Terminal Bench |
30
+ |--------------|----------|--------|
31
+ | Qwen3-Coder-480B | Qwen3-Coder-480B | 19.7% |
32
+ | Orca-Agent-v0.1-14B | Qwen3-Coder-30B | 18.25% |
33
+ | Qwen3-14B | Qwen3-Coder-30B | 7.0% |
34
+
35
+ This places Orca-Agent-v0.1 (14B) + Qwen3-Coder-Flash (30B MoE), within striking distance of [Qwen3-Coder-480B](https://huggingface.co/Qwen/Qwen3-Coder-480B-A35B-Instruct) using the same architecture which placed #26 on TerminalBench when it was published recently in [my other project](https://github.com/Danau5tin/multi-agent-coding-system).
36
+
37
+ ## Example model output
38
+ ```xml
39
+ <think>
40
+ Okay, now with task_003 identifying a backup, the next step would be to attempt recovery using this identified file.
41
+ </think>
42
+
43
+ <task_create>
44
+ agent_type: 'coder'
45
+ title: 'Attempt recovery using the identified backup file'
46
+ description: |
47
+ Execute the following steps to recover using the backup identified in task_003:
48
+ 1. Move the identified backup file (/tmp/terraform_work/.terraform.tfstate.tmp) to the target location (/infrastructure/recovered_state.json).
49
+ 2. Verify the operation by checking:
50
+ - The existence of the new file at the target location
51
+ - The size matches the original backup file
52
+ - The file permissions are set appropriately (rw-r--r--)
53
+ max_turns: 10
54
+ context_bootstrap:
55
+ - path: '/tmp/terraform_work/.terraform.tfstate.tmp'
56
+ reason: 'The backup file identified in task_003'
57
+ context_refs:
58
+ - 'task_003'
59
+ </task_create>
60
+ ```
61
+
62
+ ## Model training Overview
63
+ - Full fine-tune of Qwen3-14b
64
+ - 32x H100s
65
+ - 16x for training
66
+ - 8x inference for Orca-Agent
67
+ - 8x inference for subagent (Qwen3-Coder-30B-A3B)
68
+ - Trained with GRPO + curriculum learning (2 stages of RL, with increasing task difficulty each stage)
69
+ - Batch size 256, 64 rollouts per task
70
+ - More details [here](https://github.com/Danau5tin/Orca-Agent-RL)
71
+
72
+ ## Serving model
73
+
74
+ **vLLM**
75
+ ```bash
76
+ vllm serve Danau5tin/Orca-Agent-v0.1
77
+ ```
78
+
79
+ **SGLang**
80
+ ```bash
81
+ python -m sglang.launch_server \
82
+ --model-path Danau5tin/Orca-Agent-v0.1
83
+ ```
84
+
85
+ The agent's orchestration code can be found [here](https://github.com/Danau5tin/multi-agent-coding-system).