shiv207 commited on
Commit
b75d1aa
·
verified ·
1 Parent(s): 863b58f

Update README.md

Browse files
Files changed (1) hide show
  1. README.md +71 -70
README.md CHANGED
@@ -1,28 +1,23 @@
1
  ---
2
-
3
  base_model: unsloth/gpt-oss-20b-unsloth-bnb-4bit
4
  license: apache-2.0
5
-
6
  language:
7
-
8
- * en
9
-
10
  tags:
11
-
12
- * gpt-oss
13
- * agent
14
- * tool-calling
15
- * react
16
- * lora
17
- * unsloth
18
- * trl
19
- * reasoning
20
- * harmony
21
- * text-generation
22
-
23
  pipeline_tag: text-generation
24
-
25
- ## library_name: transformers
26
 
27
  # GPT-OSS AgentBoi
28
 
@@ -34,27 +29,27 @@ This model was fine-tuned using LoRA adapters on the ReAct subset of Agent-FLAN
34
 
35
  Large language models are often strong conversationalists but can struggle with:
36
 
37
- * Multi-step planning
38
- * Tool selection and invocation
39
- * ReAct-style reasoning workflows
40
- * Structured action generation
41
- * Separating reasoning from final responses
42
 
43
  GPT-OSS AgentBoi adapts GPT-OSS-20B toward these agent-oriented tasks while remaining trainable on consumer hardware through parameter-efficient fine-tuning.
44
 
45
  ## Model Details
46
 
47
- | Item | Value |
48
- | --------------- | ------------------------------------ |
49
- | Model Name | GPT-OSS AgentBoi |
50
- | Author | shiv207 |
51
- | Base Model | unsloth/gpt-oss-20b-unsloth-bnb-4bit |
52
- | Training Method | LoRA |
53
- | Framework | Unsloth |
54
- | Dataset | Agent-FLAN (ReAct subset) |
55
- | Primary Task | Agentic Tool Use |
56
- | Language | English |
57
- | License | Apache 2.0 |
58
 
59
  ## Training Data
60
 
@@ -62,21 +57,21 @@ The model was fine-tuned using examples from the Agent-FLAN dataset, specificall
62
 
63
  These examples teach the model to:
64
 
65
- * Break complex tasks into intermediate steps
66
- * Decide when tool usage is appropriate
67
- * Generate structured actions
68
- * Follow action-observation loops
69
- * Produce concise final responses
70
 
71
  ## Training Setup
72
 
73
  Training was performed using:
74
 
75
- * GPT-OSS-20B
76
- * Unsloth
77
- * TRL
78
- * LoRA adapters
79
- * Google Colab Tesla T4 GPU
80
 
81
  The objective was to improve agentic behavior while keeping training accessible on limited hardware.
82
 
@@ -84,20 +79,20 @@ The objective was to improve agentic behavior while keeping training accessible
84
 
85
  This model is intended for:
86
 
87
- * AI agents
88
- * Tool-calling systems
89
- * Research assistants
90
- * Retrieval-augmented generation workflows
91
- * Multi-step planning tasks
92
- * Agentic reasoning experiments
93
 
94
  Potential applications include:
95
 
96
- * Search agents
97
- * Knowledge retrieval systems
98
- * Function-calling assistants
99
- * Research copilots
100
- * Workflow automation agents
101
 
102
  ## Example
103
 
@@ -109,11 +104,11 @@ Search for the latest SpaceX launch and summarize it.
109
 
110
  ### Expected Agent Behavior
111
 
112
- 1. Analyze the request
113
- 2. Determine that external information is required
114
- 3. Generate a structured search action
115
- 4. Process retrieved information
116
- 5. Produce a concise final answer
117
 
118
  The fine-tuning objective is to increase consistency in these workflows compared to the base model.
119
 
@@ -129,22 +124,28 @@ model, tokenizer = FastLanguageModel.from_pretrained(
129
 
130
  ## Limitations
131
 
132
- * Evaluated primarily through qualitative testing.
133
- * No formal benchmark suite was used.
134
- * Training utilized only a subset of Agent-FLAN.
135
- * Performance may vary on unseen tool schemas.
136
- * Not optimized for general-purpose instruction tuning beyond agent-oriented tasks.
137
 
138
  ## Acknowledgments
139
 
140
  This project builds upon the work of:
141
 
142
- * OpenAI for GPT-OSS and the Harmony conversation format
143
- * Unsloth for efficient GPT-OSS fine-tuning support
144
- * InternLM for the Agent-FLAN dataset
 
 
 
 
 
 
145
 
146
  ## Author
147
 
148
  **shiv207**
149
 
150
- If you find this project useful, feel free to open issues, share feedback, or build on top of it.
 
1
  ---
2
+ model_name: GPT-OSS AgentBoi
3
  base_model: unsloth/gpt-oss-20b-unsloth-bnb-4bit
4
  license: apache-2.0
 
5
  language:
6
+ - en
 
 
7
  tags:
8
+ - gpt-oss
9
+ - agent
10
+ - tool-calling
11
+ - react
12
+ - lora
13
+ - unsloth
14
+ - trl
15
+ - reasoning
16
+ - harmony
17
+ - text-generation
 
 
18
  pipeline_tag: text-generation
19
+ library_name: transformers
20
+ ---
21
 
22
  # GPT-OSS AgentBoi
23
 
 
29
 
30
  Large language models are often strong conversationalists but can struggle with:
31
 
32
+ - Multi-step planning
33
+ - Tool selection and invocation
34
+ - ReAct-style reasoning workflows
35
+ - Structured action generation
36
+ - Separating reasoning from final responses
37
 
38
  GPT-OSS AgentBoi adapts GPT-OSS-20B toward these agent-oriented tasks while remaining trainable on consumer hardware through parameter-efficient fine-tuning.
39
 
40
  ## Model Details
41
 
42
+ | Item | Value |
43
+ |--------|--------|
44
+ | Model Name | GPT-OSS AgentBoi |
45
+ | Author | shiv207 |
46
+ | Base Model | unsloth/gpt-oss-20b-unsloth-bnb-4bit |
47
+ | Training Method | LoRA |
48
+ | Framework | Unsloth |
49
+ | Dataset | Agent-FLAN (ReAct subset) |
50
+ | Primary Task | Agentic Tool Use |
51
+ | Language | English |
52
+ | License | Apache 2.0 |
53
 
54
  ## Training Data
55
 
 
57
 
58
  These examples teach the model to:
59
 
60
+ - Break complex tasks into intermediate steps
61
+ - Decide when tool usage is appropriate
62
+ - Generate structured actions
63
+ - Follow action-observation loops
64
+ - Produce concise final responses
65
 
66
  ## Training Setup
67
 
68
  Training was performed using:
69
 
70
+ - GPT-OSS-20B
71
+ - Unsloth
72
+ - TRL
73
+ - LoRA adapters
74
+ - Google Colab Tesla T4 GPU
75
 
76
  The objective was to improve agentic behavior while keeping training accessible on limited hardware.
77
 
 
79
 
80
  This model is intended for:
81
 
82
+ - AI agents
83
+ - Tool-calling systems
84
+ - Research assistants
85
+ - Retrieval-augmented generation workflows
86
+ - Multi-step planning tasks
87
+ - Agentic reasoning experiments
88
 
89
  Potential applications include:
90
 
91
+ - Search agents
92
+ - Knowledge retrieval systems
93
+ - Function-calling assistants
94
+ - Research copilots
95
+ - Workflow automation agents
96
 
97
  ## Example
98
 
 
104
 
105
  ### Expected Agent Behavior
106
 
107
+ 1. Analyze the request.
108
+ 2. Determine that external information is required.
109
+ 3. Generate a structured search action.
110
+ 4. Process retrieved information.
111
+ 5. Produce a concise final answer.
112
 
113
  The fine-tuning objective is to increase consistency in these workflows compared to the base model.
114
 
 
124
 
125
  ## Limitations
126
 
127
+ - Evaluated primarily through qualitative testing.
128
+ - No formal benchmark suite was used.
129
+ - Training utilized only a subset of Agent-FLAN.
130
+ - Performance may vary on unseen tool schemas.
131
+ - Not optimized for general-purpose instruction tuning beyond agent-oriented tasks.
132
 
133
  ## Acknowledgments
134
 
135
  This project builds upon the work of:
136
 
137
+ - OpenAI for GPT-OSS and the Harmony conversation format.
138
+ - Unsloth for efficient GPT-OSS fine-tuning support.
139
+ - InternLM for the Agent-FLAN dataset.
140
+
141
+ ## Repository
142
+
143
+ Source code and training notebook:
144
+
145
+ GitHub: https://github.com/shiv207
146
 
147
  ## Author
148
 
149
  **shiv207**
150
 
151
+ If you find this project useful, feel free to open issues, share feedback, or build on top of it.