Text Generation
Transformers
Safetensors
English
gpt-oss
agent
tool-calling
react
lora
unsloth
trl
reasoning
harmony
Instructions to use shiv207/gpt_oss_AGENTBOI with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use shiv207/gpt_oss_AGENTBOI with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("text-generation", model="shiv207/gpt_oss_AGENTBOI")# pip install -U transformers accelerate # Load model directly from transformers import AutoModel model = AutoModel.from_pretrained("shiv207/gpt_oss_AGENTBOI", device_map="auto") - Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- vLLM
How to use shiv207/gpt_oss_AGENTBOI with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "shiv207/gpt_oss_AGENTBOI" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "shiv207/gpt_oss_AGENTBOI", "prompt": "Once upon a time,", "max_tokens": 512, "temperature": 0.5 }'Use Docker
docker model run hf.co/shiv207/gpt_oss_AGENTBOI
- SGLang
How to use shiv207/gpt_oss_AGENTBOI with SGLang:
Install from pip and serve model
# Install SGLang from pip: pip install sglang # Start the SGLang server: python3 -m sglang.launch_server \ --model-path "shiv207/gpt_oss_AGENTBOI" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "shiv207/gpt_oss_AGENTBOI", "prompt": "Once upon a time,", "max_tokens": 512, "temperature": 0.5 }'Use Docker images
docker run --gpus all \ --shm-size 32g \ -p 30000:30000 \ -v ~/.cache/huggingface:/root/.cache/huggingface \ --env "HF_TOKEN=<secret>" \ --ipc=host \ lmsysorg/sglang:latest \ python3 -m sglang.launch_server \ --model-path "shiv207/gpt_oss_AGENTBOI" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "shiv207/gpt_oss_AGENTBOI", "prompt": "Once upon a time,", "max_tokens": 512, "temperature": 0.5 }' - Unsloth Desktop
- Docker Model Runner
How to use shiv207/gpt_oss_AGENTBOI with Docker Model Runner:
docker model run hf.co/shiv207/gpt_oss_AGENTBOI
Update README.md
Browse files
README.md
CHANGED
|
@@ -1,28 +1,23 @@
|
|
| 1 |
---
|
| 2 |
-
|
| 3 |
base_model: unsloth/gpt-oss-20b-unsloth-bnb-4bit
|
| 4 |
license: apache-2.0
|
| 5 |
-
|
| 6 |
language:
|
| 7 |
-
|
| 8 |
-
* en
|
| 9 |
-
|
| 10 |
tags:
|
| 11 |
-
|
| 12 |
-
|
| 13 |
-
|
| 14 |
-
|
| 15 |
-
|
| 16 |
-
|
| 17 |
-
|
| 18 |
-
|
| 19 |
-
|
| 20 |
-
|
| 21 |
-
* text-generation
|
| 22 |
-
|
| 23 |
pipeline_tag: text-generation
|
| 24 |
-
|
| 25 |
-
|
| 26 |
|
| 27 |
# GPT-OSS AgentBoi
|
| 28 |
|
|
@@ -34,27 +29,27 @@ This model was fine-tuned using LoRA adapters on the ReAct subset of Agent-FLAN
|
|
| 34 |
|
| 35 |
Large language models are often strong conversationalists but can struggle with:
|
| 36 |
|
| 37 |
-
|
| 38 |
-
|
| 39 |
-
|
| 40 |
-
|
| 41 |
-
|
| 42 |
|
| 43 |
GPT-OSS AgentBoi adapts GPT-OSS-20B toward these agent-oriented tasks while remaining trainable on consumer hardware through parameter-efficient fine-tuning.
|
| 44 |
|
| 45 |
## Model Details
|
| 46 |
|
| 47 |
-
| Item
|
| 48 |
-
|
|
| 49 |
-
| Model Name
|
| 50 |
-
| Author
|
| 51 |
-
| Base Model
|
| 52 |
-
| Training Method | LoRA
|
| 53 |
-
| Framework
|
| 54 |
-
| Dataset
|
| 55 |
-
| Primary Task
|
| 56 |
-
| Language
|
| 57 |
-
| License
|
| 58 |
|
| 59 |
## Training Data
|
| 60 |
|
|
@@ -62,21 +57,21 @@ The model was fine-tuned using examples from the Agent-FLAN dataset, specificall
|
|
| 62 |
|
| 63 |
These examples teach the model to:
|
| 64 |
|
| 65 |
-
|
| 66 |
-
|
| 67 |
-
|
| 68 |
-
|
| 69 |
-
|
| 70 |
|
| 71 |
## Training Setup
|
| 72 |
|
| 73 |
Training was performed using:
|
| 74 |
|
| 75 |
-
|
| 76 |
-
|
| 77 |
-
|
| 78 |
-
|
| 79 |
-
|
| 80 |
|
| 81 |
The objective was to improve agentic behavior while keeping training accessible on limited hardware.
|
| 82 |
|
|
@@ -84,20 +79,20 @@ The objective was to improve agentic behavior while keeping training accessible
|
|
| 84 |
|
| 85 |
This model is intended for:
|
| 86 |
|
| 87 |
-
|
| 88 |
-
|
| 89 |
-
|
| 90 |
-
|
| 91 |
-
|
| 92 |
-
|
| 93 |
|
| 94 |
Potential applications include:
|
| 95 |
|
| 96 |
-
|
| 97 |
-
|
| 98 |
-
|
| 99 |
-
|
| 100 |
-
|
| 101 |
|
| 102 |
## Example
|
| 103 |
|
|
@@ -109,11 +104,11 @@ Search for the latest SpaceX launch and summarize it.
|
|
| 109 |
|
| 110 |
### Expected Agent Behavior
|
| 111 |
|
| 112 |
-
1. Analyze the request
|
| 113 |
-
2. Determine that external information is required
|
| 114 |
-
3. Generate a structured search action
|
| 115 |
-
4. Process retrieved information
|
| 116 |
-
5. Produce a concise final answer
|
| 117 |
|
| 118 |
The fine-tuning objective is to increase consistency in these workflows compared to the base model.
|
| 119 |
|
|
@@ -129,22 +124,28 @@ model, tokenizer = FastLanguageModel.from_pretrained(
|
|
| 129 |
|
| 130 |
## Limitations
|
| 131 |
|
| 132 |
-
|
| 133 |
-
|
| 134 |
-
|
| 135 |
-
|
| 136 |
-
|
| 137 |
|
| 138 |
## Acknowledgments
|
| 139 |
|
| 140 |
This project builds upon the work of:
|
| 141 |
|
| 142 |
-
|
| 143 |
-
|
| 144 |
-
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 145 |
|
| 146 |
## Author
|
| 147 |
|
| 148 |
**shiv207**
|
| 149 |
|
| 150 |
-
If you find this project useful, feel free to open issues, share feedback, or build on top of it.
|
|
|
|
| 1 |
---
|
| 2 |
+
model_name: GPT-OSS AgentBoi
|
| 3 |
base_model: unsloth/gpt-oss-20b-unsloth-bnb-4bit
|
| 4 |
license: apache-2.0
|
|
|
|
| 5 |
language:
|
| 6 |
+
- en
|
|
|
|
|
|
|
| 7 |
tags:
|
| 8 |
+
- gpt-oss
|
| 9 |
+
- agent
|
| 10 |
+
- tool-calling
|
| 11 |
+
- react
|
| 12 |
+
- lora
|
| 13 |
+
- unsloth
|
| 14 |
+
- trl
|
| 15 |
+
- reasoning
|
| 16 |
+
- harmony
|
| 17 |
+
- text-generation
|
|
|
|
|
|
|
| 18 |
pipeline_tag: text-generation
|
| 19 |
+
library_name: transformers
|
| 20 |
+
---
|
| 21 |
|
| 22 |
# GPT-OSS AgentBoi
|
| 23 |
|
|
|
|
| 29 |
|
| 30 |
Large language models are often strong conversationalists but can struggle with:
|
| 31 |
|
| 32 |
+
- Multi-step planning
|
| 33 |
+
- Tool selection and invocation
|
| 34 |
+
- ReAct-style reasoning workflows
|
| 35 |
+
- Structured action generation
|
| 36 |
+
- Separating reasoning from final responses
|
| 37 |
|
| 38 |
GPT-OSS AgentBoi adapts GPT-OSS-20B toward these agent-oriented tasks while remaining trainable on consumer hardware through parameter-efficient fine-tuning.
|
| 39 |
|
| 40 |
## Model Details
|
| 41 |
|
| 42 |
+
| Item | Value |
|
| 43 |
+
|--------|--------|
|
| 44 |
+
| Model Name | GPT-OSS AgentBoi |
|
| 45 |
+
| Author | shiv207 |
|
| 46 |
+
| Base Model | unsloth/gpt-oss-20b-unsloth-bnb-4bit |
|
| 47 |
+
| Training Method | LoRA |
|
| 48 |
+
| Framework | Unsloth |
|
| 49 |
+
| Dataset | Agent-FLAN (ReAct subset) |
|
| 50 |
+
| Primary Task | Agentic Tool Use |
|
| 51 |
+
| Language | English |
|
| 52 |
+
| License | Apache 2.0 |
|
| 53 |
|
| 54 |
## Training Data
|
| 55 |
|
|
|
|
| 57 |
|
| 58 |
These examples teach the model to:
|
| 59 |
|
| 60 |
+
- Break complex tasks into intermediate steps
|
| 61 |
+
- Decide when tool usage is appropriate
|
| 62 |
+
- Generate structured actions
|
| 63 |
+
- Follow action-observation loops
|
| 64 |
+
- Produce concise final responses
|
| 65 |
|
| 66 |
## Training Setup
|
| 67 |
|
| 68 |
Training was performed using:
|
| 69 |
|
| 70 |
+
- GPT-OSS-20B
|
| 71 |
+
- Unsloth
|
| 72 |
+
- TRL
|
| 73 |
+
- LoRA adapters
|
| 74 |
+
- Google Colab Tesla T4 GPU
|
| 75 |
|
| 76 |
The objective was to improve agentic behavior while keeping training accessible on limited hardware.
|
| 77 |
|
|
|
|
| 79 |
|
| 80 |
This model is intended for:
|
| 81 |
|
| 82 |
+
- AI agents
|
| 83 |
+
- Tool-calling systems
|
| 84 |
+
- Research assistants
|
| 85 |
+
- Retrieval-augmented generation workflows
|
| 86 |
+
- Multi-step planning tasks
|
| 87 |
+
- Agentic reasoning experiments
|
| 88 |
|
| 89 |
Potential applications include:
|
| 90 |
|
| 91 |
+
- Search agents
|
| 92 |
+
- Knowledge retrieval systems
|
| 93 |
+
- Function-calling assistants
|
| 94 |
+
- Research copilots
|
| 95 |
+
- Workflow automation agents
|
| 96 |
|
| 97 |
## Example
|
| 98 |
|
|
|
|
| 104 |
|
| 105 |
### Expected Agent Behavior
|
| 106 |
|
| 107 |
+
1. Analyze the request.
|
| 108 |
+
2. Determine that external information is required.
|
| 109 |
+
3. Generate a structured search action.
|
| 110 |
+
4. Process retrieved information.
|
| 111 |
+
5. Produce a concise final answer.
|
| 112 |
|
| 113 |
The fine-tuning objective is to increase consistency in these workflows compared to the base model.
|
| 114 |
|
|
|
|
| 124 |
|
| 125 |
## Limitations
|
| 126 |
|
| 127 |
+
- Evaluated primarily through qualitative testing.
|
| 128 |
+
- No formal benchmark suite was used.
|
| 129 |
+
- Training utilized only a subset of Agent-FLAN.
|
| 130 |
+
- Performance may vary on unseen tool schemas.
|
| 131 |
+
- Not optimized for general-purpose instruction tuning beyond agent-oriented tasks.
|
| 132 |
|
| 133 |
## Acknowledgments
|
| 134 |
|
| 135 |
This project builds upon the work of:
|
| 136 |
|
| 137 |
+
- OpenAI for GPT-OSS and the Harmony conversation format.
|
| 138 |
+
- Unsloth for efficient GPT-OSS fine-tuning support.
|
| 139 |
+
- InternLM for the Agent-FLAN dataset.
|
| 140 |
+
|
| 141 |
+
## Repository
|
| 142 |
+
|
| 143 |
+
Source code and training notebook:
|
| 144 |
+
|
| 145 |
+
GitHub: https://github.com/shiv207
|
| 146 |
|
| 147 |
## Author
|
| 148 |
|
| 149 |
**shiv207**
|
| 150 |
|
| 151 |
+
If you find this project useful, feel free to open issues, share feedback, or build on top of it.
|