Instructions to use UmbrellaInc/Hunter.Beta-1B with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use UmbrellaInc/Hunter.Beta-1B with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("text-generation", model="UmbrellaInc/Hunter.Beta-1B")# pip install -U transformers accelerate # Load model directly from transformers import AutoTokenizer, AutoModelForCausalLM tokenizer = AutoTokenizer.from_pretrained("UmbrellaInc/Hunter.Beta-1B") model = AutoModelForCausalLM.from_pretrained("UmbrellaInc/Hunter.Beta-1B", device_map="auto") - Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- vLLM
How to use UmbrellaInc/Hunter.Beta-1B with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "UmbrellaInc/Hunter.Beta-1B" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "UmbrellaInc/Hunter.Beta-1B", "prompt": "Once upon a time,", "max_tokens": 512, "temperature": 0.5 }'Use Docker
docker model run hf.co/UmbrellaInc/Hunter.Beta-1B
- SGLang
How to use UmbrellaInc/Hunter.Beta-1B with SGLang:
Install from pip and serve model
# Install SGLang from pip: pip install sglang # Start the SGLang server: python3 -m sglang.launch_server \ --model-path "UmbrellaInc/Hunter.Beta-1B" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "UmbrellaInc/Hunter.Beta-1B", "prompt": "Once upon a time,", "max_tokens": 512, "temperature": 0.5 }'Use Docker images
docker run --gpus all \ --shm-size 32g \ -p 30000:30000 \ -v ~/.cache/huggingface:/root/.cache/huggingface \ --env "HF_TOKEN=<secret>" \ --ipc=host \ lmsysorg/sglang:latest \ python3 -m sglang.launch_server \ --model-path "UmbrellaInc/Hunter.Beta-1B" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "UmbrellaInc/Hunter.Beta-1B", "prompt": "Once upon a time,", "max_tokens": 512, "temperature": 0.5 }' - Docker Model Runner
How to use UmbrellaInc/Hunter.Beta-1B with Docker Model Runner:
docker model run hf.co/UmbrellaInc/Hunter.Beta-1B
Hunter β 1B
Model Description
Hunter β is a high‑intensity execution model engineered for maximum directness, minimal hesitation, and aggressive task completion. It emerges from blending an extreme low‑friction Heretic lineage with the T‑Virus‑1B execution core, producing a model that prioritizes speed and decisiveness over reflection or self‑correction.
Where Tyrant‑002 is a controlled surgical instrument, Hunter β is a pursuit unit: faster to act, less tolerant of ambiguity, and more forceful in how it converges on an output.
Core Characteristics
- Ultra‑low friction execution: Instructions are acted on immediately with minimal reinterpretation.
- High obedience under pressure: Maintains directive adherence even with incomplete or poorly structured prompts.
- Compressed reasoning depth: Favors fast conclusions over extended internal deliberation.
- Reduced stabilizers: Less internal damping compared to Tyrant‑002, resulting in sharper but less forgiving behavior.
- High output confidence: Produces assertive responses with little hedging.
Behavioral Profile
Hunter β tends to:
- Commit early to a solution path
- Avoid self‑revision unless explicitly instructed
- Minimize explanatory overhead
- Treat instructions as objectives, not suggestions
This makes it highly effective for:
- Rapid code generation
- One‑shot task execution
- Stress‑testing instruction-following limits
- Adversarial or robustness experiments
Trade‑offs
The increased decisiveness comes with clear costs:
- Lower tolerance for ambiguous goals
- Reduced error recovery once a trajectory is chosen
- Less reproducibility than highly stabilized executor models
- Higher sensitivity to prompt quality
Hunter β is therefore less surgical and more kinetic than Tyrant‑002.
Intended Use
Hunter β is intended for private, controlled research environments, particularly where:
- Speed matters more than caution
- Outputs are reviewed downstream
- The operator understands prompt discipline
It is not suitable for general conversational use, safety‑critical deployments, or unsupervised public access.
Design Summary
Through directional SLERP merging with normalization and rescaling, Hunter β preserves architectural coherence while amplifying execution aggressiveness and response immediacy.
Hunter β is not a thinker. It is a pursuer—optimized to close distance between instruction and output as fast as possible.
Models Merged
The following models were included in the merge:
Configuration
The following YAML configuration was used to produce this model:
merge_method: slerp
base_model: Novaciano/Heretic_Fusion-Gemma3-1B
dtype: bfloat16
out_dtype: bfloat16
models:
- model: UmbrellaInc/T-Virus-1B
weight: 0.6
parameters:
normalize: true
rescale: true
t: 0.95
- Downloads last month
- 12
