Instructions to use psikosen/canopy-258m-r3 with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use psikosen/canopy-258m-r3 with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("text-generation", model="psikosen/canopy-258m-r3", trust_remote_code=True) messages = [ {"role": "user", "content": "Who are you?"}, ] pipe(messages)# Load model directly from transformers import AutoModelForCausalLM model = AutoModelForCausalLM.from_pretrained("psikosen/canopy-258m-r3", trust_remote_code=True, device_map="auto") - Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- vLLM
How to use psikosen/canopy-258m-r3 with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "psikosen/canopy-258m-r3" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "psikosen/canopy-258m-r3", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker
docker model run hf.co/psikosen/canopy-258m-r3
- SGLang
How to use psikosen/canopy-258m-r3 with SGLang:
Install from pip and serve model
# Install SGLang from pip: pip install sglang # Start the SGLang server: python3 -m sglang.launch_server \ --model-path "psikosen/canopy-258m-r3" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "psikosen/canopy-258m-r3", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker images
docker run --gpus all \ --shm-size 32g \ -p 30000:30000 \ -v ~/.cache/huggingface:/root/.cache/huggingface \ --env "HF_TOKEN=<secret>" \ --ipc=host \ lmsysorg/sglang:latest \ python3 -m sglang.launch_server \ --model-path "psikosen/canopy-258m-r3" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "psikosen/canopy-258m-r3", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }' - Docker Model Runner
How to use psikosen/canopy-258m-r3 with Docker Model Runner:
docker model run hf.co/psikosen/canopy-258m-r3
Canopy-258M-R3 v6: Autonomous Tri-Engine Browser Agent & Swarm Coordinator
Canopy-258M-R3 v6 is a 258.56M parameter Recurrent Mixture-of-Experts (MoE) browser automation agent optimized for high-speed edge navigation, anti-bot stealth, and verified persistence state read-backs.
v6 introduces a modular Tri-Engine Execution Ecosystem, an invariant URL Token Normalization Preprocessor, and deep syntheses of 2026 frontier multi-agent reasoning literature.
Key Innovations in v6
1. Cooperative Tri-Engine Browser Swarm
Modern web navigation presents divergent requirements: speed, visual fidelity, and anti-bot evasion. v6 unifies three specialized engines via TriEngineSwarmCoordinator:
- Lightpanda Engine (Headless Zig): Ultra-fast headless execution (~27.5 MB standalone RSS, ~15ms cold navigation, 2.04x speedup). Deployed for high-throughput initial crawling, link harvesting, and structural AXTree indexing.
- Obscura Engine (Native Rust CDP): Stealth browser with native TLS fingerprint impersonation, Canvas/Audio spoofing, and native PNG rasterization (~49.3 MB RSS vs ~656 MiB Chromium). Deployed for visual Set-of-Marks grounding and form submissions on protected sites.
- Chromium Engine: Desktop Blink fallback for heavy client-side single-page applications.
2. URL Token Normalization Preprocessor
As documented in controlled sensitivity evaluations, small language models (258M) exhibit sensitivity to ephemeral address tokens (such as random loopback ports http://127.0.0.1:40809 vs http://127.0.0.1:33987), which alter prompt tokenizations and can divert causal attention.
v6 implements invariant URL normalization:
The agent reasoning context sees deterministic canonical tokens, while the execution layer resolves targets to live page origins transparently.
3. Frontier Multi-Agent Reasoning Syntheses
- Thought Communication Bus (CMU / Meta AI): Disentangles shared coordination thoughts $\hat{Z}{\text{shared}}$ from agent-private intent $\hat{Z}{\text{private}}$, avoiding brittle string serialization.
- Flow Reasoning Refiner (Georgia Tech / MIT, EqR): Iterative recurrent flow refinement toward stable spatial coordinate attractors, eliminating coordinate hallucinations.
- Graph Machine DOM Referral Engine (Iter Labs): $O(n)$ referral graph with dynamic 2-hop pointer chasing for $O(1)$ element retrieval.
- Saved-Record State Verifier: Explicit goal-purpose/field/value binding and persistent read-back contract, eliminating false completion claims ($18 \to 0$).
Live Performance & Swarm Telemetry
| Engine / Modality | Memory (RSS) | Navigation / Step Latency | Visual Receipts | Anti-Bot Stealth |
|---|---|---|---|---|
| Lightpanda (Scraping) | 27.5 MB | ~15 - 35 ms (2.04x speedup) | No (Text Only) | Basic |
| Obscura (Interactions) | 49.3 MB | ~85 - 120 ms | Yes (Full PNG SoM) | Advanced (Sannysoft 100%) |
| Chromium (Baseline) | 656.0 MiB | ~180 - 320 ms | Yes | Standard CDP |
| Cooperative Swarm | <80 MB combined | Optimal Split | Yes | Active |
Model Architecture Specifications
| Hyperparameter | Value | Description |
|---|---|---|
| Total Parameters | 258,555,654 | Standalone weights with tied embeddings |
| Active Parameters | ~112,000,000 | Active parameter compute per token |
| Recurrent Layers | 18 effective layers | 3 Prelude + 6 Recurrent (visited 2x) + 3 Coda |
| Recurrent Scaling | $1/\sqrt{2} \approx 0.7071$ | SMELT recurrence variance stabilization |
| KV-Cache Engine | Prefix Sliding | 128 prefix tokens + 512 sliding window tokens |
| MoE Routing | Top-2 of 8 Experts | Dense first 3 layers, MoE middle/coda layers |
| Context Window | 2,048 tokens | RoPE position embeddings |
| Vocabulary Size | 49,152 | Byte-level BPE tokenizer (Cosmo-2) |
Quickstart: Python Swarm Inference
import asyncio
from miniswardbower.browser.swarm_coordinator import TriEngineSwarmCoordinator
from miniswardbower.core.schemas import BrowserAction, BrowserActionType
async def run_swarm():
swarm = TriEngineSwarmCoordinator()
try:
# Stage 1: High-speed scraping via Lightpanda (~25ms)
tree, text = await swarm.scrape_and_index("https://news.ycombinator.com")
print(f"Indexed {len(tree.elements)} interactive elements via Lightpanda.")
# Stage 2: Stealth interaction & visual audit via Obscura
actions = [
BrowserAction(op=BrowserActionType.TYPE, target="input[name='q']", text="Canopy MoE"),
BrowserAction(op=BrowserActionType.PRESS, key="Enter")
]
results, receipt = await swarm.visual_interact_and_submit(
"https://news.ycombinator.com", actions, capture_audit_screenshot=True
)
print(f"Executed {len(results)} actions. Visual receipt: {len(receipt)} bytes.")
finally:
await swarm.close_all()
asyncio.run(run_swarm())
License
Released under the Apache 2.0 License.
- Downloads last month
- 1,198