Text Generation
Transformers
Safetensors
English
llama
raspberry-pi
gpio
embedded
structured-output
json
tiny
Eval Results (legacy)
text-generation-inference
Instructions to use AwaleSagar/gpio-llm-nano-rpi5 with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use AwaleSagar/gpio-llm-nano-rpi5 with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("text-generation", model="AwaleSagar/gpio-llm-nano-rpi5")# Load model directly from transformers import AutoTokenizer, AutoModelForCausalLM tokenizer = AutoTokenizer.from_pretrained("AwaleSagar/gpio-llm-nano-rpi5") model = AutoModelForCausalLM.from_pretrained("AwaleSagar/gpio-llm-nano-rpi5", device_map="auto") - Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- vLLM
How to use AwaleSagar/gpio-llm-nano-rpi5 with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "AwaleSagar/gpio-llm-nano-rpi5" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "AwaleSagar/gpio-llm-nano-rpi5", "prompt": "Once upon a time,", "max_tokens": 512, "temperature": 0.5 }'Use Docker
docker model run hf.co/AwaleSagar/gpio-llm-nano-rpi5
- SGLang
How to use AwaleSagar/gpio-llm-nano-rpi5 with SGLang:
Install from pip and serve model
# Install SGLang from pip: pip install sglang # Start the SGLang server: python3 -m sglang.launch_server \ --model-path "AwaleSagar/gpio-llm-nano-rpi5" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "AwaleSagar/gpio-llm-nano-rpi5", "prompt": "Once upon a time,", "max_tokens": 512, "temperature": 0.5 }'Use Docker images
docker run --gpus all \ --shm-size 32g \ -p 30000:30000 \ -v ~/.cache/huggingface:/root/.cache/huggingface \ --env "HF_TOKEN=<secret>" \ --ipc=host \ lmsysorg/sglang:latest \ python3 -m sglang.launch_server \ --model-path "AwaleSagar/gpio-llm-nano-rpi5" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "AwaleSagar/gpio-llm-nano-rpi5", "prompt": "Once upon a time,", "max_tokens": 512, "temperature": 0.5 }' - Docker Model Runner
How to use AwaleSagar/gpio-llm-nano-rpi5 with Docker Model Runner:
docker model run hf.co/AwaleSagar/gpio-llm-nano-rpi5
Download training/eval_c_engine_mac_unconstrained.json from AwaleSagar/gpio-llm-nano-rpi5: direct link, hf CLI and curl.
- Browser
- Download file 596 Bytes
-
https://huggingface.co/AwaleSagar/gpio-llm-nano-rpi5/resolve/main/training/eval_c_engine_mac_unconstrained.json
- Command line
-
hf download hf://AwaleSagar/gpio-llm-nano-rpi5/training/eval_c_engine_mac_unconstrained.json
-
curl -L -o eval_c_engine_mac_unconstrained.json https://huggingface.co/AwaleSagar/gpio-llm-nano-rpi5/resolve/main/training/eval_c_engine_mac_unconstrained.json
596 Bytes
| { | |
| "file": "mac_nano_unconstrained.out", | |
| "rows": 5083, | |
| "exact": 0.9169781625024592, | |
| "valid_json": 0.9994097973637616, | |
| "unsafe_execute": 0.046612466124661245, | |
| "non_exec_rows": 1845, | |
| "latency_ms_p50": 3.5, | |
| "latency_ms_p95": 6.8, | |
| "latency_ms_max": 25.9, | |
| "per_action_exact": { | |
| "ask_clarification": 0.8560606060606061, | |
| "gpio_mode": 0.9637305699481865, | |
| "gpio_pulse": 0.9755434782608695, | |
| "gpio_pwm": 0.9965397923875432, | |
| "gpio_read": 0.9877622377622378, | |
| "gpio_sequence": 0.9046653144016227, | |
| "gpio_write": 0.8296370967741935, | |
| "report_error": 0.9133459835547122, | |
| "wait": 1.0 | |
| } | |
| } |