AI & ML interests
High-performance local inference, NVIDIA RTX 5080 optimisation, CUDA, quantisation, long-context inference, multimodal models, speculative decoding, and reproducible model builds.
Recent Activity
NInfer RTX 5080
Project-maintained NInfer runtime builds, validation records and model artifacts optimized for NVIDIA RTX 5080 16 GB GPUs.
Current focus: Qwen3.8-27B at true 131,072-token context with:
- Q4 KV cache
- MTP-3 speculative decoding
- Vision support
- mixed Q3/Q4/Q5 quantization
- approximately 3.953 effective BPW
Official model
Qwen3.8-27B for NInfer — RTX 5080 16 GB
Validated artifact:
qwen3_8_27b.ninfer
SHA-256:
c4a7e9ab...ec6b050e21
Source and validation
Canonical source, benchmarks, release records and reproducibility documentation:
github.com/toddballinger/ninfer-5080
The project separates model-artifact identity from runtime releases so later NInfer optimizations can be qualified without implying that the validated model file has changed.