AI & ML interests

High-performance local inference, NVIDIA RTX 5080 optimisation, CUDA, quantisation, long-context inference, multimodal models, speculative decoding, and reproducible model builds.

Recent Activity

tmballin  updated a model 7 days ago
ninfer-5080/Qwen3.8-27B-RTX5080
tmballin  updated a Space 15 days ago
ninfer-5080/README
tmballin  published a Space 15 days ago
ninfer-5080/README
View all activity

Organization Card

NInfer RTX 5080

Project-maintained NInfer runtime builds, validation records and model artifacts optimized for NVIDIA RTX 5080 16 GB GPUs.

Current focus: Qwen3.8-27B at true 131,072-token context with:

  • Q4 KV cache
  • MTP-3 speculative decoding
  • Vision support
  • mixed Q3/Q4/Q5 quantization
  • approximately 3.953 effective BPW

Official model

Qwen3.8-27B for NInfer — RTX 5080 16 GB

Validated artifact:

qwen3_8_27b.ninfer

SHA-256:

c4a7e9ab...ec6b050e21

Source and validation

Canonical source, benchmarks, release records and reproducibility documentation:

github.com/toddballinger/ninfer-5080

The project separates model-artifact identity from runtime releases so later NInfer optimizations can be qualified without implying that the validated model file has changed.

datasets 0

None public yet