Measured peak <192 GB; estimated 256GB fit.
AI & ML interests
Fast local LLM inference and tested model releases for Apple Silicon and NVIDIA CUDA.
Recent Activity
View all activity
Organization Card
OPEN-SOURCE LOCAL INFERENCE
TensorFold
Fast local LLM inference and tested model releases for Apple Silicon and NVIDIA CUDA.
TensorFold is an open-source inference runtime with an OpenAI-compatible API. This Hugging Face organization publishes compatible checkpoints and supporting assets tested on real hardware.
What you will find here
- Apple Silicon: MLX quantized checkpoints with measured speed and memory use.
- Drafted decoding: MTP and DFlash assets when the upstream model provides a compatible drafter.
- NVIDIA CUDA: tested recipes for supported GPUs and DGX Spark systems.
- Clear release notes: runtime versions, setup steps, known limits, licences, and upstream credit.
Run TensorFold
curl -fsSL https://tensorfold.dev/install.sh | sh
TensorFold does not train the base models. Model design, training, evaluations, and upstream documentation remain the work of the original authors and contributors.
models 66
TensorFold/detectra-v1
Image Classification • 21.8M • Updated • 1
TensorFold/Solar-Open2-250B-MLX-8bit
Text Generation • 250B • Updated • 94
TensorFold/Solar-Open2-250B-MLX-6bit
Text Generation • 250B • Updated • 385 • 1
TensorFold/Solar-Open2-250B-MLX-4bit
Text Generation • 250B • Updated • 104
TensorFold/Qwen3.8-Flash-Next-oQ8
Updated
TensorFold/Qwen3.8-Flash-Next-MLX-oQ8-MTP
Image-Text-to-Text • 180B • Updated • 1.44k • 5
TensorFold/Qwen3.8-Flash-Next-MLX-oQ6-MTP
Image-Text-to-Text • 180B • Updated • 944 • 3
TensorFold/Qwen3.8-Flash-Next-MLX-oQ6
Image-Text-to-Text • 177B • Updated • 423 • 1
TensorFold/Qwen3.8-Flash-Next-MLX-oQ4-MTP
Image-Text-to-Text • 180B • Updated • 4.68k • 14
TensorFold/Qwen3.8-Flash-Next-MLX-oQ4
Image-Text-to-Text • 177B • Updated • 704 • 1
datasets 0
None public yet