Foq Reflex 8B
Reference reflex model of the Foq decision engine โ a 100 % local, System 1 arbitration engine for AI agents.
- Format: GGUF, ternary PQ2_0 quantization (2.13 bpw), 2.2 GB
- Runtime:
llama-server(llama.cpp), 4 concurrent slots, Flash Attention - Hardware: any GPU with 4 GB VRAM, Apple Silicon, or x86/ARM CPU
- License: Apache 2.0 (derived from Apache-2.0 open-weight families โ see notices in the GGUF metadata and the repository's THIRD-PARTY-LICENSES.md)
- Optional adapter:
foq_decision_8b.gguf(LoRA, 61 MB) โ decision specialization, +10.7 points measured on the blind exam bench
Install
pip install foq
foq setup
foq setup downloads this model (SHA-256 verified), detects llama-server,
and optionally installs the decision adapter.
Measured results
Production exam (150 cases: business, cognitive traps, sensitive content, unseen fresh cases): 150/150 (100 %) with the adapter, P50 latency 26 ms on a consumer GPU. Full methodology and raw data: github.com/yohanargentina-oss/Foq.
- Downloads last month
- 280
Hardware compatibility
Log In to add your hardware
2-bit
Inference Providers NEW
This model isn't deployed by any Inference Provider. ๐ Ask for provider support