Foq Reflex 8B

Reference reflex model of the Foq decision engine โ€” a 100 % local, System 1 arbitration engine for AI agents.

  • Format: GGUF, ternary PQ2_0 quantization (2.13 bpw), 2.2 GB
  • Runtime: llama-server (llama.cpp), 4 concurrent slots, Flash Attention
  • Hardware: any GPU with 4 GB VRAM, Apple Silicon, or x86/ARM CPU
  • License: Apache 2.0 (derived from Apache-2.0 open-weight families โ€” see notices in the GGUF metadata and the repository's THIRD-PARTY-LICENSES.md)
  • Optional adapter: foq_decision_8b.gguf (LoRA, 61 MB) โ€” decision specialization, +10.7 points measured on the blind exam bench

Install

pip install foq
foq setup

foq setup downloads this model (SHA-256 verified), detects llama-server, and optionally installs the decision adapter.

Measured results

Production exam (150 cases: business, cognitive traps, sensitive content, unseen fresh cases): 150/150 (100 %) with the adapter, P50 latency 26 ms on a consumer GPU. Full methodology and raw data: github.com/yohanargentina-oss/Foq.

Downloads last month
280
GGUF
Model size
15.3M params
Architecture
qwen3
Hardware compatibility
Log In to add your hardware

2-bit

Inference Providers NEW
This model isn't deployed by any Inference Provider. ๐Ÿ™‹ Ask for provider support