APUS-OpenJev-v1-9B / README.md
gump2049's picture
Add GGUF / MLX versions and Frozen80 results
82c9c56 verified
|
Raw
History Blame Contribute Delete
4.1 kB
metadata
library_name: transformers
license: apache-2.0
base_model: Qwen/Qwen3.5-9B
base_model_relation: finetune
pipeline_tag: text-generation
language:
  - en
  - zh
tags:
  - apus-openjev
  - decision-model
  - structured-output
  - bf16

APUS-OpenJev-v1-9B

English | 涓枃 路 Collection 路 Model family 路 Technical Report 路 Runtime

A Qwen3.5-based decision model for browser action selection, workflow routing, and natural-language principle judgments. This repository contains 9B checkpoint-3000 merged BF16 weights, ready to download independently without a separate LoRA adapter.

Highlights

  • Score dynamic candidates supplied with each request and return their distribution.
  • The included native runtime supports effort=low (16 layers) and effort=high (32 layers). Use high for text generation.
  • Reuse Qwen language representations and vocabulary projection; application code can assemble decisions into structured workflow outputs.

Quick start

python -m pip install huggingface_hub
hf download apus-ailab/APUS-OpenJev-v1-9B --local-dir ./APUS-OpenJev-v1-9B
cd APUS-OpenJev-v1-9B
python -m pip install -r requirements.txt
python examples.py . --device cuda:0 --effort high

Evaluation and training

The merged model in this repository scores 68/80 (85.00%) at full depth on the Frozen80 development panel, covering Browser, HelpSteer3, BoolQ, MNLI, and attribute decisions. See merged-evaluation.json and training.md. This reused engineering panel is not an independent blind benchmark or an end-to-end browser success rate.

Candidate probabilities express relative preference; calibration is required before interpreting them as correctness probabilities. BF16 merging changes some probabilities; see runtime evidence and limitations. The series 9B release selects checkpoint-3000, corresponding to the 85% development-panel result. Checkpoint-5949 remains available in the original family repository.

Series and downloads

The 4B, 9B, and 35B-A3B models use independent repositories grouped in a Collection. Hugging Face displays downloads per model. The original family repository and legacy paths remain available. The 35B-A3B merged release is subject to its own validation and publication status.

GGUF and MLX versions

Quantized versions of this model (Frozen80 68/80) for Ollama / llama.cpp / LM Studio and Apple Silicon Macs: GGUF collection 路 MLX collection.

Version Frozen80 Decisions = BF16
GGUF Q8_0 68/80 80/80
GGUF Q4_K_M 68/80 78/80
MLX 8bit 69/80 79/80
MLX 4bit 68/80 75/80
ollama run hf.co/apus-ailab/APUS-OpenJev-v1-9B-GGUF:Q8_0 --think=false

Ollama needs thinking disabled (--think=false, or "think": false in the API). Per-question results and usage are in each repository.

License and acknowledgments

We thank the Qwen/Qwen3.5-9B team. See LICENSE, provenance, and the pinned source and file identities in release-manifest.json.

Authors: gumpcheng (xDAN2099), zhangxu, APUS AI-LAB.