APUS-OpenJev-v1 / 4B-5949 /README.md
gump2049's picture
Add evaluation dataset link and vLLM serving example
c33a214 verified
|
Raw History Blame
1.92 kB

APUS-OpenJev-v1 路 4B

A decision model for browser agents and business workflows. This directory contains standalone BF16 weights and a runtime with selectable effort="low" and effort="high" compute budgets.

Model family 路 Architecture 路 Runtime guide

Quick start

Use a CUDA-capable PyTorch environment.

python -m pip install huggingface_hub
hf auth login
hf download apus-ailab/APUS-OpenJev-v1 \
  --include "4B-5949/*" --local-dir ./APUS-OpenJev-v1
cd ./APUS-OpenJev-v1/4B-5949
python -m pip install -r requirements.txt
python examples.py . --device cuda:0 --effort high

The included runtime provides compute-budget selection. Use high for text generation.

Evaluation

With the full compute budget, this merged model scores 66/80 (82.50%) on the Frozen80 development panel: browser action selection, principle-based judgment, evidence-based questions, natural language inference, and attribute decisions.

This reused development panel is an engineering reference, not an independent benchmark. BF16 merging changes some candidate probabilities; decision thresholds require revalidation. See evaluation results and runtime checks for details.

Provenance

We thank the Qwen team for the Qwen3.5-4B base model. Training details and source records are in training.md; artifact hashes are in release-manifest.json. Consult LICENSE and the base-model and source-project notices.

Authors: gumpcheng (https://huggingface.co/xDAN2099), zhangxu, APUS AI-LAB