Instructions to use apus-ailab/APUS-OpenJev-v1 with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use apus-ailab/APUS-OpenJev-v1 with Transformers:
# Load model directly from transformers import AutoModel model = AutoModel.from_pretrained("apus-ailab/APUS-OpenJev-v1", device_map="auto") - Notebooks
- Google Colab
- Kaggle
File size: 3,671 Bytes
f03f2db | 1 | <svg xmlns="http://www.w3.org/2000/svg" width="1200" height="720" viewBox="0 0 1200 720" role="img" aria-labelledby="title desc"><title id="title">Compute effort and verified HTTP serving reference</title><desc id="desc">4B same-model accuracy comparison: high82.50 percent, low76.25 percent. Separate9B high vLLM HTTP serving reference:25.58/222.36/275.58ms. One RTX PRO6000 Blackwell96GB GPU. No low HTTP latency is inferred.</desc><rect width="1200" height="720" rx="20" fill="#f6f8fc"/><g font-family="Arial, Helvetica, sans-serif"><text x="36" y="47" font-size="28" fill="#14243c" font-weight="700">Choose the compute budget. Measure the response.</text><text x="36" y="78" font-size="16" fill="#60708a" font-weight="400">Selectable effort: low / high 路 single-GPU deployment</text><rect x="24" y="101" width="376" height="105" rx="8" fill="white" stroke="#dce3ee"/><text x="42" y="146" font-size="29" fill="#087f73" font-weight="700">low / high</text><text x="42" y="181" font-size="15" fill="#60708a" font-weight="400">Two selectable compute budgets</text><rect x="416" y="101" width="376" height="105" rx="8" fill="white" stroke="#dce3ee"/><text x="434" y="146" font-size="29" fill="#087f73" font-weight="700">6.25 pp</text><text x="434" y="181" font-size="15" fill="#60708a" font-weight="400">4B measured accuracy trade-off</text><rect x="808" y="101" width="376" height="105" rx="8" fill="white" stroke="#dce3ee"/><text x="826" y="146" font-size="29" fill="#087f73" font-weight="700">1脳 RTX PRO 6000</text><text x="826" y="181" font-size="15" fill="#60708a" font-weight="400">Blackwell 路 96GB GPU</text><rect x="24" y="232" width="1152" height="192" rx="8" fill="white" stroke="#dce3ee"/><text x="44" y="264" font-size="22" fill="#14243c" font-weight="700">4B 路 same-model decision accuracy</text><text x="44" y="316" font-size="19" fill="#14243c" font-weight="700">effort: high</text><rect x="248" y="295" width="594.0" height="29" rx="3" fill="#245cdb"/><text x="854.0" y="317" font-size="17" fill="#245cdb" font-weight="700">82.50% 路 66 / 80</text><text x="44" y="373" font-size="19" fill="#14243c" font-weight="700">effort: low</text><rect x="248" y="352" width="549.0" height="29" rx="3" fill="#087f73"/><text x="809.0" y="374" font-size="17" fill="#087f73" font-weight="700">76.25% 路 61 / 80</text><text x="248" y="406" font-size="12" fill="#60708a" font-weight="400">0</text><text x="944" y="406" font-size="12" fill="#60708a" font-weight="400">100%</text><rect x="24" y="442" width="1152" height="161" rx="8" fill="white" stroke="#dce3ee"/><text x="44" y="475" font-size="22" fill="#14243c" font-weight="700">9B 路 effort: high 路 vLLM HTTP API</text><text x="44" y="518" font-size="16" fill="#60708a" font-weight="400">P50</text><text x="44" y="562" font-size="31" fill="#087f73" font-weight="700">25.58 ms</text><text x="419" y="518" font-size="16" fill="#60708a" font-weight="400">P95</text><text x="419" y="562" font-size="31" fill="#087f73" font-weight="700">222.36 ms</text><text x="794" y="518" font-size="16" fill="#60708a" font-weight="400">P99</text><text x="794" y="562" font-size="31" fill="#087f73" font-weight="700">275.58 ms</text><text x="36" y="635" font-size="16" fill="#14243c" font-weight="700">HTTP full-response latency 路 1脳 RTX PRO 6000 Blackwell 96GB GPU</text><text x="36" y="666" font-size="14" fill="#60708a" font-weight="400">4B quality and 9B HTTP latency are separate measurements. No low-effort HTTP timing is inferred.</text><text x="36" y="693" font-size="13" fill="#60708a" font-weight="400">Quality: same 4B-5949 model, 80 fixed questions. HTTP: supplied 9B high run snapshot, 333/333 success.</text></g></svg> |