Instructions to use apus-ailab/APUS-OpenJev-v1 with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use apus-ailab/APUS-OpenJev-v1 with Transformers:
# Load model directly from transformers import AutoModel model = AutoModel.from_pretrained("apus-ailab/APUS-OpenJev-v1", device_map="auto") - Notebooks
- Google Colab
- Kaggle
Download ARCHITECTURE.md from apus-ailab/APUS-OpenJev-v1: direct link, hf CLI and curl.
- Browser
- Download file 5.31 kB
-
https://huggingface.co/apus-ailab/APUS-OpenJev-v1/resolve/68e5880df6be3bd820345b9233032e8e325bf4c1/ARCHITECTURE.md
- Command line
-
hf download hf://apus-ailab/APUS-OpenJev-v1@68e5880df6be3bd820345b9233032e8e325bf4c1/ARCHITECTURE.md
-
curl -L -o ARCHITECTURE.md https://huggingface.co/apus-ailab/APUS-OpenJev-v1/resolve/68e5880df6be3bd820345b9233032e8e325bf4c1/ARCHITECTURE.md
Architecture: Decisions Across Compute Budgets
Technical Report (PDF) · Full report (Markdown)
APUS-OpenJev addresses a practical question: how can a language model make useful decisions without requiring the same amount of computation for every application? A workflow may need a choice among a few actions, while another task needs more extensive interpretation of the evidence. The design combines a shared language backbone, joint training across compute budgets, and an output path aligned with candidate decisions.
One shared semantic backbone
The model retains Qwen's language understanding and existing network blocks. Each request supplies the task, evidence, and candidate descriptions in natural language. The candidates define what the application can do; the model learns to interpret their meaning in context. It does not require a new classifier with a permanent set of business labels for every workflow.
The same backbone supports a shorter path and a full path. Both use the model's existing normalization and language-model output head to read decisions. This creates two ways to use the same learned representation, rather than maintaining separate models with unrelated decision rules.
flowchart TD
T[Task, evidence, and dynamic candidates] --> B[Shared semantic backbone]
B --> S[Short-path decision]
B --> F[Full-path decision]
S -. supervised learning .-> L[Training: correctness and cross-depth consistency]
F -. supervision and guidance .-> L
S --> R[Inference: one caller-selected budget]
F --> R
The diagram shows the paths available during training and inference, not a requirement to execute both for every request.
Joint learning across compute budgets
An early exit is useful only if the representation at that point is ready to support the task. Simply stopping a pretrained model sooner does not ensure that its intermediate features can produce a reliable decision.
Training therefore supervises both paths on the same decision examples. The short path must learn to identify the correct candidate using the computation available to it. The full path also learns from the reference answer, preserving a direct objective for the larger budget. Shared trainable parameters connect these objectives: adaptation must support useful decisions at more than one point in the network.
This is the central training constraint. The shorter path is a trained decision path, not an arbitrary cutoff. It can still be weaker on difficult tasks; training creates a usable tradeoff rather than eliminating that tradeoff.
Cross-depth decision consistency
Correct-answer supervision provides the target choice, but the full path also expresses relative preferences across the other candidates. Those preferences provide an additional learning signal for the short path.
A self-distillation objective encourages the short-path distribution to approach the full-path distribution for the same input. Gradients stop through the full-path target in this term; the full path continues learning through its own supervised objective. No external teacher is required for this mechanism.
Consistency is encouraged, not guaranteed. The objective does not make the two budgets interchangeable or turn their probabilities into calibrated confidence. It gives the shorter path guidance about the decision structure learned with more computation.
Decision-aligned output and execution control
At inference, the model scores short labels associated with the supplied candidates. Host code maps the result back to a candidate ID and assembles the response. This avoids generating the syntax of a structured answer token by token while keeping task interpretation inside the language model.
Applications select effort="low" or effort="high" to balance compute cost and decision quality. Low uses the trained shorter path; high uses the full path. Each call follows its selected budget while using the same shared model parameters.
Efficiency can come from executing fewer blocks and avoiding unnecessary output generation. Long-context input processing can still dominate cost, and these mechanisms do not establish a universal speedup. Matched end-to-end measurements are needed for deployment claims.
Scope and practical tradeoffs
The design combines task conditioning, training constraints, and execution control within a shared decision model. It is intended for choices such as static browser actions, workflow routing, and evidence-based judgments. Applications still need useful candidates and task-specific evaluation.
Open-ended TYPE text received full-path supervision; use high for text generation. Candidate scores are relative preferences, and merged weights require their own threshold validation. Quality results and timing boundaries are reported on the main page.
For reproducible details, see the model's training notes, runtime, and merge evaluation. Each model directory preserves its own evidence.
Authors: gumpcheng (https://huggingface.co/xDAN2099), zhangxu, APUS AI-LAB