APUS-OpenJev-v1 / assets /learning-framework.svg
gump2049's picture
Publish complete APUS-OpenJev-v1 models and Technical Report v1.1
68e5880 verified
|
Raw History Blame Contribute Delete
10.8 kB
<svg xmlns="http://www.w3.org/2000/svg" viewBox="0 0 1440 1260" width="1440" height="1260" role="img" aria-labelledby="training-architecture-title training-architecture-desc"><title id="training-architecture-title">Learning decisions across compute budgets.</title><desc id="training-architecture-desc">Training reuses one native backbone and one normalization and output vocabulary head at two depths, with supervised decision losses and stop-gradient deep-to-shallow distillation; Jet-DCRL is a separate research training stage.</desc><defs><marker id="arrow" viewBox="0 0 10 10" refX="9" refY="5" markerWidth="7" markerHeight="7" orient="auto-start-reverse"><path d="M 0 0 L 10 5 L 0 10 z" fill="#526b78"/></marker></defs><rect width="1440" height="1260" fill="white"/><text x="40" y="43" fill="#008477" font-size="15" font-weight="700" font-family="Arial, Helvetica, sans-serif">APUS–OPENJEV · ARCHITECTURE STUDY</text><text x="40" y="86" fill="#17313f" font-size="36" font-weight="700" font-family="Arial, Helvetica, sans-serif">Learning decisions across compute budgets.</text><text x="40" y="122" fill="#526b78" font-size="21" font-weight="400" font-family="Arial, Helvetica, sans-serif">Candidate semantics enter as language; depth becomes a learned decision interface.</text><rect x="336" y="164" width="1056" height="188" rx="8" fill="#f6f9f9" stroke="#cedbdc" stroke-width="1.5"/><text x="356" y="194" fill="#526b78" font-size="15" font-weight="700" font-family="Arial, Helvetica, sans-serif">SHARED SEMANTIC BACKBONE · ONE FORWARD TRAVERSAL</text><path d="M 288 266 H 376" fill="none" stroke="#526b78" stroke-width="2" marker-end="url(#arrow)"/><path d="M 776 266 H 864" fill="none" stroke="#526b78" stroke-width="2" marker-end="url(#arrow)"/><path d="M 576 316 V 428" fill="none" stroke="#526b78" stroke-width="2" marker-end="url(#arrow)"/><path d="M 1088 316 V 428" fill="none" stroke="#526b78" stroke-width="2" marker-end="url(#arrow)"/><path d="M 864 516 H 776" fill="none" stroke="#526b78" stroke-width="2" marker-end="url(#arrow)" stroke-dasharray="6 5"/><path d="M 544 572 V 680" fill="none" stroke="#526b78" stroke-width="2" marker-end="url(#arrow)"/><path d="M 1120 572 V 680" fill="none" stroke="#526b78" stroke-width="2" marker-end="url(#arrow)"/><rect x="40" y="216" width="248" height="100" rx="8" fill="white" stroke="#cedbdc" stroke-width="1.5"/><text x="60" y="250" fill="#17313f" font-size="23" font-weight="700" font-family="Arial, Helvetica, sans-serif">Semantic task input</text><text x="60" y="281" fill="#526b78" font-size="18" font-weight="400" font-family="Arial, Helvetica, sans-serif">Context + question</text><text x="60" y="306" fill="#526b78" font-size="18" font-weight="400" font-family="Arial, Helvetica, sans-serif">Candidate descriptions</text><rect x="376" y="216" width="400" height="100" rx="8" fill="white" stroke="#cedbdc" stroke-width="1.5"/><text x="396" y="250" fill="#17313f" font-size="23" font-weight="700" font-family="Arial, Helvetica, sans-serif">Shared lower segment</text><text x="396" y="281" fill="#526b78" font-size="18" font-weight="400" font-family="Arial, Helvetica, sans-serif">Native blocks 1 … d</text><text x="396" y="306" fill="#526b78" font-size="18" font-weight="400" font-family="Arial, Helvetica, sans-serif">Retain the residual state h(d)</text><rect x="864" y="216" width="448" height="100" rx="8" fill="white" stroke="#cedbdc" stroke-width="1.5"/><text x="884" y="250" fill="#17313f" font-size="23" font-weight="700" font-family="Arial, Helvetica, sans-serif">Remaining upper segment</text><text x="884" y="281" fill="#526b78" font-size="18" font-weight="400" font-family="Arial, Helvetica, sans-serif">Native blocks d+1 … D</text><text x="884" y="306" fill="#526b78" font-size="18" font-weight="400" font-family="Arial, Helvetica, sans-serif">Continue from h(d), without replay</text><rect x="376" y="428" width="400" height="144" rx="8" fill="#eaf7f3" stroke="#008477" stroke-width="1.5"/><text x="396" y="462" fill="#17313f" font-size="23" font-weight="700" font-family="Arial, Helvetica, sans-serif">Low-budget decision distribution</text><text x="396" y="493" fill="#526b78" font-size="18" font-weight="400" font-family="Arial, Helvetica, sans-serif">Shared final norm + vocabulary head</text><text x="396" y="518" fill="#526b78" font-size="18" font-weight="400" font-family="Arial, Helvetica, sans-serif">Project candidate rows at answer position</text><text x="396" y="543" fill="#526b78" font-size="18" font-weight="400" font-family="Arial, Helvetica, sans-serif">p(low) over the request’s candidates</text><rect x="864" y="428" width="448" height="144" rx="8" fill="#eaf7f3" stroke="#008477" stroke-width="1.5"/><text x="884" y="462" fill="#17313f" font-size="23" font-weight="700" font-family="Arial, Helvetica, sans-serif">High-budget decision distribution</text><text x="884" y="493" fill="#526b78" font-size="18" font-weight="400" font-family="Arial, Helvetica, sans-serif">Same norm and output-head parameters</text><text x="884" y="518" fill="#526b78" font-size="18" font-weight="400" font-family="Arial, Helvetica, sans-serif">Same candidate semantics and target space</text><text x="884" y="543" fill="#526b78" font-size="18" font-weight="400" font-family="Arial, Helvetica, sans-serif">p(high) also provides a training target</text><rect x="784" y="474" width="72" height="25" rx="8" fill="white" stroke="white" stroke-width="1.5"/><text x="792" y="493" fill="#008477" font-size="18" font-weight="700" font-family="Arial, Helvetica, sans-serif">sg(·)</text><text x="40" y="465" fill="#526b78" font-size="15" font-weight="700" font-family="Arial, Helvetica, sans-serif">LANGUAGE RETENTION</text><text x="40" y="503" fill="#17313f" font-size="19" font-weight="400" font-family="Arial, Helvetica, sans-serif">Open-ended TYPE examples</text><text x="40" y="532" fill="#17313f" font-size="19" font-weight="400" font-family="Arial, Helvetica, sans-serif">use full-path token CE.</text><text x="40" y="576" fill="#526b78" font-size="17" font-weight="400" font-family="Arial, Helvetica, sans-serif">Interleaved training batches</text><text x="40" y="602" fill="#526b78" font-size="17" font-weight="400" font-family="Arial, Helvetica, sans-serif">preserve a text output path.</text><text x="568" y="625" fill="#526b78" font-size="18" font-weight="400" font-family="Arial, Helvetica, sans-serif">Labels supervise both exits; gold answers are not prompt tokens.</text><rect x="376" y="680" width="936" height="132" rx="8" fill="white" stroke="#cedbdc" stroke-width="1.5"/><text x="396" y="714" fill="#17313f" font-size="23" font-weight="700" font-family="Arial, Helvetica, sans-serif">Joint decision learning</text><text x="396" y="745" fill="#526b78" font-size="18" font-weight="400" font-family="Arial, Helvetica, sans-serif">L = 0.5 CE(target, p(low)) + 0.5 CE(target, p(high))</text><text x="396" y="770" fill="#526b78" font-size="18" font-weight="400" font-family="Arial, Helvetica, sans-serif"> + 0.1 KL(stopgrad[p(high)] || p(low))</text><text x="40" y="727" fill="#526b78" font-size="15" font-weight="700" font-family="Arial, Helvetica, sans-serif">MULTI-OBJECTIVE</text><text x="40" y="762" fill="#17313f" font-size="19" font-weight="400" font-family="Arial, Helvetica, sans-serif">Decision supervision</text><text x="40" y="790" fill="#17313f" font-size="19" font-weight="400" font-family="Arial, Helvetica, sans-serif">+ text retention</text><text x="376" y="850" fill="#526b78" font-size="19" font-weight="400" font-family="Arial, Helvetica, sans-serif">The deep distribution teaches the early exit; both exits remain anchored to the supervised target.</text><text x="40" y="918" fill="#008477" font-size="15" font-weight="700" font-family="Arial, Helvetica, sans-serif">TRAINING PROGRAM · RELEASE LINEAGE AND NEXT STAGE</text><path d="M 312 1026 H 360" fill="none" stroke="#526b78" stroke-width="2" marker-end="url(#arrow)"/><path d="M 752 1026 H 800" fill="none" stroke="#526b78" stroke-width="2" marker-end="url(#arrow)" stroke-dasharray="6 5"/><rect x="40" y="956" width="272" height="152" rx="8" fill="white" stroke="#cedbdc" stroke-width="1.5"/><text x="60" y="990" fill="#17313f" font-size="23" font-weight="700" font-family="Arial, Helvetica, sans-serif">Initialization</text><text x="60" y="1021" fill="#526b78" font-size="18" font-weight="400" font-family="Arial, Helvetica, sans-serif">4B: decision SFT warm-start</text><text x="60" y="1046" fill="#526b78" font-size="18" font-weight="400" font-family="Arial, Helvetica, sans-serif">9B: base initialization</text><text x="60" y="1071" fill="#526b78" font-size="18" font-weight="400" font-family="Arial, Helvetica, sans-serif">Separate model lineages</text><rect x="360" y="956" width="392" height="152" rx="8" fill="white" stroke="#cedbdc" stroke-width="1.5"/><text x="380" y="990" fill="#17313f" font-size="23" font-weight="700" font-family="Arial, Helvetica, sans-serif">Joint supervised curriculum</text><text x="380" y="1021" fill="#526b78" font-size="18" font-weight="400" font-family="Arial, Helvetica, sans-serif">Dual-exit CE + cross-depth KL</text><text x="380" y="1046" fill="#526b78" font-size="18" font-weight="400" font-family="Arial, Helvetica, sans-serif">Interleaved full-path text retention</text><text x="380" y="1071" fill="#526b78" font-size="18" font-weight="400" font-family="Arial, Helvetica, sans-serif">Implemented release training</text><rect x="800" y="956" width="592" height="152" rx="8" fill="white" stroke="#cedbdc" stroke-width="1.5" stroke-dasharray="7 5"/><text x="820" y="990" fill="#17313f" font-size="23" font-weight="700" font-family="Arial, Helvetica, sans-serif">Jet-DCRL · research stage</text><text x="820" y="1021" fill="#526b78" font-size="18" font-weight="400" font-family="Arial, Helvetica, sans-serif">Decision-Calibrated Reinforcement Learning</text><text x="820" y="1046" fill="#526b78" font-size="18" font-weight="400" font-family="Arial, Helvetica, sans-serif">Decision utility + probability quality + reference anchoring</text><text x="820" y="1071" fill="#526b78" font-size="18" font-weight="400" font-family="Arial, Helvetica, sans-serif">Proposed follow-on training; evaluate calibrated budget policies</text><text x="40" y="1149" fill="#526b78" font-size="18" font-weight="400" font-family="Arial, Helvetica, sans-serif">CE and cross-depth distillation are one joint optimization objective, not separate completed stages.</text><line x1="40" y1="1192" x2="1392" y2="1192" stroke="#cedbdc"/><text x="40" y="1225" fill="#526b78" font-size="17" font-weight="400" font-family="Arial, Helvetica, sans-serif">Solid arrows: implemented flow Dashed teacher arrow: stop-gradient Dashed stage: research program</text></svg>