--- library_name: portallib license: apache-2.0 base_model: Qwen/Qwen3-8B datasets: - RampPublic/portallib-tasks tags: - portal - hypernetwork - lora - peft - multiple-choice --- # PorTAL refit for Qwen3-8B This is a native [PorTAL](https://github.com/ramp-public/portallib) artifact refitted onto `Qwen/Qwen3-8B`. Its 14-task latent table and canonical LoRA-generating core were learned jointly from Qwen3-1.7B and Qwen3-4B and frozen during refitting. Only a fresh Qwen3-8B alignment was trained. The artifact generates rank-8 LoRA factors for the query and value projections of every decoder layer. ## Evaluation One seed was evaluated on the complete 14-task validation suite using continuation log-probability divided by character length (`acc_norm`). Gold continuation token-mean NLL was tracked separately for checkpoint selection. | Model | Macro `acc_norm` | |---|---:| | Frozen Qwen3-8B | 0.6681 | | PorTAL-adapted | 0.7767 | | Absolute lift | +0.1086 | These are research benchmark results for this exact artifact and evaluation recipe, not a general performance guarantee. ## Refit recipe - Base: `Qwen/Qwen3-8B` at `b968826d9c46dd6066d109eabc6255188de91218` - Frozen source carrier: `RampPublic/portal-qwen3-4b` - Dataset: `RampPublic/portallib-tasks` at `ffc3c0e44f529bf64a5ae62ed5db090952db97ea` - Refit data: deterministic seeded sample of up to 1,000 examples per task from the complete training pool - Optimization: 5 epochs, batch size 4, alignment LR `1e-3`, linear decay with 10% warmup, seed 0 - Trainable parameters: target-base alignment only; task latents and canonical core remain frozen - Checkpoint: maximum macro validation `acc_norm`, with lower gold NLL as the tie-breaker - Architecture: q/v targets, rank 8, alpha 16, task latent 256, layer embedding 32, hidden 512, canonical width 1024 ## Usage ```python from portallib import PortalModel portal = PortalModel.from_pretrained( "RampPublic/portal-qwen3-8b", revision="v0.2.0", ) portal.export_peft("rte", "./portal-rte-qwen3-8b") ``` See the [release recipe](https://github.com/ramp-public/portallib/blob/main/REPRODUCING.md) for the full task list, evaluation definition, and refitting procedure. The artifact is Apache-2.0; the benchmark dataset contains components under multiple upstream licenses documented on its dataset card.