Ariadne-Laya-TD

A small typed-decision adapter for frozen Laya, with a more stable training recipe.

Ariadne is a lightweight input interface that specializes Laya while keeping its pretrained embeddings, encoder and decision heads frozen. Training updates one identity-initialized affine transformation, containing 1,049,600 parameters, immediately before the encoder stack.

The experiment asks: Instead of adapting the representation to the task, can we adapt the task to the representation? The working hypothesis is that a small learned transformation can make useful structure in the existing frozen model more accessible for a particular task. The benchmark results below measure how well that approach works.

One shared base, many specializations. A deployment can retain the same approximately 421M-parameter Laya base and change only the 1.05M-parameter interface for each specialty. Compatible specialists differ in only the interface's weight matrix and bias, so switching specialties needs to replace those tensors while the backbone stays loaded as opposed to swapping in the entire 400M parameters for each fine-tuned speciality.

This release supplies the Typed Decisions member of that family. The exported FP32 interface is 4.2 MB; ten such interfaces would add about 42 MB alongside one shared base. The design supports a family of domain adapters with a common tokenizer, encoder and decision pathway. Every interface must target the same pinned base revision and compatible configuration.

The included checkpoint scores 75.45% on the local Typed Decisions test, compared with 76.70% for Official Laya-TD. It is seed 4, epoch 8, selected by validation loss. This release uses LR 0.0001.

Across five deterministic runs, typed-decision accuracy was 75.28% ± 0.93 percentage points (sample SD). Official Laya-TD scored 76.70% on the same cases. The table below reports all eight tasks and every seed.

Lower learning rate improved training stability. Changing LR from 3e-4 to 1e-4 raised mean typed-decision accuracy from 70.23% to 75.28% and reduced the across-seed standard deviation from 6.68 to 0.93 points. Sample variance fell by 98.1% in this five-seed comparison. The complete learning-rate comparison is below.

Consistency across seeds. Typed-decision accuracy ranged from 74.25% to 76.55%. AG News stayed between 93.67% and 95.17%, and MASSIVE between 75.00% and 79.00%. These are measurements of this recipe on five seeds and the fixed benchmark bundle.

Architecture

The map is x′ = Wx + b, initialized with W = I, b = 0, immediately after Laya's original token embedding and normalization. Original embeddings, normalization, all 28 ModernBERT blocks, and decision heads remain frozen and in evaluation mode. Gradients pass through the frozen stack to update only the map. There is no added LLM and no extra BERT forward pass.

The combined model has 422,343,427 parameters; 1,049,600 were trainable (about 0.25%). Each FP32 adapter checkpoint is approximately 4.2 MB. The original model and tokenizer are downloaded separately.

Base: convaiinnovations/laya, revision 55cf4c4ebb4ebe31b2550e8bdf3bd21b99753851. Source weights were hashed before training and at checkpoint saves; all pretrained tensors remained unchanged.

Evaluation

The main comparison is Official Laya-TD, the evaluation reference. Our released interface was trained on base Laya. The training source and evaluation reference are recorded separately. Both use the same RTX6000 runtime, BF16 autocast, strict deterministic settings, 1024/256 token budgets and benchmark cases. Each model retains its own native SDK calibration.

All scores below are local measurements on identical frozen cases. Accuracy is the highest-probability label accuracy, including ordinal questions; failed decisions count as incorrect. Values are percentages; variance is sample variance in percentage-points squared. Five seeds are used for the mean and SD. These are not the upstream model card's historical benchmark numbers.

Benchmark Decisions Official Laya-TD Adapter mean ± SD Variance Selected checkpoint Selected − reference (pp)
Typed decisions 2000 76.70 75.28 ± 0.93 0.8620 75.45 -1.25
AG News 600 94.67 94.47 ± 0.58 0.3389 95.17 +0.50
BoolQ 600 82.00 82.10 ± 0.89 0.7861 83.17 +1.17
DAIR Emotion 600 58.17 57.20 ± 1.00 1.0056 57.83 -0.33
Prompt injections 116 67.24 74.48 ± 3.93 15.4578 74.14 +6.90
SST-5 600 48.00 41.00 ± 2.46 6.0278 41.00 -7.00
MASSIVE intent EN 300 81.00 76.80 ± 1.56 2.4222 77.00 -4.00
XNLI EN 300 86.33 87.27 ± 1.04 1.0778 86.33 +0.00

What changed from the first release

The earlier adapter used LR 3e-4. This version uses 1e-4 with the same seeds, data, initialization, batch, stopping rule and deterministic runtime. The table compares all five runs of each recipe. Different stopping epochs are a consequence of the shared validation rule.

Benchmark Earlier LR 3e-4 mean ± SD This LR 1e-4 mean ± SD
Typed decisions 70.23 ± 6.68 75.28 ± 0.93
AG News 54.57 ± 35.43 94.47 ± 0.58
BoolQ 63.27 ± 16.02 82.10 ± 0.89
DAIR Emotion 30.33 ± 24.82 57.20 ± 1.00
Prompt injections 57.76 ± 15.54 74.48 ± 3.93
SST-5 30.53 ± 13.61 41.00 ± 2.46
MASSIVE intent EN 32.53 ± 39.58 76.80 ± 1.56
XNLI EN 55.53 ± 29.33 87.27 ± 1.04

The earlier published checkpoint was seed 1, epoch 14. The new validation-selected checkpoint improves AG News from 93.17% to 95.17%, BoolQ from 79.83% to 83.17%, and MASSIVE from 74.33% to 77.00%. Its typed-decision score is 75.45%, compared with 76.95% for the earlier checkpoint. The five-seed means above measure the improvement in the training recipe.

Individual seeds

Seed Typed AG News BoolQ Emotion Injection SST-5 MASSIVE XNLI Best epoch Stop epoch
0 74.25 94.17 80.83 58.50 69.83 38.17 79.00 89.00 11 14
1 75.65 94.83 82.67 57.17 78.45 44.83 75.67 87.00 9 12
2 74.50 94.50 81.83 56.00 71.55 39.83 75.00 87.33 13 16
3 76.55 93.67 82.00 56.50 78.45 41.17 77.33 86.67 8 11
4 75.45 95.17 83.17 57.83 74.14 41.00 77.00 86.33 8 11

Selected checkpoint versus Official Laya-TD: paired uncertainty

These descriptive 95% intervals resample cases 10,000 times, keeping each typed case's five decisions together. They describe benchmark-case uncertainty for the selected model, separately from the across-seed SD above. They have no multiple-comparison correction and were not used to select the checkpoint.

Benchmark Change from Official Laya-TD (pp) Paired 95% interval (pp)
Typed decisions -1.25 [-3.00, +0.50]
AG News +0.50 [-0.50, +1.50]
BoolQ +1.17 [-0.83, +3.17]
DAIR Emotion -0.33 [-2.67, +1.83]
Prompt injections +6.90 [+2.59, +12.07]
SST-5 -7.00 [-10.50, -3.50]
MASSIVE intent EN -4.00 [-7.67, -0.67]
XNLI EN +0.00 [-2.67, +2.67]

Checkpoint selection

The default package contains seed 4, epoch 8, selected by the lowest validation soft cross-entropy across the five seeds (0.854020). This selection rule was recorded before the sweep. Test performance was not used to choose the checkpoint. Only the selected adapter weights are included; all five seeds’ results are reported.

Benchmark scope

  • Typed Decisions: 400 test cases / 2,000 decisions from four synthetic workflows. Training and validation case IDs are disjoint from test IDs. The test task itself is the adaptation task.
  • AG News, BoolQ, Emotion and SST-5: fixed prefixes of 600 examples each. Prompt injections: all 116 test examples. MASSIVE and XNLI: first 300 English test examples each.
  • MASSIVE uses the gold intent plus 19 seeded distractors (20 options), not all intents. XNLI uses three labels.
  • These seven retention tasks were excluded from adapter training; some were present in the base model's training. These are held-out examples for this adaptation, not proof of wholly unseen-task generalization.
  • The test suite was inspected during earlier exploratory experiments. This recipe was selected adaptively after earlier results; seed 0 began as the lower-LR pilot and was retained for the five-seed confirmation. Checkpoint selection used validation only, but the research process is not a blind test. Earlier runs with changing stopping rules are excluded from the main table.
  • No downstream deployment, multilingual retention, or broad safety claim is established by these small English benchmark subsets.

Training

Training data: LocalLLaMA/typed-decisions, revision f7a2487edd7a043a5441a5e9ccc7fe5ddbd9ebe8. A case-level split with seed 42 yields 1,080 training cases (5,400 decisions) and 120 validation cases (600 decisions). Questions from one case never cross splits.

  • Seeds 0, 1, 2, 3, 4; RTX PRO 6000 Blackwell only, sequential runs.
  • Every map starts at the same identity transformation and native dropout stays off. Seeds control training-example shuffling; this comparison measures sensitivity to data order with deterministic computation.
  • AdamW, LR 0.0001, weight decay 0.01, gradient clipping at norm 1; effective batch 32 (microbatch 8 × accumulation 4). No learning-rate warmup or scheduler.
  • Soft-target cross-entropy on raw logits. The pretrained stack stays in eval mode; its dropout is off. The affine map is computed in FP32; the remaining CUDA forward uses BF16 autocast.
  • Minimum 10 completed epochs; stop after three consecutive epochs without a strictly lower validation loss. Patience counts from epoch 1, with stopping gated by the minimum. No fixed maximum. Keep the best trained checkpoint; epoch zero is a separate baseline.
  • Input/head token budgets: 1024/256, matched between Official Laya-TD and our adapter. These differ from the original base Laya config's defaults.

Reproducibility

Strict deterministic algorithms were enabled with errors on unsupported operations, cuDNN benchmarking and cuDNN SDPA were disabled, and float32 matmul precision was set to highest. Python, NumPy and PyTorch seeds were set. Processes used PYTHONHASHSEED=0 and CUBLAS_WORKSPACE_CONFIG=:4096:8.

Two independent two-epoch seed-0 verification runs (338 optimizer updates each) matched losses, validation metrics, and checkpoint tensors bit for bit. The reported seed-0 run matched that two-epoch metric prefix. This verifies the checked environment and prefix; it does not promise bitwise agreement across GPUs, libraries or operating systems. See environment.json and evidence/replay-check.json.

Before training, the identity adapter matched every native-base answer object on all 5,116 benchmark decisions. All five final runs were independently rescored and checked for unchanged pretrained weights. Export verification is recorded separately in export-verification.json after testing the portable package.

Confidence and temperature warning

The pinned base checkpoint ships choice:11+ temperature 0.1005828. Laya 0.3.20 clamps it to 0.5 and emits a warning on load. Our adapter and its untrained base control use that same behavior. The official Laya-TD reference retains its own checkpoint calibration. Only the 300 MASSIVE examples use this bucket in this suite. These temperatures are not used in the training loss.

Temperature scaling preserves class order mathematically but changes probability sharpness and calibration metrics. Calibration was not refitted after adaptation. Confidence calibration should be evaluated for the intended application; the base model's calibration claims have not been established for this adapter.

Load locally

This is a custom adapter. Use the bundled loader: laya.load(repo_id) and AutoModel.from_pretrained(repo_id) do not install this input transformation.

Download the repository files with the Hugging Face UI or hf download GoatHerder/Ariadne-Laya-TD --local-dir ariadne-laya-td, then change into that directory. The learning rate and selected seed are recorded in training_protocol.json and adapter_config.json.

Install a PyTorch build for your GPU and the versions in requirements.txt. The measured environment used PyTorch 2.13.0+cu130; package versions and CUDA/cuDNN details are in environment.json. Run from this downloaded repository directory:

from load_adapter import load

model = load(device="cuda:0")  # defaults to the validation-selected seed
answer = model.predict("The customer says the delivery arrived damaged.", {
    "route": {"type": "choice", "instructions": "Choose the support queue.",
              "criteria": {"delivery": "Delivery and damaged items",
                           "billing": "Payments and invoices"}}
})
print(answer)
model.close()

Only the validation-selected seed is included in this compact package. Use device='cpu' for CPU inference. The loader downloads the pinned base, verifies base configuration/tokenizer hashes and model tensors, checks adapter weights, then freezes all parameters for inference. local_files_only=True uses cached weights; base_path can point to an existing copy of the pinned base.

Switching specializations

The shared-base design makes the interface the unit of specialization: keep the base resident and replace the compatible affine weight and bias. This preserves the original embeddings, all 28 encoder blocks and the decision heads. The adapter adds no additional encoder pass.

The current load() helper constructs one model instance with a selected interface. In-place switching within a running service remains a serving integration step. Such a service must validate the base revision and interface shape and coordinate switching with active requests. This release does not include measured switching latency.

Reproduce preparation and one training run

The bundled ariadne_bench source contains the actual tokenizer/prompt preparation, training loop, adapters and scorer. Dataset files are downloaded from their pinned sources; evaluation data is not redistributed.

python -m ariadne_bench.experiments.prepare_alignment --out data/alignment
python -m ariadne_bench prepare --profile laya-core \
  --suites typed-decisions,ag-news,emotion,boolq,sst5,prompt-injections,massive-intent.en,xnli.en \
  --out data/heldout-eight

The expected case-file hashes and immutable dataset revisions are in training_protocol.json and benchmarks/sources.lock.json. make_training_config.py resolves the pinned base into a local training config:

python make_training_config.py --seed 0 --device cuda:0 --out train-config.json
PYTHONHASHSEED=0 CUBLAS_WORKSPACE_CONFIG=:4096:8 USE_TF=0 OMP_NUM_THREADS=4 \
  python -m ariadne_bench.experiments.align --config train-config.json \
  --train data/alignment/train --validation data/alignment/validation \
  --out training-seed-0 --epochs 0 --min-epochs 10 --early-stopping-patience 3 \
  --batch-size 8 --accumulation 4 --lr 0.0001 --keep-best-only

PYTHONHASHSEED=0 CUBLAS_WORKSPACE_CONFIG=:4096:8 USE_TF=0 OMP_NUM_THREADS=4 \
  python -m ariadne_bench run --data data/heldout-eight \
  --config training-seed-0/benchmark-config.json --out evaluation-seed-0 --warmup 5

License and attribution

Apache-2.0. The base-model card and Typed Decisions dataset card list Apache-2.0. Base weights are downloaded from their original repository. Evaluation datasets retain their own licenses. Prompt preparation credits and upstream notices are in THIRD_PARTY.md and licenses/.

This is an independent adaptation of Laya. The upstream authors did not produce these adapter weights or experimental results.

Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for GoatHerder/Ariadne-Laya-TD

Adapter
(12)
this model

Dataset used to train GoatHerder/Ariadne-Laya-TD