--- license: apache-2.0 pipeline_tag: feature-extraction language: - en datasets: - devrim/goodwiki_long_synthetic_ir tags: - retrieval - long-documents - document-embeddings - reign base_model: thenlper/gte-small --- # reign-tiny-l1_gn-gte-small_s384_val-selected REIGN `tiny-l1` cross-chunk encoder trained on GoodWiki-Long-Synthetic over a frozen GTE-small guidance network — a long-document bi-encoder that reads a sequence of cached chunk embeddings instead of tokens. ## Configuration | | | | --- | --- | | REIGN encoder | `tiny-l1` — 1 layer, d = 192, 3 heads, FFN 768, **0.56M** trainable parameters | | Guidance network (frozen) | [`thenlper/gte-small`](https://huggingface.co/thenlper/gte-small) — GTE-small, 33M | | Chunk size *K* | 512 (the guidance network's context window) | | Training stride *S* | 384 — stride tag in the checkpoint name | | Chunk-position signal | none — the encoder is a permutation-equivariant set function | | Pooling | mean over the chunk sequence | | Training data | [`devrim/goodwiki_long_synthetic_ir`](https://huggingface.co/datasets/devrim/goodwiki_long_synthetic_ir) | | Training recipe | released cosine recipe (below) | | Checkpoint selection | best validation nDCG@10 on the `val` qrels split | ## Reported results Every number below is the value the paper reports for **this exact checkpoint**; nothing is re-derived here. | Benchmark | Metric | Eval stride | Value | Paper | | --- | --- | --- | ---: | --- | | GoodWiki-Long test | nDCG@10 | s384 | 64.46 | Table 7 | ## Usage The checkpoint holds only the REIGN cross-chunk encoder. The guidance network is loaded separately and stays frozen, so both must be named at construction time. ```bash pip install git+https://github.com/devrimcavusoglu/reign.git ``` ```python import numpy as np from huggingface_hub import snapshot_download from reign.encoders.reign import ReignBaselineEncoder checkpoint_path = snapshot_download("devrim/reign-tiny-l1_gn-gte-small_s384_val-selected") encoder = ReignBaselineEncoder( checkpoint_path=checkpoint_path, gn_model="thenlper/gte-small", chunk_size=512, stride=384, ) docs = [open("doc_a.txt").read(), open("doc_b.txt").read()] emb = encoder.encode(docs, batch_size=8) # (2, hidden_size), L2-normalised print(float(np.dot(emb[0], emb[1]))) # cosine similarity ``` `ReignBaselineEncoder` returns L2-normalised vectors, so the cosine is a dot product. `chunk_size` is the guidance network's sliding-window size — 512 for every released checkpoint, matching its context window — and `stride` controls the overlap, with `stride == chunk_size` giving non-overlapping chunking. The evaluation-time stride is a runtime argument, and the paper's headline tables report the best-performing stride per guidance network. For the lower-level surface, `ReignModel` (a `PreTrainedModel` consuming `inputs_embeds`) and `ReignFeatureExtractor` (the guidance-network wrapper, with the on-disk embedding cache) are importable from `reign` and `reign.feature_extractor`. ## Operating regime REIGN targets **multi-chunk inputs** and primarily document-to-document retrieval. Inputs shorter than the chunk size collapse to a single chunk embedding, leaving the cross-chunk encoder nothing to aggregate — that regime is served by the guidance network alone, and this checkpoint should not be used for it. ## Training recipe The released checkpoints use a three-way cosine embedding loss over document pairs carrying graded targets *s* ∈ {1, 0, −1} — positive, partial, negative — with partial weight λ = 0.5. | Setting | Value | | --- | --- | | Objective | three-way cosine embedding loss, λ = 0.5 | | Batch construction | 18 anchors × (1 positive + 2 partials + 17 in-batch negatives) = 360 pairs / step | | Optimiser | AdamW, lr 1e-5, weight decay 1e-4, cosine annealing | | Epochs | 50, validating every 4 | | Selection | best validation nDCG@10 | | Precision | 16-mixed | | Seed | 42 | | Guidance-network embeddings | precomputed and cached | | Hardware | single 24 GB consumer GPU | A second, separate protocol (InfoNCE at τ = 0.07, batch 48, 20 epochs) is used **only** for the paper's positional-encoding and training-objective ablations. Those arms sit below this released operating point by construction and are not comparable to it. See `docs/TRAINING.md` in the code repository. Because 16-mixed training is not bit-reproducible even at a fixed seed, a retrained checkpoint will not match these weights bit-for-bit; compare metrics, not weights. ## Files - `config.json` — `ReignModel` configuration - `model.safetensors` — encoder weights (float32) ## Links - Code: - Project page: - Dataset: - Paper: *REIGN: Refurbished Embeddings with Integrated Guidance Networks for Efficient Context-Length Scaling*, Findings of the Association for Computational Linguistics: EMNLP 2026 (to appear). ## Citation ```bibtex @inproceedings{cavusoglu2026reign, title = {{REIGN}: Refurbished Embeddings with Integrated Guidance Networks for Efficient Context-Length Scaling}, author = {{\c{C}}avu{\c{s}}o{\u{g}}lu, Devrim and Akba{\c{s}}, Emre}, booktitle = {Findings of the Association for Computational Linguistics: {EMNLP} 2026}, year = {2026}, publisher = {Association for Computational Linguistics}, note = {To appear} } ``` ## License Apache License 2.0. The `devrim/goodwiki_long_synthetic_ir` dataset is released under CC BY-SA 4.0, preserving the share-alike licensing and attribution of GoodWiki and the underlying Wikipedia text.