# Training data attribution Loom Spark 2 was trained on several openly licensed corpora. Some of these licences require attribution; this file satisfies that requirement and must be kept with any redistribution. ## SQuAD 2.0 — CC BY-SA 4.0 Rajpurkar, Jia & Liang. "Know What You Don't Know: Unanswerable Questions for SQuAD." https://rajpurkar.github.io/SQuAD-explorer/ Used for grounded reading, and — via its unanswerable questions — for teaching the model to say when a result does not contain the answer. ## MASSIVE — CC BY 4.0 Amazon. https://github.com/alexa/massive Derived from SLURP, also CC BY 4.0. Used for tool-decision training. ## CLINC150 — CC BY 3.0 Larson et al. "An Evaluation Dataset for Intent Classification and Out-of-Scope Prediction." https://github.com/clinc/oos-eval Used for tool-decision training. ## databricks-dolly-15k — CC BY-SA 3.0 Databricks. https://huggingface.co/datasets/databricks/databricks-dolly-15k Used for instruction following. ## OpenAssistant OASST1 — Apache 2.0 LAION / OpenAssistant. https://huggingface.co/datasets/OpenAssistant/oasst1 Used for multi-turn dialogue structure. Only English conversations with short assistant replies were kept. ## Persona curriculum — Textile Labs Identity, limits, warmth and brevity were written for Loom and are not derived from any public dataset. Only English portions were used. No source text was altered except for truncation of passages to a realistic tool-result length, and surface augmentation (casing, punctuation, filler) applied to user turns in training copies only.