Eliovp's picture
Qwen3.8 27B rotated W3 INT3 weights for the Paiton RDNA4 runtime
278486d verified
|
Raw History Blame Contribute Delete
5.16 kB

Calibration data

This file covers the calibration data of the INT3 weights in this repository (w3rot-manifest.json, SHA-256 335c5c834fb3f455203908d2b5aff8131aca2fe64ad73d59feddf32247d41a70). Every source carries a permissive license: Apache-2.0, MIT, BSD-3-Clause, ODC-By 1.0 or CC BY 4.0. The manifest's sources.datasets section records the revision, license and token count of all 33 entries.

Token counts

  • Calibration: 292,864 tokens (123 sequences of 2,048 tokens plus 5 long-context sequences of 8,192 tokens).
  • Held-out: 32,768 tokens (2 × 2,048 per category plus 1 × 8,192 long context), disjoint from the calibration rows.
  • Category shares follow the calibration mixture published for ISTA-DASLab/Qwen3.8-27B-3Bit-GSQ.

Data mix

Category Source License Tokens Share
Math AI-MO/NuminaMath-CoT Apache-2.0 29,208 10.0 %
Math EleutherAI/hendrycks_math, intermediate algebra ¹ MIT 9,704 3.3 %
Code m-a-p/CodeFeedback-Filtered-Instruction Apache-2.0 29,118 9.9 %
Code Source files: Transformers (3), NumPy (5), PyTorch C++ headers (2) Apache-2.0; BSD-3-Clause 22,082 7.5 %
Science HuggingFaceTB/cosmopedia, openstax chemistry, biology and physics Apache-2.0 45,119 15.4 %
Science allenai/qasc CC BY 4.0 14,273 4.9 %
General HuggingFaceFW/fineweb-edu, sample-10BT ODC-By 1.0 22,010 7.5 %
General OpenAssistant/oasst2, English first turns Apache-2.0 14,854 5.1 %
Multilingual HuggingFaceFW/fineweb-2, 18 languages ² ODC-By 1.0 25,673 8.8 %
Multilingual CohereLabs/aya_dataset Apache-2.0 11,191 3.8 %
Agentic SWE-Gym/OpenHands-SFT-Trajectories MIT 19,453 6.6 %
Agentic NousResearch/hermes-function-calling-v1 Apache-2.0 9,219 3.1 %
Long context HuggingFaceFW/fineweb-edu, documents over 40,000 characters ODC-By 1.0 24,576 8.4 %
Long context HuggingFaceFW/finepdfs-edu, English ODC-By 1.0 16,384 5.6 %
Total 292,864 100 %

CodeFeedback-Filtered-Instruction is Apache-2.0; its card notes that part of its content was generated with third-party language models.

¹ The EleutherAI/hendrycks_math dataset card declares MIT. The original upload, hendrycks/competition_math, was disabled on the Hugging Face Hub after a DMCA notice. Its problems derive from AMC, AIME and Art of Problem Solving material, as do NuminaMath-CoT's amc_aime, aops_forum and math subsets.

² Italian, Hindi, Spanish, Portuguese, German, Korean, Russian, Japanese, Indonesian, Chinese (Mandarin), Vietnamese, Turkish, Thai, Polish, Ukrainian, Dutch, French and Arabic.

Rows were sampled at seeded random offsets. The data was used only to collect activation statistics for quantization, and it is not redistributed.

Decontamination

We dropped every candidate row that shares a 13-word sequence (lower-cased) with GSM8K (train and test), HumanEval, MMLU-Pro (test and validation), our 38 chat evaluation prompts or the needle-test prompts. This removed 5 math rows. The packed sequences share no 13-word sequence with those sets. The manifest records hashes of the reference sets. A license check also excluded source files whose header named any other license.

Required attributions

  • QASC. Tushar Khot, Peter Clark, Michal Guerquin, Peter Jansen and Ashish Sabharwal, QASC: A Dataset for Question Answering via Sentence Composition (Allen Institute for AI), allenai/qasc. It is licensed under CC BY 4.0 and provided as is. We reformatted its facts, questions and answer options into calibration prompts, and no QASC text is redistributed.
  • FineWeb-Edu, FineWeb-2 and FinePDFs-Edu. The calibration contains information from FineWeb-Edu, FineWeb-2 and FinePDFs-Edu (Hugging Face). They are made available under the ODC Attribution License 1.0, and their use is also subject to Common Crawl's Terms of Use.