Update README.md
Browse files
README.md
CHANGED
|
@@ -13,9 +13,29 @@ language:
|
|
| 13 |
Classifier weights and lexical resources for the **LittleCurriculum** five-stage
|
| 14 |
K–5 text filter.
|
| 15 |
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 16 |
These files are the runtime dependencies of
|
| 17 |
[`littlelearner-ll/littlecurriculum-filter`](https://github.com/littlelearner-ll/littlecurriculum-filter).
|
| 18 |
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 19 |
## Usage
|
| 20 |
|
| 21 |
```bash
|
|
@@ -26,6 +46,7 @@ python download_artifacts.py # fetches this repo
|
|
| 26 |
python filter_k5.py --in shard.parquet --out kept.parquet
|
| 27 |
```
|
| 28 |
|
|
|
|
| 29 |
## Contents
|
| 30 |
|
| 31 |
| Path | Size | Used by | What it is |
|
|
@@ -48,17 +69,6 @@ annotation of FineWeb-Edu would have been prohibitively expensive, which is
|
|
| 48 |
what motivates the cascaded design: a cheap lexical stage, then fastText, then
|
| 49 |
the ~50× more expensive ModernBERT.
|
| 50 |
|
| 51 |
-
## Limitations
|
| 52 |
-
|
| 53 |
-
**The classifiers are trained for web prose.**
|
| 54 |
-
Stages 2 and 3 were trained on labels generated from FineWeb-Edu documents. Performance may degrade on substantially different data distributions, in which case retraining the classifiers is recommended.
|
| 55 |
-
|
| 56 |
-
**The grade boundary is fixed to K–5.**
|
| 57 |
-
Stages 2 and 3 use classifiers specifically trained to distinguish K–5 from higher-grade content. Retargeting the pipeline to a different grade boundary therefore requires retraining these classifiers.
|
| 58 |
-
|
| 59 |
-
**The classifiers expect whole documents.**
|
| 60 |
-
They estimate the overall grade level of a document and have little context to work with for very short snippets. We recommend applying the pipeline to full documents rather than individual sentences.
|
| 61 |
-
|
| 62 |
|
| 63 |
## Citation
|
| 64 |
|
|
|
|
| 13 |
Classifier weights and lexical resources for the **LittleCurriculum** five-stage
|
| 14 |
K–5 text filter.
|
| 15 |
|
| 16 |
+
LittleCurriculum is produced from FineWeb-Edu using five sequential stages:
|
| 17 |
+
|
| 18 |
+
1. Age-of-Acquisition and word-frequency pre-filtering
|
| 19 |
+
2. fastText grade-level classification
|
| 20 |
+
3. ModernBERT grade-level classification
|
| 21 |
+
4. Advanced mathematical and symbolic notation filtering
|
| 22 |
+
5. Frequency sampling based on Beyond-K–5-associated vocabulary
|
| 23 |
+
|
| 24 |
These files are the runtime dependencies of
|
| 25 |
[`littlelearner-ll/littlecurriculum-filter`](https://github.com/littlelearner-ll/littlecurriculum-filter).
|
| 26 |
|
| 27 |
+
## Intended Use
|
| 28 |
+
|
| 29 |
+
**The classifiers are trained for web prose.**
|
| 30 |
+
Stages 2 and 3 were trained on labels generated from FineWeb-Edu documents. Performance may degrade on substantially different data distributions, in which case retraining the classifiers is recommended.
|
| 31 |
+
|
| 32 |
+
**The grade boundary is fixed to K–5.**
|
| 33 |
+
Stages 2 and 3 use classifiers specifically trained to distinguish K–5 from higher-grade content. Retargeting the pipeline to a different grade boundary therefore requires retraining these classifiers.
|
| 34 |
+
|
| 35 |
+
**The classifiers expect whole documents.**
|
| 36 |
+
They estimate the overall grade level of a document and have little context to work with for very short snippets. We recommend applying the pipeline to full documents rather than individual sentences.
|
| 37 |
+
|
| 38 |
+
|
| 39 |
## Usage
|
| 40 |
|
| 41 |
```bash
|
|
|
|
| 46 |
python filter_k5.py --in shard.parquet --out kept.parquet
|
| 47 |
```
|
| 48 |
|
| 49 |
+
|
| 50 |
## Contents
|
| 51 |
|
| 52 |
| Path | Size | Used by | What it is |
|
|
|
|
| 69 |
what motivates the cascaded design: a cheap lexical stage, then fastText, then
|
| 70 |
the ~50× more expensive ModernBERT.
|
| 71 |
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 72 |
|
| 73 |
## Citation
|
| 74 |
|