ffli commited on
Commit
1062e8f
·
verified ·
1 Parent(s): d7ec4ad

Update README.md

Browse files
Files changed (1) hide show
  1. README.md +21 -11
README.md CHANGED
@@ -13,9 +13,29 @@ language:
13
  Classifier weights and lexical resources for the **LittleCurriculum** five-stage
14
  K–5 text filter.
15
 
 
 
 
 
 
 
 
 
16
  These files are the runtime dependencies of
17
  [`littlelearner-ll/littlecurriculum-filter`](https://github.com/littlelearner-ll/littlecurriculum-filter).
18
 
 
 
 
 
 
 
 
 
 
 
 
 
19
  ## Usage
20
 
21
  ```bash
@@ -26,6 +46,7 @@ python download_artifacts.py # fetches this repo
26
  python filter_k5.py --in shard.parquet --out kept.parquet
27
  ```
28
 
 
29
  ## Contents
30
 
31
  | Path | Size | Used by | What it is |
@@ -48,17 +69,6 @@ annotation of FineWeb-Edu would have been prohibitively expensive, which is
48
  what motivates the cascaded design: a cheap lexical stage, then fastText, then
49
  the ~50× more expensive ModernBERT.
50
 
51
- ## Limitations
52
-
53
- **The classifiers are trained for web prose.**
54
- Stages 2 and 3 were trained on labels generated from FineWeb-Edu documents. Performance may degrade on substantially different data distributions, in which case retraining the classifiers is recommended.
55
-
56
- **The grade boundary is fixed to K–5.**
57
- Stages 2 and 3 use classifiers specifically trained to distinguish K–5 from higher-grade content. Retargeting the pipeline to a different grade boundary therefore requires retraining these classifiers.
58
-
59
- **The classifiers expect whole documents.**
60
- They estimate the overall grade level of a document and have little context to work with for very short snippets. We recommend applying the pipeline to full documents rather than individual sentences.
61
-
62
 
63
  ## Citation
64
 
 
13
  Classifier weights and lexical resources for the **LittleCurriculum** five-stage
14
  K–5 text filter.
15
 
16
+ LittleCurriculum is produced from FineWeb-Edu using five sequential stages:
17
+
18
+ 1. Age-of-Acquisition and word-frequency pre-filtering
19
+ 2. fastText grade-level classification
20
+ 3. ModernBERT grade-level classification
21
+ 4. Advanced mathematical and symbolic notation filtering
22
+ 5. Frequency sampling based on Beyond-K–5-associated vocabulary
23
+
24
  These files are the runtime dependencies of
25
  [`littlelearner-ll/littlecurriculum-filter`](https://github.com/littlelearner-ll/littlecurriculum-filter).
26
 
27
+ ## Intended Use
28
+
29
+ **The classifiers are trained for web prose.**
30
+ Stages 2 and 3 were trained on labels generated from FineWeb-Edu documents. Performance may degrade on substantially different data distributions, in which case retraining the classifiers is recommended.
31
+
32
+ **The grade boundary is fixed to K–5.**
33
+ Stages 2 and 3 use classifiers specifically trained to distinguish K–5 from higher-grade content. Retargeting the pipeline to a different grade boundary therefore requires retraining these classifiers.
34
+
35
+ **The classifiers expect whole documents.**
36
+ They estimate the overall grade level of a document and have little context to work with for very short snippets. We recommend applying the pipeline to full documents rather than individual sentences.
37
+
38
+
39
  ## Usage
40
 
41
  ```bash
 
46
  python filter_k5.py --in shard.parquet --out kept.parquet
47
  ```
48
 
49
+
50
  ## Contents
51
 
52
  | Path | Size | Used by | What it is |
 
69
  what motivates the cascaded design: a cheap lexical stage, then fastText, then
70
  the ~50× more expensive ModernBERT.
71
 
 
 
 
 
 
 
 
 
 
 
 
72
 
73
  ## Citation
74