Corrected content-token measurements over the frozen audit dataset: 84,005,795 documents, 94 Common Crawl snapshots, 233 distinct score values · paper: DOI 10.5281/zenodo.21740081 · code & data · NM AI Research, ORCID 0009-0003-4213-7769 · CC BY 4.0
Correction notice. The truncation series below excludes the CLS and SEP special tokens. The linked DOI and PDF remain the historical release. The separate fixed-score extension is retained in the repository as historical data and is unavailable here pending a raw-text recalculation. See correction details.
The slider steps through every score that occurs in the sample, with 233 distinct values observed in this sample, consistent with bfloat16 output quantisation. The shipped filter admits a document when its rounded score reaches 3, which on this grid means a raw score strictly above 2.5. The rounding rule changes the answer at one place only: documents sitting exactly on the cut, which matters at 2.5 and 4.5 and nowhere else.
Hover a bar for the exact grid value.
Rounding half-to-even sends 2.5 down to 2 and 4.5 down to 4. Rounding half-up sends them to 3 and 5. Only the first of those changes what enters the corpus, and the shipped filter uses the rule that cuts. Switch the rule above and this cohort appears and disappears.
The score is produced from at most 510 content tokens. Measured truncation is higher in later snapshots of this sample. This descriptive association does not isolate the cause or establish a monotonic increase.
Hover a point for the snapshot.
The rounding result needs none of the downloads and none of this page. Read classification/run_edu_bert.py in Hugging Face’s Cosmopedia repository, then run one line of NumPy against any FineWeb-Edu shard. The button below runs the same identity in your browser against the 233 grid values embedded in this page, weighted by how many documents sit on each.
(np.round(np.clip(df.score, 0, 5)).astype(int) == df.int_score).all() # True (np.floor(np.clip(df.score, 0, 5) + 0.5).astype(int) == df.int_score).all() # False
not run yet
The first line is the rounding NumPy performs and the one the classifier's own inference script uses. The second is the rounding most readers assume. They disagree about two grid points, and one of those two sits on the admission boundary.
What this is. A front-end over the frozen dataset behind What actually admits a document to FineWeb-Edu. The score grid, corrected per-snapshot truncation series and corrected retained-document summary are embedded in this file. Nothing is fetched at runtime. The corrected measurements differ from the historical DOI release as documented in the correction notice.
Sample. One random shard from each Common Crawl dump in HuggingFaceFW/fineweb-edu-score-2, seeded and pinned before download. The 16 dumps covering 2024 and 2025 partition documents across shards by score, so a single shard from them is not a random sample; they are excluded from the headline, leaving 94 dumps and 84,005,795 documents spanning 2013 to 2023.
Guardrails. fineweb-edu-score-2 retains int_score >= 2 only, so the grid starts at a raw 1.5 and nothing here speaks to documents below it; the 975,821 documents at exactly 1.5 are a floor, not a full count. Retention percentages on this page are shares of that already-filtered population, not of raw Common Crawl. Truncation uses 188,000 sampled documents, 2,000 per snapshot, clustered within selected shards. No population confidence interval is claimed. The historical fixed-score series used capped, first-eligible document pools; fixing the score does not remove all composition or selection effects. Its correction is pending and it is not plotted.
What this does not claim. No downstream model was trained or evaluated here, so nothing on this page says the corpus is worse for any of it. The finding is that a filter described in one line makes three decisions the line does not mention, and rounding admits most of the retained documents in this sample, as quantified above.
Disclosure. No third party reviewed, funded or directed this work. Claude, made by Anthropic, assisted the original analysis. OpenAI GPT-6 assisted the correction. Both providers compete with model providers discussed in the research. Model assistance does not constitute independent acceptance.
Reproducibility. The embedded tables are regenerated from full_scores.parquet and the analysis outputs by build_tool.py. The full pipeline, the pinned shard manifest and the write-up are in the Zenodo record and the GitHub repository. Change the data, run the build, the page moves. · More from NM AI Research