Cookbook · LLMs
Data engineering
4 mindata-engineeringdeduplicationminhashpretraining
The idea, in one analogy
If you learned to write by reading a library where every tenth book was a spam email and half the books had entire chapters duplicated, you'd waste a huge amount of your reading budget on noise, and you'd overlearn whatever got repeated most. A language model has exactly the same problem: raw scraped text is full of junk, duplication, and boilerplate, and every token spent on it is a token not spent learning something useful. Data engineering is the filtering pass that fixes this before training ever starts.
The usual pipeline
Language identification comes first. Web scrapes are multilingual by default, so if you're training a single-language model, everything else gets filtered out (a fastText classifier is the common tool for this).
Next, quality heuristics: cheap, rule-based filters. Documents that are mostly boilerplate, too short, too repetitive, or fail basic "does this look like real prose" checks get dropped. Not perfect, but cheap, and it catches a lot of the worst material. A few concrete examples of what typically gets dropped versus kept:
| Drop | Keep |
|---|---|
| "Buy cheap Rolex watches free shipping click here now !!!" | "The paper describes a new approach to federated learning that reduces communication overhead..." |
| "Lorem ipsum dolor sit amet consectetur adipiscing elit sed do eiusmod..." | "Scientists discovered that the protein responsible for cellular repair also plays a role in..." |
| " Copyright © 2018 All Rights Reserved | Privacy Policy" | |
| "This article is the same as this article. This article is the same as this article." |
The pattern across the "drop" column: spam, template boilerplate, and repetition. None of it is grammatically broken, which is exactly why a rule-based filter (not a grammar checker) is what catches it.
Then deduplication. Exact-hash dedup removes byte-identical documents (surprisingly common in web scrapes: the same article gets mirrored across many domains). Near-duplicate detection (commonly MinHash plus locality-sensitive hashing) catches documents that are almost identical but not byte-for-byte, which exact hashing misses entirely.
Last, re-tokenization. If you're training your own tokenizer (see tokenization), the filtered corpus gets tokenized with it as the last step before training. Filtering before tokenizing means the tokenizer itself is trained on cleaner data too.
Why this step gets skipped, and why that's a mistake
It's the least glamorous part of the whole pipeline and the one most likely to get rushed. It's also frequently the highest-leverage step available: a smaller model trained on well-filtered data routinely beats a larger model trained on raw scrape, because so much of a raw scrape's tokens are actively harmful (duplication makes the model overfit toward repeated content), not merely neutral.
There's also an easy way to fool yourself here. If you change your data pipeline and your corpus size changes at the same time, you can't tell whether an improvement came from cleaner data or just having more of it. Any honest before/after comparison needs to hold corpus size fixed, or explicitly name that it didn't, rather than quietly crediting "the new pipeline is better."
Where to look further
- The FineWeb technical report: a detailed, empirical walkthrough of this pipeline at scale, with ablations showing what each filtering step is worth.
datatrove: Hugging Face's library for large-scale text filtering/dedup pipelines.- This project's own filtering pipeline and results:
slm-from-scratchPhase 2.