{"ok":true,"source":"tensorfeed.ai","lastUpdated":"2026-09-14","count":29,"datasets":[{"id":"fineweb","name":"FineWeb","publisher":"Hugging Face","stage":"pretraining","contentType":"web text","tokens":"18.5T","items":"25.9B documents","license":"ODC-BY-1.0","languages":"English","released":"2024-04","url":"https://huggingface.co/datasets/HuggingFaceFW/fineweb","notes":"Filtered and deduplicated CommonCrawl, processed with the datatrove library. Launched at 15T gpt2 tokens; v1.4.0 (July 2025) added the January to June 2025 snapshots and the card now lists more than 18.5T tokens. Some domains were removed in January 2025 in response to a cease-and-desist notice."},{"id":"fineweb-edu","name":"FineWeb-Edu","publisher":"Hugging Face","stage":"pretraining","contentType":"educational web text","tokens":"1.3T","items":"1.5B documents","license":"ODC-BY-1.0","languages":"English","released":"2024-05","url":"https://huggingface.co/datasets/HuggingFaceFW/fineweb-edu","notes":"FineWeb subset kept by an educational-quality classifier trained on Llama-3-70B-Instruct annotations. A looser threshold variant (FineWeb-Edu-score-2) holds 5.4T tokens. Hugging Face ablations report it outperforming full FineWeb on knowledge and reasoning benchmarks such as MMLU and ARC."},{"id":"common-crawl","name":"Common Crawl","publisher":"Common Crawl Foundation","stage":"pretraining","contentType":"web text","tokens":"petabyte-scale raw","items":"300B+ webpages","license":"Common Crawl Terms of Use (limited license; crawled content keeps third-party rights)","languages":"multilingual (200+)","released":"ongoing, 3-5B new pages added monthly","url":"https://commoncrawl.org","notes":"The raw web crawl most open web pretraining datasets are built on. Not a public-domain dedication: the terms grant a limited license to the service and require users to respect the copyrights of the crawled material. Most labs filter it heavily before training."},{"id":"redpajama-v2","name":"RedPajama v2","publisher":"Together AI","stage":"pretraining","contentType":"web text","tokens":"30T","items":"100B+ documents","license":"Common Crawl Terms of Use (data); Apache-2.0 (code)","languages":"English, German, French, Spanish, Italian","released":"2023-10","url":"https://huggingface.co/datasets/togethercomputer/RedPajama-Data-V2","notes":"Built from 84 CommonCrawl snapshots with the CCNet pipeline. About 30B documents carry pre-computed quality signals and duplicate ids are published, so labs can apply their own filtering thresholds; the deduplicated head and middle buckets total roughly 30T tokens."},{"id":"the-pile","name":"The Pile","publisher":"EleutherAI","stage":"pretraining","contentType":"mixed (web + books + papers + code)","tokens":"825GB / 300B","items":"22 sub-corpora","license":"mixed (per-subcorpus)","languages":"English-focused","released":"2020","url":"https://pile.eleuther.ai","notes":"The original open pretraining corpus, now mostly of historical interest. The Books3 subset was taken offline in 2023 after DMCA notices from the Danish Rights Alliance, and the original download host no longer serves the full corpus. EleutherAI released the Common Pile v0.1 in 2025 as its openly licensed successor."},{"id":"common-pile-v0.1","name":"Common Pile v0.1","publisher":"EleutherAI (with University of Toronto, Vector Institute, Hugging Face, AI2, and others)","stage":"pretraining","contentType":"mixed (papers, code, books, government text, educational materials, transcripts)","tokens":"8TB","items":"30 sources","license":"public domain and openly licensed text (per-source licenses)","languages":"English-focused","released":"2025-06","url":"https://huggingface.co/common-pile","notes":"Successor to The Pile built only from public-domain and openly licensed text, including an openly licensed subset of The Stack v2 and about 300,000 public-domain books. Released alongside Comma v0.1-1T and Comma v0.1-2T, 7B models trained on 1T and 2T tokens of it."},{"id":"dolma","name":"Dolma","publisher":"Allen AI (AI2)","stage":"pretraining","contentType":"web + code + books + papers","tokens":"3T","items":"mixed","license":"ODC-BY-1.0","languages":"English","released":"2024-04","url":"https://huggingface.co/datasets/allenai/dolma","notes":"AI2's open pretraining corpus behind the original OLMo models, with documented source provenance. Relicensed from the AI2 ImpACT license to ODC-BY in April 2024, when v1.7 became the default version. Succeeded by the Dolma 3 mix used for OLMo 3."},{"id":"dolma-3-mix","name":"Dolma 3 Mix (6T)","publisher":"Allen AI (AI2)","stage":"pretraining","contentType":"web + papers + code","tokens":"6T","items":null,"license":"ODC-BY-1.0","languages":"English","released":"2025-11","url":"https://huggingface.co/datasets/allenai/dolma3_mix-6T","notes":"The pretraining mix used for Olmo-3-1125-32B, mostly drawn from Common Crawl. AI2 also publishes a 150B-token sample with the same upsampling strategy for smaller experiments."},{"id":"fineweb-2","name":"FineWeb2","publisher":"Hugging Face","stage":"pretraining","contentType":"multilingual web text","tokens":"20TB / 3T+ words","items":"5B documents","license":"ODC-BY-1.0","languages":"1000+ languages","released":"2024-12","url":"https://huggingface.co/datasets/HuggingFaceFW/fineweb-2","notes":"Multilingual successor to FineWeb built from 96 CommonCrawl snapshots (2013 to April 2024). The card reports words rather than tokens because token counts vary widely by tokenizer and script. v2.1.0 (June 2025) aligned filtering with the published paper and grew the dataset."},{"id":"finepdfs","name":"FinePDFs","publisher":"Hugging Face","stage":"pretraining","contentType":"text extracted from PDFs","tokens":"3T","items":"475M documents","license":"ODC-BY-1.0","languages":"1733 language-script pairs","released":"2025-09","url":"https://huggingface.co/datasets/HuggingFaceFW/finepdfs","notes":"Described by its publisher as the largest public corpus sourced only from PDFs. Fully reproducible, with extraction code released alongside the data."},{"id":"finetranslations","name":"FineTranslations","publisher":"Hugging Face","stage":"pretraining","contentType":"parallel text (web text translated into English)","tokens":"1T+","items":null,"license":"ODC-BY-1.0","languages":"English paired with 500+ languages","released":"2026-01","url":"https://huggingface.co/datasets/HuggingFaceFW/finetranslations","notes":"FineWeb2 documents translated into English with Gemma 3 27B using a synthetic-data pipeline in datatrove. An educational-filtered variant is published separately."},{"id":"ultra-fineweb-l3","name":"Ultra-FineWeb-L3","publisher":"OpenBMB","stage":"pretraining","contentType":"synthetic web-derived text (Q&A pairs, multi-style rewrites)","tokens":"600B+ (400B+ English, 200B+ Chinese)","items":null,"license":"Apache-2.0","languages":"English, Chinese","released":"2026-02","url":"https://huggingface.co/datasets/openbmb/Ultra-FineWeb-L3","notes":"Refined tier of the UltraData framework, generated from Ultra-FineWeb with MiniCPM4 and Qwen3 and used in the decay phase of MiniCPM5-1B. The L1 filtered tier (1T+ tokens) followed in August 2026."},{"id":"refinedweb","name":"RefinedWeb","publisher":"TII","stage":"pretraining","contentType":"web text","tokens":"500-650B (public extract)","items":"968M documents (public extract)","license":"ODC-BY-1.0","languages":"English","released":"2023-06","url":"https://huggingface.co/datasets/tiiuae/falcon-refinedweb","notes":"Behind the Falcon series. Only a public extract of the full corpus was released; use is also subject to the CommonCrawl terms. Aggressive deduplication and filtering; older than FineWeb."},{"id":"the-stack-v3","name":"The Stack v3","publisher":"Hugging Face","stage":"pretraining","contentType":"source code (grouped by repository)","tokens":"3.6T (train subset)","items":"173M repositories (train subset)","license":"ODC-BY-1.0 (files keep their original repository licenses)","languages":"713 programming languages (train subset)","released":"2026-07","url":"https://huggingface.co/datasets/HuggingFaceCode/stack-v3-train","notes":"Successor to The Stack v2, crawled from GitHub default branches as of August 2025 with file contents inline. The train subset is 15.9TB; v3.1 removed leaked exact duplicates, cutting it from 4.9T to 3.6T tokens. The full corpus is 113.7TB across 770 languages and 224M repositories, and developers can opt out."},{"id":"the-stack-v2","name":"The Stack v2","publisher":"BigCode (HuggingFace + ServiceNow)","stage":"pretraining","contentType":"source code","tokens":"67.5TB full / ~900B (train-full)","items":"3.28B unique files from 104.2M repositories","license":"original repository licenses (gated terms of use)","languages":"600+ programming languages","released":"2024-02","url":"https://huggingface.co/datasets/bigcode/the-stack-v2","notes":"Sourced from the Software Heritage archive and deduplicated at file and near-duplicate level (the dedup version is 32.1TB). Behind StarCoder 2. Access is gated and use must follow the original licenses. Succeeded by The Stack v3 in 2026."},{"id":"starcoderdata","name":"StarCoderData","publisher":"BigCode","stage":"pretraining","contentType":"source code + jupyter + GitHub issues + commits","tokens":"~250B","items":"86 languages","license":"mixed permissive","languages":"86 programming languages","released":"2023-05","url":"https://huggingface.co/datasets/bigcode/starcoderdata","notes":"Pre-cleaned subset of The Stack used to train StarCoder: 783GB of code plus GitHub issues, Jupyter notebooks, and commits. Smaller than The Stack v2 and easier to work with for small-scale code-model experiments."},{"id":"tulu-3-sft-mix","name":"Tulu 3 SFT Mixture","publisher":"Allen AI","stage":"instruction-tuning","contentType":"instruction-response pairs","tokens":null,"items":"939K instructions","license":"ODC-BY-1.0","languages":"English","released":"2024-11","url":"https://huggingface.co/datasets/allenai/tulu-3-sft-mixture","notes":"SFT data for AI2's Tulu 3 post-training recipe. Combines FLAN v2, OpenAssistant, WildChat, Aya, NuminaMath, persona-generated math, code, and instruction-following prompts, and safety sets. Some subsets carry non-commercial licenses."},{"id":"open-hermes-2.5","name":"OpenHermes 2.5","publisher":"Teknium / Nous Research","stage":"instruction-tuning","contentType":"multi-turn conversations","tokens":null,"items":"1M conversations","license":"mixed (per-source)","languages":"English","released":"2023-12","url":"https://huggingface.co/datasets/teknium/OpenHermes-2.5","notes":"Compiled from GPT-4 outputs, Airoboros, ShareGPT, and other open and synthetic sets. The dataset behind OpenHermes 2.5, Nous Hermes 2, and many community fine-tunes."},{"id":"openthoughts3","name":"OpenThoughts3-1.2M","publisher":"Open Thoughts","stage":"instruction-tuning","contentType":"reasoning traces (math, code, science)","tokens":null,"items":"1.2M examples (850K math, 250K code, 100K science)","license":"Apache-2.0","languages":"English","released":"2025-06","url":"https://huggingface.co/datasets/open-thoughts/OpenThoughts3-1.2M","notes":"Third OpenThoughts release, built through an ablation-driven pipeline over question sourcing, selection, and answer generation. Used to train OpenThinker3-7B and as part of the SmolLM3 mid-training mix."},{"id":"glaive-function-calling","name":"Glaive Function Calling v2","publisher":"Glaive AI","stage":"instruction-tuning","contentType":"function-calling traces","tokens":null,"items":"113K examples","license":"Apache-2.0","languages":"English","released":"2023-12","url":"https://huggingface.co/datasets/glaiveai/glaive-function-calling-v2","notes":"Synthetic function-calling traces. A common starting point for fine-tuning open models on JSON tool use."},{"id":"open-orca","name":"OpenOrca","publisher":"Open Orca team","stage":"instruction-tuning","contentType":"reasoning traces","tokens":null,"items":"~1M GPT-4 + ~3.2M GPT-3.5 completions","license":"MIT","languages":"English","released":"2023-07","url":"https://huggingface.co/datasets/Open-Orca/OpenOrca","notes":"Reproduction of the Microsoft Orca paper. Augments FLAN tasks with GPT-4 and GPT-3.5 step-by-step reasoning traces. Strong base for math and reasoning fine-tunes."},{"id":"agent-instruct","name":"AgentInstruct (orca-agentinstruct-1M-v1)","publisher":"Microsoft Research","stage":"instruction-tuning","contentType":"synthetic instruction pairs","tokens":null,"items":"~1M instruction pairs (public subset of ~25M)","license":"CDLA-Permissive-2.0","languages":"English","released":"2024-10","url":"https://huggingface.co/datasets/microsoft/orca-agentinstruct-1M-v1","notes":"Fully synthetic data generated with the AgentInstruct agentic framework from public web text seeds: text editing, creative writing, code, reading comprehension, RAG, brain teasers, and more. The full ~25M-pair set was used to post-train Orca-3-Mistral; Microsoft shares the 1M subset for research."},{"id":"ultrafeedback","name":"UltraFeedback","publisher":"OpenBMB","stage":"dpo","contentType":"preference pairs","tokens":null,"items":"64K instructions x 4 model responses","license":"MIT","languages":"English","released":"2023-10","url":"https://huggingface.co/datasets/openbmb/UltraFeedback","notes":"Widely used DPO dataset for open models. GPT-4 rates 4 candidate responses per prompt across helpfulness, honesty, instruction-following, and truthfulness."},{"id":"tulu-3-pref","name":"Tulu 3 Preference Mixture (70B)","publisher":"Allen AI","stage":"dpo","contentType":"preference pairs","tokens":null,"items":"337K preference pairs","license":"ODC-BY-1.0","languages":"English","released":"2024-11","url":"https://huggingface.co/datasets/allenai/llama-3.1-tulu-3-70b-preference-mixture","notes":"Preference pairs for the Tulu 3 70B DPO stage, with responses generated from a pool of open models including Tulu 2 and InternLM2.5. Some portions carry non-commercial terms."},{"id":"helpsteer-2","name":"HelpSteer 2","publisher":"NVIDIA","stage":"dpo","contentType":"preference pairs (5-attribute)","tokens":null,"items":"21K conversations","license":"CC-BY-4.0","languages":"English","released":"2024-06","url":"https://huggingface.co/datasets/nvidia/HelpSteer2","notes":"NVIDIA preference data labeled across helpfulness, correctness, coherence, complexity, verbosity. Strong for multi-attribute reward modeling. Followed by HelpSteer3 in 2025."},{"id":"helpsteer-3","name":"HelpSteer3","publisher":"NVIDIA","stage":"dpo","contentType":"preference pairs + feedback + edits","tokens":null,"items":"40K preference samples (plus Feedback and Edit subsets)","license":"CC-BY-4.0","languages":"multilingual (general, STEM, code, and multilingual domains)","released":"2025-03","url":"https://huggingface.co/datasets/nvidia/HelpSteer3","notes":"Human-annotated preferences over responses from about 20 commercially permissive open models, with no outputs from proprietary providers. The Feedback and Edit subsets support inference-time scaling research."},{"id":"laion-5b","name":"LAION-5B (Re-LAION-5B)","publisher":"LAION","stage":"multimodal","contentType":"image-text pairs (links + alt text)","tokens":null,"items":"~5.8B image-text pairs","license":"Apache-2.0 (Re-LAION-5B)","languages":"multilingual","released":"2024-08","url":"https://laion.ai/blog/relaion-5b/","notes":"Behind Stable Diffusion and many open image models. The original 2022 release was withdrawn in December 2023 after a Stanford Internet Observatory report found links to CSAM. Re-LAION-5B (August 2024) removed 2,236 flagged links and ships gated on Hugging Face in research and research-safe versions."},{"id":"datacomp-1b","name":"DataComp-1B","publisher":"DataComp","stage":"multimodal","contentType":"image-text pairs","tokens":null,"items":"1.4B filtered pairs","license":"CC-BY-4.0 (url-text metadata; images keep their own copyrights)","languages":"multilingual","released":"2023","url":"https://www.datacomp.ai","notes":"Filtered subset of CommonPool. Reproducible filtering recipes are part of the contribution. Behind several open CLIP-class models."},{"id":"finevision","name":"FineVision","publisher":"Hugging Face","stage":"multimodal","contentType":"image-text instruction data for vision-language models","tokens":"9.5B answer tokens","items":"24.3M samples, 17.3M images","license":"per sub-dataset (prompts CC-BY-4.0)","languages":"per sub-dataset","released":"2025-09","url":"https://huggingface.co/datasets/HuggingFaceM4/FineVision","notes":"Unifies more than 200 public sources into 185 subsets with 88.9M turns, deduplicated and decontaminated against 66 benchmarks, including agentic and GUI tasks. Each source keeps its own license."}]}