Hugging Face Releases Largest Open Code Dataset, The Stack v3
Hugging Face and BigCode developed The Stack series to make the training of code-focused large language models more open, reproducible and transparent. The latest version embeds source files directly and preserves full-repository context, giving researchers infrastructure for next-generation open code models and cyber-defense tools. Its release also lands amid growing debate over licensing, attribution and the acceptable boundaries of distilling capabilities from existing AI models.
Hugging Face released The Stack v3 on July 23, 2026. The full corpus contains 113.7 terabytes of code from 224 million GitHub repositories across 770 languages. A filtered, near-deduplicated training set totals 15.9 terabytes and about 4.9 trillion tokens, covering 173 million repositories and 713 languages. The dataset reflects public GitHub repositories as of August 7, 2025, and is distributed under the Open Data Commons Attribution License.
All Coverage
1 original reportsThe Backstory
The history behind this eventNo historical echoes for this signal
Subscribe to Mark Radar Weekly
Every Friday, the week's strongest signals in your inbox. Unsubscribe anytime.