Mark RadarMARK RADAR
EN

Hugging Face Releases Largest Open Code Dataset, The Stack v3

1 reports · First detected 2026-07-24 · Last active 2026-07-24

Hugging Face and BigCode developed The Stack series to make the training of code-focused large language models more open, reproducible and transparent. The latest version embeds source files directly and preserves full-repository context, giving researchers infrastructure for next-generation open code models and cyber-defense tools. Its release also lands amid growing debate over licensing, attribution and the acceptable boundaries of distilling capabilities from existing AI models.

Hugging Face released The Stack v3 on July 23, 2026. The full corpus contains 113.7 terabytes of code from 224 million GitHub repositories across 770 languages. A filtered, near-deduplicated training set totals 15.9 terabytes and about 4.9 trillion tokens, covering 173 million repositories and 713 languages. The dataset reflects public GitHub repositories as of August 7, 2025, and is distributed under the Open Data Commons Attribution License.

All Coverage

1 original reports
NEWS.SMOL.AI 2026-07-23
not much happened today

The Backstory

The history behind this event

No historical echoes for this signal

Mark Radar|MARK RADAR
All times are in Taipei time (GMT+8)