Retro AI Model talkie Tests Reasoning Using Only Pre-1930 Data
Large language models blend modern internet content with historical texts, often making it difficult to distinguish reasoning from memorization. Researchers Nick Levine and David Duvenaud, along with former OpenAI scientist Alec Radford, created talkie to address that problem. Using a public-domain collection assembled by the Institutional Data Initiative at Harvard Law School Library, they applied a “time-slice” approach to examine how knowledge limited to a particular era shapes a model's capabilities.
The team released talkie-1930-13b-base in April 2026. The 13-billion-parameter model was trained on 260 billion English-language tokens dating from before January 1, 1931, with a model using the same architecture but trained on modern data serving as a control. Tests showed comparable performance in language comprehension and basic mathematics. Although talkie lagged significantly on knowledge questions, excluding anachronistic questions narrowed the gap by about half, indicating that knowledge gaps accounted for most of its weakness.
All Coverage
2 original reportsThe Backstory
The history behind this eventNo historical echoes for this signal
Subscribe to Mark Radar Weekly
Every Friday, the week's strongest signals in your inbox. Unsubscribe anytime.
If you search news on Google, you can set Mark Radar as a preferred source—our coverage will show up more often in your results. Set as preferred source on Google →