Ai2 Open-Sources Tulu 3 Post-Training Pipeline
The Allen Institute for AI, or Ai2, built Tulu 3 on Meta’s Llama 3.1 base models and released the data, weights, training recipes and evaluation tools. The project matters because post-training — the stage that turns a base model into a useful assistant — is often less transparent than pretraining. Tulu 3 sets out a reproducible sequence of supervised fine-tuning, or SFT, Direct Preference Optimization, or DPO, and Reinforcement Learning with Verifiable Rewards, or RLVR, to improve instruction following, preference alignment and reasoning.
Ai2 released the full Tulu 3 stack on Nov. 22, 2024. The latest tutorial reproduces the three-stage pipeline in Open Instruct, adds Group Relative Policy Optimization and verifier-based evaluation, and pares a multi-GPU setup down for a roughly 16GB single-machine runtime. Ai2 said RLVR improved scores over the DPO checkpoint by as much as 1.7 points on MATH, 3.3 on GSM8K and 1.3 on IFEval. It followed on Feb. 12, 2025, with an 8-billion-parameter GRPO model that beat the earlier Tulu 3 8B version on almost all internal evaluations.
All Coverage
1 original reportsThe Backstory
The history behind this eventNo historical echoes for this signal
Subscribe to Mark Radar Weekly
Every Friday, the week's strongest signals in your inbox. Unsubscribe anytime.