Mark RadarMARK RADAR
About
EN
Sign in

Ai2 Open-Sources Tulu 3 Post-Training Pipeline

1 reports · First detected 2026-08-13 · Last active 2026-08-13

The Allen Institute for AI, or Ai2, built Tulu 3 on Meta’s Llama 3.1 base models and released the data, weights, training recipes and evaluation tools. The project matters because post-training — the stage that turns a base model into a useful assistant — is often less transparent than pretraining. Tulu 3 sets out a reproducible sequence of supervised fine-tuning, or SFT, Direct Preference Optimization, or DPO, and Reinforcement Learning with Verifiable Rewards, or RLVR, to improve instruction following, preference alignment and reasoning.

Ai2 released the full Tulu 3 stack on Nov. 22, 2024. The latest tutorial reproduces the three-stage pipeline in Open Instruct, adds Group Relative Policy Optimization and verifier-based evaluation, and pares a multi-GPU setup down for a roughly 16GB single-machine runtime. Ai2 said RLVR improved scores over the DPO checkpoint by as much as 1.7 points on MATH, 3.3 on GSM8K and 1.3 on IFEval. It followed on Feb. 12, 2025, with an 8-billion-parameter GRPO model that beat the earlier Tulu 3 8B version on almost all internal evaluations.

All Coverage

1 original reports

The Backstory

The history behind this event

No historical echoes for this signal

Mark Radar|MARK RADAR
All times are in Taipei time (GMT+8)