The s1 recipe for making models think longer

With 1,000 examples and a forced "Wait", s1 got o1-style test-time scaling from an open model. Controlling how long models think is now a standard setting.

Written Updated 3 min read

The paper

  • s1: Simple test-time scalingMuennighoff et al. (Stanford, University of Washington, Allen Institute for AI, Contextual AI) · EMNLP 2025 · January 2025

OpenAI's o1 showed in late 2024 that a model does better on hard problems when it spends more tokens reasoning before it answers. OpenAI described its method only as large-scale reinforcement learning. A group from Stanford, the University of Washington, the Allen Institute for AI and Contextual AI went looking for the simplest way to get the same behaviour, and published the result as s1.

Their recipe has two parts. The first is a dataset of 1,000 hard questions (s1K), picked from about 59,000 candidates for difficulty, diversity and quality, each paired with a worked reasoning trace generated by Google's Gemini 2.0 Flash Thinking. They fine-tuned Qwen2.5-32B-Instruct on it with ordinary supervised learning, which took 26 minutes on 16 H100 GPUs, about 7 GPU-hours in all.

The second is a decoding trick they call budget forcing. To cut the model's thinking short, you insert the end-of-thinking marker. To extend it, you suppress that marker when the model tries to stop and append "Wait", which often leads it to re-check its work and fix mistakes.

The resulting model, s1-32B, beat OpenAI's o1-preview on competition math (MATH and AIME 2024) by up to 27%, and forcing it to think longer lifted its AIME 2024 score from 50% to 57%. The authors describe this as the first open reproduction of o1's test-time scaling curve. They also report its limits. Gains flatten out once the model has been forced to continue about six times, and suppressing the stop too often sends it into repetitive loops. And DeepSeek's distilled R1 model of the same size scored higher, after training on 800 times more examples.

Descriptions of s1 as frontier AI for $50 left out what it was built on: a strong open base model and reasoning traces from Gemini. The word "Wait" also got over-read. In the paper's comparison, appending "Wait" scored 53.3% on AIME 2024, against 50.0% for "Hmm", "Alternatively" or no string at all. AIME 2024 has 30 problems, so that gap is one question, and on the other two benchmarks "Wait" tied "Hmm". What mattered was forcing more thinking, much more than the choice of word.

Why I think it matters

s1 put test-time scaling within reach of anyone with an open model and a few GPU-hours, which made it easy to experiment with instead of something only frontier labs could study. My bet then, and still, is that reasoning is becoming a commodity capability, so domain expertise, data quality and application design matter more than which model you have access to.

Since then

s1 was published at EMNLP 2025, and the first author lists a best paper award from an ICLR 2025 reasoning workshop (homepage). A follow-up, s1.1, reused the same 1,000 questions with reasoning traces from DeepSeek-R1 instead of Gemini and scored better (GitHub).

Budget control became a product feature within months. Anthropic let developers cap Claude 3.7 Sonnet's thinking at any number of tokens up to 128,000 (February 2025). Google gave Gemini 2.5 Flash a thinking budget from 0 to 24,576 tokens (April 2025), and Qwen3 shipped thinking budget control the same month (Qwen).

By 2026 the vendors had moved from token counts to effort levels, with more of the decision left to the model. Anthropic's February 2026 release notes put it plainly: "Effort replaces budget_tokens for controlling thinking depth on new models" (Claude release notes). Google now recommends thinking levels over budgets (Gemini docs).

Test-time compute also drove the headline results of 2025. Google DeepMind's Gemini Deep Think reached an officially certified gold-medal standard at the International Mathematical Olympiad in July, scoring 35 out of 42 (Google DeepMind).

More thinking doesn't always help, which fits s1's own plateau. Researchers from Anthropic's fellows program found tasks where longer reasoning lowers accuracy (Inverse Scaling in Test-Time Compute, July 2025). How to spend thinking well, rather than simply spending more of it, still looks like an open problem to me.

For the reinforcement-learning route to the same behaviour, see DeepSeek R1. s1K is also a small example of training on model-written data.