10 notes

Research notes

Notes on papers I wanted to understand properly. Each one says what the paper showed, why I think it might matter outside the lab, and what is still uncertain. Most were written in November 2025 and have short updates where things have moved.

Cheaper, smaller models

Why the cost of a capable model keeps falling, and what that does to the economics.

  • The s1 recipe for making models think longer

    With 1,000 examples and a forced "Wait", s1 got o1-style test-time scaling from an open model. Controlling how long models think is now a standard setting.

    Published result · Paper January 2025

  • What DeepSeek V3 and R1 actually showed

    The famous $5.6 million covered one final training run, not the whole bill. The lasting lessons were efficient engineering and reasoning learned from rewards.

    In wide use · Paper January 2025 + 1 more

  • Running large models in 4 bits

    Storing weights in 4 bits instead of 16 cuts memory about fourfold with little loss. It decides what runs where, and it's now part of how models ship.

    In wide use · Paper October 2022 + 2 more

Learning from generated data and experience

Training models and agents on text other models wrote, or on what happened when they tried things.

  • Training agents in an imagined environment

    DreamGym has a language model imagine how a website would respond, so an agent can practise with reinforcement learning. Real gains on web benchmarks, nothing physical yet.

    Early research · Paper November 2025

  • Training agents on their own early experience

    Let an agent try alternatives to the expert's actions and learn from what happens, with no reward signal. It beat plain imitation in eight environments.

    Published result · Paper October 2025

  • Training models on synthetic data

    Small models trained partly on text written by bigger models can beat much larger ones. Synthetic data can lower privacy risk, but it isn't anonymous by default.

    In wide use · Paper December 2022 + 2 more

New architectures to watch

Ideas that change how models predict or remember. Promising at small scale, unproven at large scale.

  • Nested Learning and models that keep learning

    Google's Nested Learning gives models memories that update at different speeds, so they might keep learning. Promising at 1.3B parameters, unproven beyond.

    Early research · Paper December 2025

  • CALM and predicting several tokens at once

    Tencent's CALM predicts one vector that stands for four tokens, matching a baseline with about a third less compute at small scale. Nobody has shown yet that it scales.

    Early research · Paper October 2025

  • LLM-JEPA and LeCun's bet on predicting meaning

    LeCun argues models should predict meaning rather than words. LLM-JEPA was a first small test on language models, and he has since left Meta to pursue the idea.

    Early research · Paper September 2025

Robots and the physical world

Connecting language models to perception and action, and how far that has really come.

  • Robot foundation models and where they stand

    Vision-language-action models let robots borrow knowledge from the web. The progress since RT-2 is real, but home robots are still mostly pre-orders and pilots.

    Published result · Paper July 2023 + 2 more