Skip to content
Road to Intelligence

You Get What You Reward

Train a four-response policy with real gradient updates. Reward outcomes, valid steps or length, and watch what becomes more likely.

From Chapter 14: Reasoning Models

Try it · toy model

You Get What You Reward

Know well8 min