You Get What You Reward
Train a four-response policy with real gradient updates. Reward outcomes, valid steps or length, and watch what becomes more likely.
From Chapter 14: Reasoning Models
Try it · toy model
You Get What You Reward
Know well8 min
Read the concepts