intuitively makes sense, by rewriting training data to optimize for surprise only where it matters the backprop is given much cleaner signal. with this line of thinking one could expect GRPO to not work that well when using N human provided samples instead of on policy generated
SFT is not dead! 🥳
We found a way to make SFT rival current prevailing posttraining methods, often generalizing better and forgetting less than RL and OPSD. 🤯
Following our prior work on reasoning with sampling, we now introduce sampling to the posttraining stack.
1/n