GRPO
The model creates several solutions to the same task. Some get better scores, others worse. You want to increase the chances of better behaviors by comparing responses within this group, without building an additional model to estimate their value.
GRPO, Group Relative Policy Optimization, is a method of learning from rewards and a response's relative result within a group. Here a policy means how the model chooses actions or successive response pieces. A reward is a numerical assessment of the result.
If responses have rewards 1, 1, 0, 0, the mean is 0.5. Better results are above the mean, worse ones below it; the method uses this relative signal when updating the model. The actual algorithm includes normalization, limiting changes, and other elements. DeepSeekMath, §4.1, presents GRPO.
TTPO used a particular branch penalizing specific responses. This is not the entire definition of GRPO. Correct learning depends on a meaningful reward: if we assess incorrectly, relative comparison within a group does not repair the criterion by itself.
Connections
Voting lies. Disagreement usually does not
See also: Reinforcement Learning — general learning of a strategy from rewards, of which GRPO is a variant.