Yoav Kor

Baseline Estimation Bounds For GRPO

PDF

GRPO assigns every token in a trajectory the same advantage, using the group mean as a baseline. Control variate theory says the variance-minimising baseline is state-dependent — roughly the value function Vπ(s)V^\pi(s) — so a per-state baseline should give a finer-grained signal and a less noisy gradient estimator.

This note derives anupper bound on the available improvement and then emprically measures it on a very specific settings.

← Work