Baseline Estimation Bounds For GRPO
GRPO assigns every token in a trajectory the same advantage, using the group mean as a baseline. Control variate theory says the variance-minimising baseline is state-dependent — roughly the value function — so a per-state baseline should give a finer-grained signal and a less noisy gradient estimator.
This note derives anupper bound on the available improvement and then emprically measures it on a very specific settings.