Work
Things I’ve built and written up.
- Running RL on Academic budget at high-end speed
Some Optimization tricks to run RL on academic (not-so)-low-end hardware.
- Baseline Estimation Bounds For GRPO
How much variance reduction is actually available from a better baseline in GRPO?