Denser rewards, nearly free: contrastive potential shaping
Add contrastive prefix-level credit assignment to outcome-based RL
Add contrastive prefix-level credit assignment to outcome-based RL
Evolution of reinforcement learning for reasoning LLMs
Showing that $0\neq 0$
Step-by-step exploration of nanochat's model
Close examination of gradients of DPO and SFT