Post-Training & Alignment

RLHF, DPO, and GRPO: the reward models and rollout loops that turn a base model into a product.

5 logs in this box