Reinforcement Learning

Reinforcement learning is a subfield of machine learning, where input is the features of the current state and output is an optimal action for current timestamp.

This action should maximize the expected average reward.

In LLMs, RL is used for alignment / fine-tuning via policy-gradient methods like Proximal Policy Optimization (PPO), GRPO, GSPO, DAPO, and preference methods like KTO.