Listwise Ranking
- A Learning to Rank family that looks at the whole list and optimizes a list-level ranking metric like NDCG.
- Best of the three: directly optimizes what you actually measure, and can weight the top of the list more heavily than Pairwise Ranking.
The catch
- Ranking metrics like NDCG are based on sorting, a step function. A tiny score change either flips two items or does not, so the metric is non-differentiable. Plain gradient descent cannot use it.
The trick: LambdaRank / LambdaMART
- Do not differentiate the metric. Define the gradient directly ("lambda").
- Take the pairwise gradient and scale it by how much swapping that pair would change NDCG. A swap at the top gets a big gradient, a swap at the bottom a tiny one.
- LambdaMART = LambdaRank + gradient-boosted trees. Dominant LTR method for years, still a strong baseline.