Area Under Curve (AUC)
- AUC score is the area under the ROC Curve
- AUC score is used to compare multiple classifiers
- Greater the AUC score, better the model
- AUC more than 0.5 is better than random classifier
- Probabilistic meaning (most important): AUC = probability that a randomly chosen positive is scored higher than a randomly chosen negative (the Mann-Whitney U statistic). So AUC is a ranking metric: it depends only on score order, not score values.
- It averages over all positive-negative pairs with equal weight, and is position-blind. A model can have high AUC yet order the top of the list poorly, which is why ranking uses NDCG / Learning to Rank instead.
- AUC is invariant to class ratio (TPR uses only positives, FPR uses only negatives, so changing the mix doesn't move the curve). But it is not fully "robust": under heavy imbalance it can be optimistic (FPR's large negative denominator hides many false positives), so prefer Area Under Precision Recall Curve (AUPRC) when positives are rare.
- AUC Score is calculated from True Positive Rate (Sensitivity) and False Positive Rate (1-Specificity)
- For different threshold, the TPR and FPR is plotted on the graph
- All the point will be connected by a line
- And the score is the Area Under the Curve of that line
- Range =
- This score from multiple model can be used to find out the best model, i.e., the more AUC score means the model is more correct or better.
- Another alternative is Area Under Precision Recall Curve (AUPRC)
- If you need to find best threshold for one model, then ROC Curve, Precision Recall Curve (PRC)
How the ROC curve is produced
- The model outputs a score per item. Pick a threshold: score โฅ threshold โ predict positive.
- Every threshold gives one (FPR, TPR) point. Sweep threshold high โ low to trace the curve.
- threshold = 1.0 (strict): predict nothing โ point (0, 0).
- threshold = 0.0 (loose): predict everything โ point (1, 1).
- Curve = a walk down the score-sorted list. For each item, going down:
- hit a positive โ step up (TPR rises).
- hit a negative โ step right (FPR rises).
Example (5 items, 3 pos / 2 neg, sorted by score)
| item | label | score |
|---|---|---|
| A | P | 0.9 |
| B | P | 0.8 |
| C | N | 0.6 |
| D | P | 0.4 |
| E | N | 0.2 |
Sweep threshold just below each score (denominators: pos=3, neg=2):
| below | predicted + | TP | FP | TPR | FPR |
|---|---|---|---|---|---|
| 0.9 | A | 1 | 0 | 0.33 | 0.0 |
| 0.8 | A,B | 2 | 0 | 0.67 | 0.0 |
| 0.6 | A,B,C | 2 | 1 | 0.67 | 0.5 |
| 0.4 | A,B,C,D | 3 | 1 | 1.00 | 0.5 |
| 0.2 | all | 3 | 2 | 1.00 | 1.0 |
Why area = P(pos ranked above neg)
- At each rightward (negative) step, the curve height = TPR = fraction of positives already passed = fraction of positives ranked above this negative.
- Area = average over all negatives of that fraction = fraction of (pos, neg) pairs ordered correctly = the ranking probability above.
- Check: correctly-ordered pairs here = 5 of 6 โ AUC = 0.833, and the staircase area = 0.833. Same number.
Pseudocode (the definition is the algorithm)
def auc(scores, labels):
pos = [s for s, y in zip(scores, labels) if y == 1]
neg = [s for s, y in zip(scores, labels) if y == 0]
wins = 0.0
for p in pos:
for n in neg:
if p > n: wins += 1.0
elif p == n: wins += 0.5 # tie = half credit
return wins / (len(pos) * len(neg))
- Brute force = O(PยทN). Fast version sorts by score and uses ranks (Mann-Whitney U) = O(n log n). Same value.

