Precision and recall are the two numbers that tell you whether a classification model is actually useful, not just accurate. Precision asks: when the model says “yes”, how often is it right? Recall asks: of all the real “yes” cases, how many did the model find? This guide explains both metrics with a worked example, shows why accuracy can mislead you, and gives a practical way to choose the right balance for your own project.
Precision and recall in one minute
Every prediction from a binary classifier lands in one of four boxes. The scikit-learn documentation lays these out as a confusion matrix:
- True positive (TP): the case was positive and the model said positive.
- False positive (FP): the case was negative but the model said positive. A false alarm.
- False negative (FN): the case was positive but the model said negative. A miss.
- True negative (TN): the case was negative and the model said negative.
From those four counts come the two metrics. In the scikit-learn definitions, precision is TP ÷ (TP + FP), the ability of the classifier not to label a negative sample as positive. Recall is TP ÷ (TP + FN), the ability of the classifier to find all the positive samples. Recall is also called sensitivity.
A worked example you can check by hand
Imagine a model that flags support tickets as “urgent”. You test it on 100 tickets, and 20 of them really are urgent. The model flags 16 tickets:
- 12 of the flagged tickets really are urgent (true positives).
- 4 of the flagged tickets are not urgent (false positives).
- 8 urgent tickets were not flagged (false negatives).
- The remaining 76 tickets were correctly left alone (true negatives).
Now calculate:
- Precision = 12 ÷ (12 + 4) = 0.75. Three out of four flags were correct.
- Recall = 12 ÷ (12 + 8) = 0.60. The model found 60% of the urgent tickets.
- Accuracy = (12 + 76) ÷ 100 = 0.88.
An accuracy of 88% sounds good, but the model still misses two in every five urgent tickets. That gap is exactly what precision and recall are designed to expose.
Why accuracy alone can fool you
Accuracy counts every correct prediction equally, so it rewards a model for getting the common class right. When one class is rare, a model can score highly by ignoring it. Google’s Machine Learning Crash Course gives the classic illustration: on a heavily imbalanced dataset, a model that always predicts the negative class can post a very high accuracy while being useless. The course advises avoiding accuracy as the main metric for imbalanced datasets and looking at precision, recall or F1 instead.
Fraud detection, defect spotting, spam filtering and lead scoring all tend to have this shape: the interesting cases are a small minority. If you only report accuracy for problems like these, you will not see the failures that matter.
The precision and recall trade-off
Most classifiers output a score or probability, and you choose a threshold above which a case counts as positive. Moving that threshold moves the two metrics in opposite directions. Google’s course explains that raising the threshold usually reduces false positives, which tends to lift precision, while increasing false negatives, which lowers recall. Lowering the threshold does the reverse.
In the ticket example, lowering the threshold might catch 17 of the 20 urgent tickets, but it would probably also flag more routine tickets by mistake. You trade cleaner flags for fewer misses, or the other way round. There is rarely a setting that maximises both.
How to pick a threshold in practice
- Score a held-out validation set with the model’s raw probabilities, not just its yes/no labels.
- Sweep the threshold from low to high and record precision and recall at each step. In scikit-learn,
precision_recall_curvedoes this for you. - Plot the curve and mark where your business requirement sits, for example “recall must be at least 0.9”.
- Choose the threshold that meets the requirement with the best value of the other metric.
- Confirm on a separate test set so you are not tuning to noise in the validation data.
When to favour precision, and when to favour recall
The right balance depends on what each kind of mistake costs you. Google’s course frames it this way: favour recall when false negatives are more expensive than false positives, and favour precision when it is very important that positive predictions are correct.
- Favour precision when a false alarm is costly or annoying. Examples: an automated email that tells a customer their account is suspended, a moderation system that hides posts, or a sales tool that sends reps to call leads. Every wrong flag costs time or trust.
- Favour recall when a miss is costly. Examples: catching security incidents, finding defective parts before shipping, or routing urgent support tickets. Here it is usually cheaper to review a few extra false alarms than to let a real case slip through.
- Balance both when the costs are similar, or when you need one summary number for comparison.
A useful habit is to write the cost of each error type down in plain words before training anything. That sentence often decides the metric for you.
F1 score: one number that balances both
When you want a single figure, the F1 score combines precision and recall using their harmonic mean. Scikit-learn describes the general F-beta measure as a weighted harmonic mean of the two, where F1 weights them equally. For the ticket model, F1 = 2 × (0.75 × 0.60) ÷ (0.75 + 0.60), which is about 0.67.
The harmonic mean punishes imbalance: a model with precision 0.95 and recall 0.10 gets a low F1, even though one number looks impressive. If one error type matters more, use F-beta instead. A beta above 1 leans towards recall, and a beta below 1 leans towards precision.
Computing the metrics in Python
Scikit-learn makes these calculations a few lines of code. A typical evaluation looks like this:
confusion_matrix(y_true, y_pred)returns the four counts so you can see where errors fall.precision_score,recall_scoreandf1_scorereturn the headline metrics.classification_reportprints precision, recall and F1 for every class in one table, which is the quickest way to spot a class the model is ignoring.precision_recall_curvereturns the values you need to plot the trade-off and choose a threshold.
For multi-class problems, check the average parameter. Macro averaging treats every class equally, which highlights weak performance on small classes; weighted averaging reflects class sizes, which can hide it.
Common mistakes to avoid
- Reporting metrics on training data. Always evaluate on data the model has not seen, or the numbers will look better than reality.
- Using the default 0.5 threshold without checking it. The default is a convention, not a recommendation for your problem.
- Comparing models at different thresholds. Compare curves, or fix the threshold rule before comparing.
- Ignoring class balance in the test set. If the test set does not reflect real-world proportions, precision in particular will not carry over to production.
- Forgetting to monitor after launch. Data drifts. Re-check precision and recall on fresh labelled samples on a regular schedule.
Uneven error rates across groups of people are also worth checking: a model can have acceptable overall recall while missing far more cases for one group than another. Our AI ethics and bias hub covers ways to break metrics down by group, and the machine learning hub has more guides on training and evaluating models.
Key takeaways
Precision tells you how trustworthy the model’s positive predictions are, and recall tells you how many real positives it finds. Accuracy hides both when one class is rare, so report precision and recall, or F1, for any imbalanced problem. Decide which mistake costs more, sweep the threshold on validation data to meet that requirement, confirm the result on a separate test set, and keep measuring after the model goes live.
Written by the TechZone AI Editorial desk.


Leave feedback about this