What matters before you compare any ratings
- Calibration aligns managers on how each rating level should be applied, so “strong” means the same thing across teams.
- The process works best when managers bring specific evidence, not just general impressions or recent memories.
- Good examples compare similar roles, similar scope, and similar expectations before any final score is agreed.
- Calibration should catch rating inflation, rating deflation, and bias caused by visibility, personality, or recency.
- In the UK, fair appraisals depend on regular documentation, clear standards, and sensible handling of adjustments and role differences.
What calibration is actually fixing
I treat calibration as a standardisation step, not a popularity contest. Its job is to reduce the gap between how different managers interpret the same rating scale. One manager’s “exceeds” might be another manager’s “meets,” and that drift becomes expensive when raises, bonuses, and promotions are involved.
The main problems it solves are usually predictable. Leniency bias pushes ratings too high, stringency bias pushes them too low, and recency bias lets the last six weeks overshadow the rest of the cycle. Calibration also helps when one team has stronger visibility than another, because the most visible employees are not always the highest performers.
- Different standards make the same behaviour look better or worse depending on the manager.
- Uneven evidence means some employees are judged on data while others are judged on memory.
- Role context can distort ratings when one team handles harder clients, larger accounts, or more technical work.
- Confidence bias can reward people who speak well more than people who deliver well.
Once you see those patterns clearly, the examples become much easier to read and use. That is where the process becomes practical rather than abstract.

Practical calibration examples from real review cycles
When I explain calibration, I start with cases that managers recognise immediately. The point is not to memorise a script; it is to see how evidence changes the final call. Here are the kinds of examples that usually matter most in performance management.
| Scenario | Initial rating | Calibrated outcome | Why it matters |
|---|---|---|---|
| A sales manager rates a rep as outstanding because they hit 108% of target. | Exceeds expectations | Still strong, but moved to solid exceeds after comparing territory quality and lead volume. | Raw percentages look impressive, but they are not fair if one rep had a much easier book of business. |
| An engineer is rated average because they are quiet in meetings and do not self-promote. | Meets expectations | Raised to exceeds after reviewing defect reduction, incident handling, and peer support. | Calibration catches visibility bias and shifts the focus back to measurable impact. |
| A support team manager gives almost everyone a top score because the team had a hard quarter. | Mostly outstanding | Split into a clearer spread after looking at first-contact resolution, escalations, and quality checks. | Good intent can still create rating inflation if the team’s performance is not separated carefully. |
| A new hire on probation is compared with employees who have been in role for twelve months. | Below average | Re-rated using a probation-specific view and a shorter evidence window. | People need to be judged against the right timeline, not against a full-year standard they have not had time to reach. |
These examples matter because they show the real job of calibration: not flattening everyone into the middle, but making sure the rating matches the evidence and the context. If the evidence changes the conclusion, the rating should change too. If it does not, the manager should be able to explain why.
The fastest way to use these examples in your own cycle is to ask one question: “Would two managers looking at the same evidence reach the same conclusion?” If the answer is no, the process is not calibrated yet.
How I would run a calibration meeting without losing the room
A calibration meeting goes wrong when it turns into a debate club. I keep it structured, evidence-led, and short enough that managers stay disciplined. For a small group, 60 to 90 minutes is usually enough; larger groups need to be split so the discussion does not drift into fatigue and anecdote.
Before the meeting
- Define each rating level in observable terms so people are arguing about behaviour, not labels.
- Ask every manager to bring the same three inputs for each employee: goal progress, behaviour examples, and impact.
- Group comparable roles together first, then compare across roles only when the scope is genuinely similar.
- Flag outliers in advance so the group is ready to challenge them instead of discovering them late.
During the meeting
- Start with the most disputed cases, because they expose weak standards fastest.
- Ask the manager to cite evidence before anyone gives an opinion.
- Test whether the rating reflects role difficulty, actual performance, or just manager habit.
- Record the reason for any change, not just the final score.
After the meeting
- Send managers the final standards and the logic behind the hardest decisions.
- Track whether certain managers, teams, or rating levels keep drifting in the same direction.
- Use the next review cycle to check whether the calibration actually improved trust and consistency.
I prefer this structure because it keeps the conversation grounded. Once managers know the meeting is about evidence and standards, not persuasion, the quality of the discussion improves quickly. That also makes it much easier to decide what counts as strong evidence in the first place.
What strong evidence looks like when managers disagree
Calibration gets easier when managers bring facts that can be checked. Vague praise can be useful as colour, but it should not drive the final rating. I want evidence that shows what someone did, how often they did it, and what changed because of it.
| Weak evidence | Stronger evidence | Why the stronger version helps |
|---|---|---|
| “She has a great attitude.” | “She led three difficult client calls, resolved two escalation issues, and kept the account renewal on track.” | It shows behaviour and impact instead of personality language. |
| “He works hard.” | “He delivered two projects on time, reduced cycle time by 12%, and documented the process for the wider team.” | It links effort to measurable outcomes. |
| “They are not leadership material.” | “They have not yet shown consistent delegation, cross-team coordination, or coaching of others.” | It turns a vague judgement into observable development gaps. |
| “Everyone likes them.” | “Peers rely on them to unblock work, and three colleagues named them in feedback as the person who solved the issue fastest.” | Popularity becomes useful only when it is tied to actual contribution. |
That distinction matters because calibration is not just about scoring the past year. It is also about deciding what someone needs next. A rating with no evidence behind it is too weak to guide development, and too shaky to support pay or promotion decisions.
Common mistakes that distort ratings
Most weak calibration sessions fail for the same reasons. The danger is not only unfairness; it is also inconsistency, because people stop trusting a system that seems to change shape depending on the room.
- Comparing unlike roles without adjusting for scope, workload, or complexity.
- Using memory instead of notes, which gives recent events too much influence.
- Letting one loud manager dominate while quieter managers accept the first confident view in the room.
- Forcing a curve too early before the evidence has been reviewed properly.
- Confusing potential with performance and letting future promise inflate the current rating.
- Ignoring context such as illness, parental leave, team restructuring, or a role change during the cycle.
The fix is not complicated, but it does require discipline. I always come back to the same rule: if the meeting cannot explain why a rating changed, it probably should not have changed. That principle becomes even more important in UK teams, where fairness and documentation carry extra weight.
How calibration fits UK performance management
In the UK, I usually think in terms of appraisals rather than appraisals plus some hidden scoring layer. Acas describes reviews as a chance to discuss what someone is doing well, where they need support, and what development they need. Calibration is the part that keeps those conversations consistent across line managers, so the same behaviour does not get treated differently from one team to another.
CIPD also notes the shift toward more flexible, ongoing review systems. That matters because calibration should not depend on one end-of-year meeting and a blurred memory of what happened. It works better when managers have regular check-ins, written notes, and clear examples from across the year.
- Keep records of goal progress, feedback, and major changes in role scope.
- Separate performance from personality, especially when communication style is being mistaken for output.
- Check for adjustment issues if someone’s performance seems inconsistent with the support or tools they received.
- Review similar roles together before comparing across departments or seniority levels.
If I were building this for a UK organisation, I would also make sure managers understand that a calibration meeting is not a substitute for good management. It only works when the underlying review notes are honest, current, and specific. That is what makes the final step useful rather than decorative.
The habits I would keep for the next review cycle
If there is one thing I would carry into the next cycle, it is this: calibration should make the review conversation clearer, not more theatrical. The most useful system is the one that gets managers to slow down, use better evidence, and explain ratings in plain language.
- Use the same rating anchors every cycle so the scale does not drift.
- Keep a short rationale for every rating change, especially the outliers.
- Compare people against role expectations, not against whichever colleague happens to be easiest to remember.
- Refresh manager training before the review season starts, not after the mistakes have already happened.
- Revisit calibration whenever the organisation changes shape through restructures, promotions, or major workload shifts.
The strongest review processes are the ones that get more honest over time. When managers can point to evidence, when ratings mean the same thing across teams, and when the final decision reflects the full year rather than the loudest moment, calibration stops feeling like admin and starts doing real work.
