← All articles

Bias Aware Calibration Meeting Agenda: Facilitator Scripts + 60/75/90

Bias Aware Calibration Meeting Agenda: Facilitator Scripts + 60/75/90

Facilitator leading a calibration score discussion

A calibration meeting works when it runs on a timed, evidence-first agenda that requires mandatory prework, assigns a facilitator and scribe, builds in live bias checks, and closes with a written decision log. Judge success by whether every final rating is backed by documented evidence and every disputed case leaves the room with a named owner and a follow-up date.


TL;DR:

  • Calibration meetings should be scheduled with rigorous prework, including documented evidence and clear ratings from managers, to avoid turning the session into an information discovery.
  • The agenda must focus on flagged and outlier cases first, and assign roles such as facilitator, scribe, and presenting manager to prevent dominance and groupthink.
  • Evidence standards require at least two independent STAR-formatted points per contested case, backed by transcripts or recordings when possible, to ensure ratings are fact-based.
  • Final ratings must be documented with a clear decision log that includes before and after scores, evidence, rationale, owner, and follow-up due dates to ensure accountability.
  • Meeting length should match group size and cases flagged, with 60 to 90 minutes appropriate; the process benefits from visuals like shared rating sheets and live evidence review.

Xl
Build Stronger Calibration Evidence
XL Roleplay helps teams review realistic conversations, scored coaching reports, transcripts, and rubrics for more consistent performance decisions.

Table of Contents

Quick calibration agenda template with time blocks and owners

A calibration session needs a structure that gets to the hard cases quickly instead of burning the first half hour on easy ones. The template below works whether you run it in 60, 75, or 90 minutes: only the middle segment, person-by-person review, changes length.

The sequence stays fixed: open with purpose and ground rules, walk through the overall rating distribution, review flagged and outlier cases first, resolve remaining disagreements, then recap decisions and owners. Reviewing outliers first, rather than saving them for the end, keeps energy and attention on the cases that actually need debate, a practice reflected in tested 90-minute calibration formats.

Assign three owners before the meeting starts, not during it:

  • Facilitator: runs the clock, enforces evidence-first discussion, and calls out bias in the moment.
  • Scribe: captures the decision log in real time, including before and after ratings and rationale.
  • Presenting manager: brings the case for each of their direct reports, evidence in hand, ready to defend the rating rather than the person.

Pre-meeting checklist and manager deliverables

The single biggest failure mode in calibration is turning the meeting into a discovery session instead of a decision session. That happens when managers walk in without documented evidence and the group spends its time asking questions that prework should have already answered. Step-by-step facilitation guidance recommends requiring manager prework 48 hours before the session so the room can move straight to decisions.

Each presenting manager should arrive with:

  1. A proposed rating for every employee under review, not a range or a placeholder.
  2. Two to three specific evidence points per employee, written in STAR form (situation, task, action, result).
  3. Clear flags on any rating the manager is unsure about, so the group knows where to spend time.

On the facilitator or HR side, prework looks different but is just as concrete:

  1. Aggregate the proposed ratings into a single distribution view before the meeting, so the group can see clustering and outliers at a glance.
  2. Identify outlier cases in advance and slot them first on the agenda rather than discovering them live.
  3. Prepare a shared rating sheet, whether a spreadsheet or a shared document, that every attendee can see and edit in real time.

Send the agenda, the distribution summary, and the shared rating sheet to all attendees 48 hours before the meeting. That window gives managers time to review peers’ cases and arrive with informed questions instead of cold reactions.

Roles, ground rules, and logistics to prevent dominance and groupthink

Calibration meetings are structurally vulnerable to two failures: a handful of assertive voices dominating the discussion, and the group drifting toward agreement just to move faster. Harvard Business School Publishing notes that calibration can introduce new bias, including groupthink and dominance by assertive personalities, when the facilitator does not enforce structured, evidence-based discussion. Clear roles are the first defense.

  • Facilitator: owns the clock, redirects any comment that isn’t grounded in evidence, and names dominance patterns out loud when they appear.
  • Scribe: documents every rating change, the evidence cited, and who raised it, in a format the group can review later.
  • Presenting manager: speaks to evidence, not personality, and defends the score rather than the person.
  • Bias observer (optional): sits outside the rating discussion specifically to flag halo effects, recency bias, or similarity bias as they surface.

Three ground rules matter more than any others: keep everything said in the room confidential, require evidence before any rating challenge, and timebox every case so no single employee eats the whole session. Groups of five to eight managers tend to move fastest; anything larger benefits from breakout rooms or a split into multiple sessions, an approach detailed in guidance on running a first calibration session.

Pro Tip: Put the shared rating sheet on a screen everyone can see during the meeting, not just before it. Live visibility keeps the scribe accurate and keeps managers honest about what they actually said.

Facilitation scripts and a stepwise method for resolving disagreements

A facilitator’s exact words set the tone for the entire meeting. Opening with a script rather than improvising signals that the session runs on structure, not personality.

Try this to open: “We’re here to align ratings against evidence, not to negotiate favors. Every challenge to a rating needs a specific example attached to it.” When a case goes quiet, prompt with: “What’s the strongest piece of evidence behind this rating, in the employee’s own words or metrics?” When one or two voices have dominated three cases in a row, redirect directly: “Let’s hear from someone who hasn’t spoken on this one yet.”

When disagreement surfaces, work it in order:

  1. Ask the presenting manager to restate the evidence behind the proposed rating, not just the rating itself.
  2. Invite anyone who disagrees to offer counter-evidence, not just a gut reaction.
  3. If the group is split, give it a two-minute microdeadline to reach a decision or flag it for escalation.
  4. Escalate unresolved cases to a follow-up conversation between the facilitator and both managers, rather than forcing a vote in the room.

Calibration meetings can introduce new biases such as groupthink or dominance by assertive personalities if the facilitator fails to enforce structured, evidence-based discussion. Harvard Business School Publishing

Not every disagreement means the rating is wrong. Sometimes it means the rubric itself has a gap, an ambiguous competency definition or a missing example for what “exceeds expectations” actually looks like. Log that as a rubric gap for the next cycle instead of forcing a rating change just to close the discussion.

Bias mitigation checklist and evidence standards

Calibration exists to reduce inconsistency across managers, but the meeting format itself can introduce fresh bias if nobody names it while it’s happening. A Harvard Business Review analysis warns that calibration can amplify bias rather than reduce it when discussions are unstructured, and recommends strict facilitation paired with firm evidence requirements.

Calibration meetings without a named facilitator and evidence rule are prone to reintroducing bias rather than removing it, according to the Harvard Business Review’s analysis of calibration discussions. The fix is a live checklist the facilitator calls out by name during discussion, not a slide shown once at the start.

  • Halo or horns: one strong or weak trait coloring the entire rating.
  • Recency: a rating built on the last month instead of the full review period.
  • Similarity: favoring employees whose style or background resembles the manager’s own.
  • Proximity: rating in-office or frequently-seen employees higher than remote peers with equal output.
  • Central tendency: clustering everyone in the middle to avoid defending an extreme score.

Evidence standards need equal weight. STAR examples, backed by more than one source when possible, keep ratings tied to documented outcomes rather than adjectives, a standard echoed in calibration rubric guidance. Set a floor of at least two independent evidence points per contested rating, and use blind scoring or an outlier pre-flag before the meeting to keep the discussion from starting with a debate over who talks first.

Make and record decisions and post-meeting follow-up actions

Every rating that changes in the room needs a paper trail, both for fairness and for the next round of appeals or audits. The decision log should capture, per employee: the before and after rating, the specific evidence cited for the change, the rationale in a sentence or two, the owner responsible for follow-up, and a due date.

Field Example entry
Employee Case reviewed in session 4
Before rating Exceeds expectations
After rating Meets expectations
Evidence cited Two quarters of missed deadlines, documented in project logs
Owner and due date Manager, two weeks

After the meeting, three things happen in order:

  • Update the official performance record with the finalized rating, not the manager’s original proposal.
  • Draft feedback notes for each manager to deliver to their employee, separate from any coaching plan.
  • Assign development or coaching follow-ups where the calibration surfaced a real skill gap rather than a rating dispute.

Track agreement rate between proposed and final ratings, the count of open rubric gaps carried into the next cycle, and how much the overall distribution shifted. Those three metrics tell you whether calibration is tightening consistency or just adding a meeting.

Sample agendas and timing presets: when to use 60 vs 75 vs 90 minutes

The right session length depends on group size and how many flagged cases you’re carrying in, not on a fixed preference for shorter meetings.

  • 60 minutes: works for small groups, three to five managers, with a limited set of clear-cut cases and strict timeboxing on every segment.
  • 75 minutes: fits a slightly larger group or a mix of clear-cut and flagged cases that need a few extra minutes each.
  • 90 minutes: suits larger groups or sessions carrying more flagged and outlier cases, giving the resolution segment real room to work.

Smaller groups move faster and produce more candid discussion, while larger groups do better splitting into multiple sessions or using breakout rooms rather than stretching one meeting past 90 minutes, a pattern confirmed in guidance on first-time calibration sessions. When a session runs long anyway, triage: finish every flagged and outlier case in the room, and move uncontested middle-of-distribution cases to an asynchronous sign-off instead of rushing them in the final five minutes.

How transcript-linked scoring speeds up evidence collection

Gathering two or three defensible evidence points per employee is the slowest part of calibration prework, especially for sales and service teams where performance shows up in live conversations rather than static output. Managers who coach against recorded or transcripted practice sessions have a head start: a scored coaching report tied to a specific call or roleplay gives a rating a citable source instead of a manager’s paraphrase.

Transcript evidence linked to rubric scoring

Those transcripts also double as rubric anchors. A well-scored example of strong objection handling or a clear discovery call becomes a shared reference point for what “exceeds expectations” looks like on that specific competency, useful the next time two managers disagree on where a borderline case lands. See our sales rubric calibration guide for more on building those anchors.

Pro Tip: Technology can speed up evidence capture, but it doesn’t replace a facilitator enforcing ground rules. Bring the transcript into the room; don’t let it substitute for the discussion.

What experienced facilitators wish they’d known sooner

The most common mistake is letting the meeting open without a stated ground rule, which lets the first dominant voice set the tone for everyone after. The second is skipping outlier cases at the start and running out of time before reaching them.

Before you start, read this aloud: state the purpose, name the one non-negotiable ground rule (evidence before opinion), and confirm the shared rating sheet is visible to everyone. That thirty-second habit prevents most of what goes wrong later.

— Adam

A faster route to calibration-ready evidence

Sales and service managers often lose the most prework time chasing down what actually happened on a call instead of writing the rating itself. A structured pilot with realistic AI roleplay and scored coaching reports can shorten that step: reps practice against a persona built from your own methodology, and the transcript and rubric score become the evidence a manager brings straight into the calibration room.

Xl

A short pilot works best in a small group with a fixed rubric, tested over a couple of weeks, to see whether transcript-linked evidence actually cuts prep time before rolling it out further. Success looks like managers arriving at calibration with cited examples instead of memory, and a growing library of rubric anchors the whole team can reuse. Check the XL Roleplay pilot or compare the Individual and Business plans if scored roleplay evidence fits how your team already runs reviews.

Sources

The agenda structure, prework requirements, and bias safeguards in this guide draw on Calibration 101 from UC Davis HR and the step-by-step calibration facilitation guide from Windmill. For a deeper look at how calibration can introduce bias, read the Harvard Business School Publishing case and the Harvard Business Review analysis.

FAQ

What are the 5 points of calibration?

There is no single universal “5 points” framework; definitions vary by organization. A common version covers purpose and ground rules, distribution overview, person-by-person evidence review, disagreement resolution, and a documented recap of final decisions.

What is a calibration meeting?

A calibration meeting is a step near the end of a performance review cycle where managers who submitted draft ratings meet as a group to align those ratings against a shared standard, as described in UC Davis HR’s calibration overview. The goal is consistency across managers, not a first look at performance.

What should I say in my performance review meeting?

Speak to specific, documented outcomes rather than general impressions, using STAR-style examples (situation, task, action, result) tied to metrics or direct quotes. That same evidence standard is what calibration committees expect managers to defend in the room.

What are the steps involved in a calibration process?

The process starts with managers submitting prework, including proposed ratings and evidence points, 48 hours before the session, as recommended in Windmill’s facilitation guide. The meeting itself follows a timed agenda covering distribution review, case-by-case discussion, and resolution, and closes with a decision log and assigned follow-up owners.