Skip to content

Moderation and marker consistency for teaching teams

Moderation6 min readUpdated DeepMarking team

Show contents

When a unit has four tutors, a student's mark should not depend on which of them picked up the script. In practice it often does, and students compare notes. This guide sets out a moderation process a small teaching team can actually run: calibrating before marking, double marking a sample, handling disagreements, and checking the distribution at the end.

What moderation is for

Moderation is the set of checks that make marks fair across markers and consistent with the standard the unit claims. It covers two different problems:

  • Inconsistency between markers. Tutor A is systematically two or three marks harder than Tutor B, or reads a criterion differently.
  • Inconsistency within a marker. The same tutor drifts over a long marking session, or marks a script differently depending on what came before it.

Most universities require some form of moderation, and the specific rules vary. Check your institution's assessment policy for what is required, such as minimum sample sizes or who signs off. What follows is a practical process that fits inside most policies.

Before marking: the calibration session

Calibration is the most effective moderation step because it prevents problems rather than correcting them. It takes 60 to 90 minutes and should happen before anyone marks for real.

Preparing

The unit coordinator selects four to six scripts from the actual submissions, chosen to span the range: an apparent fail, a borderline pass, a solid credit, a likely distinction, and one that is unusual in some way (off topic, unconventional structure, strong on one criterion and weak on another). Remove student names.

Send the scripts, rubric and any marker guide to each tutor a day or two ahead. Ask them to mark independently and bring their marks per criterion, with a short note on the evidence they used.

Running the session

  1. Put everyone's marks into a shared table before discussion, so nobody anchors on the coordinator's view.
  2. Start with the script where marks are closest. Agree quickly and move on.
  3. Spend most of the time on the scripts with the widest spread. For each, ask the highest and lowest markers to point to the evidence in the script and the descriptor words they relied on.
  4. Agree an interpretation and write it down, in one or two sentences, as a note to the marker guide.
  5. Agree the final mark for each calibration script. These become anchor scripts that all markers can refer back to.

A calibration table might look like this after round one:

ScriptTutor ATutor BTutor CCoordinatorSpread
1 (apparent fail)384235407
2 (borderline)5261505511
3 (solid)687066674
4 (strong)807684828
5 (uneven)5872606414

Scripts 2 and 5 need the most discussion. Script 5 is uneven, which is where analytic rubrics and markers most often part ways. Tutor B is consistently a little higher on middle-range work, which is worth noting gently.

What to record

Keep a one page calibration record: the agreed marks, the interpretations agreed, and any rubric wording that caused trouble. It is useful for your own moderation report, for next year's team, and if an appeal arises.

During marking: sample double marking

After calibration, tutors mark their allocations. Partway through, a second marker reviews a sample.

A common approach:

  • The coordinator or a second tutor blind marks a sample from each marker's pile. A typical sample is around 10 percent or a minimum of five scripts, whichever is larger, but follow your institution's rules if they specify a number.
  • Include every fail and a spread of scripts across grade bands, not just random picks. Fails and borderline passes are where errors matter most.
  • Do this early, after each tutor has marked a quarter of their allocation. Finding a problem at that point means remarking 20 scripts, not 80.

Blind second marking (without seeing the first mark) gives a truer check. Open review (seeing the first mark and comments) is faster and is often used for the non-fail part of the sample. Be clear which you are doing.

Handling disagreements

Decide in advance what counts as a disagreement worth acting on. A common threshold is a difference of more than 5 marks out of 100, or a difference that crosses a grade boundary.

SituationSuggested response
Difference within toleranceFirst mark stands
Difference over tolerance, one scriptThe two markers discuss and agree; if they cannot, a third marker decides
Consistent difference in one direction across the sampleReview that marker's whole allocation for the affected criteria, or apply an agreed adjustment after discussion with the coordinator
Difference on a fail or boundary scriptAlways discuss and agree; record the reasoning

Avoid simply averaging two marks as the default. Averaging hides the reason for the disagreement and can produce a mark neither marker thinks is right. It is better to talk through the evidence and agree a mark both can defend.

If a marker is consistently out of line, handle it as a professional conversation, not a reprimand. Usually the cause is a different reading of one or two descriptors, and a short conversation about the anchor scripts fixes it.

Where marks are adjusted, make sure feedback comments still match. A comment that says "strong analysis" on a script whose analysis mark was moderated down will confuse the student and invite an appeal.

After marking: checking the distribution

Once all marks are in, look at them by marker.

MarkerScriptsMeanMedianFailsHD
Tutor A6264.16545
Tutor B5870.371111
Tutor C6063.56454

Different tutorial groups can genuinely differ, so a gap is not proof of a problem. But a six mark difference in the mean, alongside what you saw in calibration, is a reason to pull five or six of Tutor B's scripts near grade boundaries and check them against the anchors.

Also look at scripts just below each boundary (49, 64, 74, 84 on a typical Australian scale). Clusters there often reflect reluctance to award the higher grade rather than a real difference in quality.

A simple moderation workflow

For a unit with a coordinator and three tutors, a full cycle looks like this:

  1. Week before submission: coordinator confirms rubric and marker guide, and shares allocation.
  2. Day 1 to 2 after submission: coordinator picks five calibration scripts; tutors mark them independently.
  3. Day 3: 75 minute calibration meeting. Update marker guide with agreed interpretations.
  4. Days 3 to 6: tutors mark first quarter of allocation.
  5. Day 7: coordinator second marks a sample from each tutor, including all fails so far. Raise any systematic issues immediately.
  6. Days 7 to 12: tutors finish marking.
  7. Day 13: coordinator reviews remaining fails, checks distributions by marker and scripts near boundaries.
  8. Day 14: marks finalised, moderation record written, marks and feedback released.

Scale the timings to your turnaround, but keep the order. Calibration first, early sampling second, distribution checks last.

Keeping consistency within a single marker

Moderation is often framed as a team problem, but sole markers drift too. The same habits help: choose anchor scripts before you start, reread one every 25 or 30 scripts, mark exam questions across the cohort rather than script by script, and look at your own grade boundaries before release. If you can, swap five scripts with a colleague in a related unit. Even a small outside check catches surprising drift.

Common questions

How many scripts should be double marked?
Requirements vary by institution, so check your assessment policy. A common practical approach is around 10 percent or at least five scripts per marker, always including every fail and a spread across grade bands.
Should disagreements be resolved by averaging the two marks?
Averaging is quick but hides why the markers differed. It is usually better for the markers to discuss the evidence and agree a mark, with a third marker deciding if they cannot.
What is the difference between calibration and moderation?
Calibration happens before marking and aligns markers on shared examples. Moderation is the wider process, including calibration, sample checking during marking and reviewing results afterwards.

DeepMarking marks against your rubric and lets you review every mark. Try it free.