Skip to content

Using AI to help mark responsibly

AI in marking6 min readUpdated DeepMarking team

Show contents

Many lecturers are now asking whether AI tools can take some of the load off marking, and some are already using them informally. The useful question is not whether AI can produce a mark (it can) but which parts of the marking process it can support without weakening the judgement students are entitled to. This guide sets out where AI help is reasonable, where it is not, and the practical safeguards around review, bias, privacy and transparency.

Where AI can reasonably help

A first pass against the rubric

Given a clear rubric and a submission, a language model can suggest a level for each criterion with a short justification pointing to parts of the text. Treated as a draft for the marker to accept, change or reject, this can speed up the routine part of marking, particularly for criteria that are fairly concrete, such as whether a report includes the required sections or whether each claim has a citation.

It is weaker on criteria that depend on disciplinary judgement: whether an argument is genuinely original, whether a design choice is sensible in context, whether a clinical reflection shows real insight. Suggestions there need closer checking.

Consistency checks

AI can be useful as a second reader. After you have marked, it can flag scripts where your mark and its suggested mark differ by more than a set amount, or where your comments do not match the level you awarded. You then look again at those scripts. This does not mean the AI is right; it means the script is worth a second look, much like a moderation sample.

Drafting feedback

Writing individual comments is one of the slowest parts of marking. An AI draft based on the rubric levels and the marker's notes can give you a starting point to edit. The marker's job is to make sure the comment is accurate, specific to the script and in a tone they are willing to put their name to.

Administrative work

Summarising common issues across a cohort, turning your notes into a comment bank, or checking that every script has feedback on every criterion are low risk uses that save real time.

Where AI should not decide

TaskWhy a human needs to decide
Final marksThe marker and the institution are accountable for the mark. A model's output cannot be questioned, cross-examined or held responsible in an appeal.
Fail and borderline decisionsThese carry the highest stakes for students and are where subtle judgement matters most.
Academic integrity findingsDeciding that a student has cheated, including through AI use, needs evidence and due process. Current AI detection tools are not reliable enough to base a finding on.
Special consideration and adjustmentsThese depend on circumstances and policy, not on the text of the submission.
Work outside the model's competenceHighly technical, creative or culturally specific work where the model may not recognise quality or may misread it.

A useful test: if a student appealed this mark, could you explain and defend it in your own words, based on your own reading of the work? If the honest answer is "the AI said so", the mark is not ready.

Human review that is real, not nominal

"A human reviews every mark" can mean anything from careful checking to clicking accept 200 times. Some habits keep review meaningful:

  • Read the submission, not just the suggestion. At minimum, skim enough of each script to know whether the suggested level is plausible.
  • Look at the evidence the suggestion cites. If a justification quotes the script, check the quote is there and says what is claimed.
  • Track how often you change suggestions. If you are accepting nearly everything, either the tool is very well suited to the task or you are not really reviewing. Spot check a few scripts from scratch to find out which.
  • Mark some scripts blind. Marking a handful without seeing the suggestion first keeps your own standard independent.
  • Be alert to automation bias. People tend to anchor on a number they have been shown. If you notice yourself adjusting your view towards the suggestion, pause and reread the rubric.

Bias and fairness

Language models learn from large amounts of text and can reflect patterns in it. In marking, possible concerns include:

  • rewarding a particular writing style, fluent academic English or familiar cultural references, which may disadvantage students writing in an additional language or from different backgrounds;
  • treating confident phrasing as stronger argument;
  • inconsistent results on very similar inputs.

Human markers have biases too, so the aim is not a perfect tool but a process that catches problems. Practical steps include checking whether suggestions drift for particular groups of students where you can do so appropriately, paying close attention to criteria about writing and expression, and never letting a suggestion override clear evidence in the script.

Student work is personal information, and often contains health details, personal reflections or names of others. Before sending any of it to an AI service:

  • Check your institution's policy. Many universities now have guidance on which AI tools are approved for student data and which are not. Some require formal approval before any student work leaves institutional systems.
  • Know where the data goes. Find out whether the provider stores submissions, for how long, where, and whether they are used to train models. A general purpose consumer chatbot account is usually not an appropriate place for student work.
  • Remove identifiers where you can. Names and student numbers are rarely needed for rubric-based suggestions.
  • Consider consent. Depending on your institution and jurisdiction, you may need to tell students, or ask them, before their work is processed by an external tool.

Being transparent with students

Students are entitled to know how their work is assessed. If AI plays any part in marking, say so in the unit outline or assessment brief, in plain terms:

Adjust it to what is actually true for your tools and your institution. If you want to add that student work is not used to train AI models, check that the tool's own terms say so first. Being open about AI use also makes it easier to ask students to be open about theirs.

How DeepMarking approaches this

DeepMarking is built around the review model described above. For each rubric criterion it suggests a mark and a comment that explains it, and flags the criteria it could not judge with confidence. You can change any mark or comment and confirm each criterion as you check it.

Two things to know when you plan your own process. First, DeepMarking records which criteria you have confirmed, but it does not stop you exporting marks you have not checked: the review step is yours to do. Second, the marking request includes the student's name and ID as read from the file names or cover page, so removing names from documents does not remove them from the request. Our Privacy Policy and the help article How student work is sent to the AI set out what is sent and kept.

DeepMarking is meant to take time off the routine part of marking, not to replace the marker's judgement, and it will not suit every task or every institution's policy. Check your institution's rules on AI tools and student data before using it or any similar tool.

A checklist before you start

  1. Your institution's policy allows the tool for this purpose and this kind of data.
  2. The rubric is clear enough that a careful human would mark consistently with it.
  3. Students have been told how AI is used in marking.
  4. You have tested the tool on a few scripts you have already marked, and know where it is weak.
  5. You will review every mark and comment, and read enough of each script to do so honestly.
  6. Fails, borderline scripts and any integrity concerns get full human marking.
  7. You can explain every final mark in your own words.

If any of those is not yet true, the time saved is probably not worth the risk.

Common questions

Can AI mark student assignments on its own?
It can produce a mark, but it should not be the final decision. The marker and institution are accountable for results, so a human should review every mark and be able to explain it.
Is it allowed to put student work into an AI tool?
That depends on your institution's policy and the tool's data handling. Many universities restrict which tools may process student data, so check before uploading any submissions.
Should I tell students if AI helps with marking?
Yes. State in the unit outline or brief how AI is used, that a human reviews every mark, and how their work is handled. Transparency supports trust and fair appeals.
Can AI detectors reliably identify AI-written student work?
Current detection tools are not reliable enough to support an academic integrity finding on their own. Treat any result as a prompt for further inquiry through your institution's process.

DeepMarking marks against your rubric and lets you review every mark. Try it free.