Skip to content

Marking programming assignments

Assessment design6 min readUpdated DeepMarking team

Show contents

Programming assignments look easy to mark objectively: either the code works or it does not. In practice, a submission that passes every test can be unreadable, and one that crashes on launch can contain the best design in the cohort. This guide covers how to set criteria that capture both, how to combine automated tests with reading the code, and how to handle suspected copying without jumping to conclusions.

Decide what the assignment is assessing

Programming tasks assess different things at different stages. A first year task might mainly assess whether students can write a loop that works. A third year task might care more about design, testing and maintainability than about any single feature.

Write down the two or three things the assignment most needs to show, then weight the criteria to match. A common failure is a rubric that is 70 percent "functionality" in a unit whose learning outcomes are about software design.

Criteria that work

Most programming rubrics are built from four or five criteria. Here is a worked example for a second year assignment: a command line inventory management tool in Python, about 400 to 800 lines, with a provided specification.

Criterion (weight)FailPassCreditDistinctionHigh Distinction
Functionality (35)Does not run, or fewer than half the core features workCore features work for normal inputAll core features work; handles common invalid input without crashingAll features work, including edge cases in the specificationHandles unusual input gracefully with clear error messages; behaviour matches specification exactly
Code quality and design (25)Hard to follow; large blocks of repeated codeCode is readable; some repetition; functions usedLogic divided into functions with clear responsibilities; meaningful namesSensible module structure; data and behaviour organised so changes are localDesign choices would make adding a new feature straightforward; consistent style throughout
Testing (20)No testsSome tests of main functions with normal inputTests cover core features including at least some invalid inputTests are organised, cover edge cases and run with one commandTests are clearly targeted at likely failure points; test names describe the behaviour checked
Documentation (10)No README or commentsREADME explains how to run the programREADME explains how to run and test; non-obvious code commentedDocstrings on public functions; README notes design decisions and limitationsDocumentation would let another student extend the code without asking the author
Version control (10)One or two commits at the endSeveral commits with vague messagesRegular commits with descriptive messagesCommits are small and focused; history shows incremental developmentHistory tells a clear story of the work, including fixes to issues found in testing

Note what is not in the rubric: marks for using a particular library, for code length, or for features beyond the specification. If you want to reward extensions, say so in the brief and cap the bonus.

Running code versus reading it

Programming markers need to do both, and the balance depends on cohort size and what you are assessing.

Running the code

Automated tests are the fastest way to check functionality, and they are consistent across markers. A good test harness:

  • runs each submission in the same clean environment, with the same Python version and dependencies;
  • uses tests the students have not seen, alongside any public tests they were given;
  • records which tests pass and fail, with output, per submission;
  • times out rather than hanging on an infinite loop.

Publish some tests in advance so students can check their basic setup, and keep the rest hidden so they cannot code to the test. Tell students you will do this.

Automated results are a starting point, not a mark. A submission might fail twenty tests because of one wrong filename or a missing import. Look at the pattern before scoring functionality. A common rule is that if a trivial fix (a renamed file, a changed import path) makes it run, the marker applies the fix, notes it, and applies a small deduction rather than scoring the work as non-functional.

Reading the code

Tests cannot judge design, readability or the quality of the tests themselves. For those you need to read the code. To keep this manageable:

  • Read with a purpose. Open the main module and one or two central functions. You do not need to read every line to judge structure and naming.
  • Use tools to find where to look. A linter report, a function length count or a duplication check shows which files deserve attention.
  • Mark by criterion across submissions. Mark functionality for everyone from the test output, then read code for design across the cohort. Your standard for "clear responsibilities" stays steadier when you compare submissions side by side.

Time budget

For the example assignment, a realistic budget is 20 to 30 minutes per submission: 2 minutes reviewing automated results, 10 to 15 minutes reading code and tests, 3 minutes checking README and commit history, and 5 minutes writing feedback. Run the test harness across the whole cohort before marking starts, not script by script.

Feedback on code

Programming feedback is most useful when it points at specific lines and suggests a concrete change.

WeakBetter
Poor structure.main() handles input, validation, file saving and printing in 120 lines. Try extracting validate_item() and save_inventory() so each function does one job.
Needs more tests.Your tests cover adding items but not removing an item that does not exist. That is the case that crashes your program (see test output).
Use better names.Variables like x, temp2 and data in update_stock make the logic hard to follow. Names such as item_id and new_quantity would make it self-explaining.
Good work.Your use of a dictionary keyed by item ID makes lookups simple and keeps find_item to three lines. Well chosen.

Handling similarity and suspected copying

Code similarity tools are widely used and useful. They are also easy to misread, and a wrong accusation is serious for the student and the marker. Treat similarity as a reason to look more closely, not as a finding.

Why code is often similar for innocent reasons

  • Starter code provided by the teaching team.
  • Small tasks with few sensible solutions. Two correct implementations of a simple function can be almost identical.
  • Code from lectures, tutorials or official documentation that students were encouraged to use.
  • Common library patterns and boilerplate.

Exclude starter code and known shared material from similarity checks where the tool allows it, and set thresholds based on what a typical cohort looks like, not on a fixed percentage.

Signals that deserve a closer look

  • Identical unusual choices: the same odd variable names, the same unnecessary step, the same misspelt comment.
  • The same bug in the same place, especially a bug that is not an obvious mistake.
  • Code whose style changes sharply partway through a file.
  • A submission far beyond what the student has shown in labs or earlier tasks, combined with an inability to explain it.
  • Commit history showing a large, complete block of code added in a single commit shortly before the deadline, with no earlier development.

None of these alone proves anything. Students work late, some start fresh repositories, and some improve quickly.

What to do

  1. Record what you observed, factually, with file and line references.
  2. Do not accuse the student or discuss your suspicion with other students.
  3. Follow your institution's academic integrity process, which usually means referring the matter to a designated officer rather than deciding it yourself.
  4. Where the process allows, a conversation in which the student walks through their code is often the fairest way to establish understanding.
  5. Mark the work on its merits in the meantime, or hold the mark, as your process specifies.

Generative AI tools add another layer. Policies on their use vary widely between institutions and units, so state clearly in the brief what is allowed, and design at least part of the assessment, such as a short code walkthrough or viva, so that understanding is demonstrated in a way that does not depend on detecting tool use.

Common questions

Should programming assignments be marked entirely by automated tests?
Automated tests are a consistent way to check functionality, but they cannot judge design, readability or test quality. Most rubrics combine test results with a marker reading the code.
What if a submission fails because of a trivial error like a wrong file name?
Many markers apply the trivial fix, note it in feedback and make a small deduction, rather than scoring the whole submission as non-functional. State this approach in the brief so it is applied consistently.
Does a high similarity score mean a student copied?
No. Similarity can come from starter code, simple tasks with few solutions or shared examples. Treat a high score as a reason to look closely and follow your institution's academic integrity process.

DeepMarking marks against your rubric and lets you review every mark. Try it free.