A lecturer marks a stack of essays and finds that the median has improved. Fewer structural problems, cleaner referencing, tidier arguments. The distribution has narrowed at both ends. Nothing is obviously wrong with any individual script, and she has no confidence that she is measuring what she thinks she is measuring.
Most of the sector response to this has been about detection and policing. That is understandable and it is losing. Detection tools are unreliable enough that acting on them is risky, and the arms race rewards whoever is willing to spend more effort on concealment. The more useful question is the one underneath: if the essay was a proxy for something, what was the something, and is there a better way to see it?
What the essay was actually for
The unseen exam and the take-home essay were never valued because universities needed more essays. They were instruments. They gave a marker indirect evidence about whether a student could hold a position, weigh competing accounts, notice what a source was doing, and construct a case that survived contact with an alternative reading.
That evidence was always indirect. The essay produced a trace of thinking, and the marker inferred the thinking from the trace. It worked because producing a good trace without doing the thinking was hard and slow. That constraint has gone. The trace is now cheap, and the inference no longer holds.
This is worth stating plainly, because a lot of assessment reform starts from the wrong place. The problem is not that students have a new tool. The problem is that a long-standing inference from artefact to capability has quietly stopped being valid, and most assessment design still depends on it.
Why the usual responses do not hold
Three responses have dominated, and each has a specific failure.
Return to invigilated handwritten exams. This restores the inference, at a cost. Timed handwriting under pressure measures recall and composure alongside understanding, and it disadvantages exactly the students who were already disadvantaged. It also assesses a way of working that no graduate will ever use again. Some of this is defensible for some subjects. Applied across a degree it narrows what a qualification means.
Detect and penalise. Setting aside accuracy, this makes the relationship adversarial and teaches students that the goal is to avoid being caught rather than to learn something. It also creates a genuine equity problem, since false positives fall disproportionately on students writing in a second language.
Ban it in some modules, permit it in others, and leave the policy vague. This is where most institutions actually are. Students receive contradictory signals from different lecturers in the same term and reasonably conclude that nobody has decided anything.
Assess the process, not only the product
The workable direction is to move evidence collection closer to the thinking and further from the artefact. This does not require abandoning written work. It requires the written work to stop carrying the entire assessment load on its own.
In practice that means asking students to show the route as well as the destination. A record of what they tried, what they rejected, where they changed their mind. A short conversation about the submission. A defence of one choice they made. These are not new inventions. Vivas, lab notebooks, studio crits and design journals have done this for decades in the disciplines that always assessed process, and those disciplines are having a considerably easier time this year than the ones that did not.
Five changes that are practical at scale
Workload is the constraint that kills most assessment reform. These are chosen because they survive large cohorts.
Add a short oral checkpoint, sampled rather than universal. Five minutes per student on a random sample of ten to fifteen per cent, with the sample announced in advance but the selection not. Students prepare as though they will be asked. The deterrent works across the whole cohort at a fraction of the marking cost.
Assess the annotated process, not just the submission. Ask for a one page account of the decisions behind the work: what was considered and dropped, which source changed their thinking, what they would do differently with another week. Mark it. Unmarked reflective addenda get treated as paperwork and written accordingly.
Set tasks that require local, recent or personal material. Work grounded in a specific placement, a seminar discussion from week four, this year’s local data or a student’s own fieldwork cannot be produced by a model that was not there. This costs design time once, not marking time every year.
Make critique of AI output the assessed task. Give students a generated answer and ask them to find where it is weak, what it has assumed, and what a specialist would notice that it missed. This assesses judgement directly and teaches a capability graduates will use.
Say what is permitted, per assessment, in one line. Not a policy document. A single sentence on the brief stating what use of AI is expected, permitted or excluded for this task, and what students must declare. Ambiguity is doing more damage than any individual rule.
What this looked like in one department
A social sciences department redesigned a second year module rather than the whole programme, on the reasonable grounds that they could not fix everything at once. They kept the 2,500 word essay and changed two things around it.
Students submitted a short decision log with the essay, covering the sources they discarded and why, and one point at which their argument changed. And the module ran fifteen minute conversations with a sample of students, in which they were asked about one paragraph of their own work.
The marking load rose by roughly a fifth in the first year, mostly because the decision logs were unfamiliar to mark consistently. Two effects were worth the cost. Seminar preparation improved noticeably, because students knew they might be asked to talk about their thinking. And the staff found the conversations diagnostically useful in a way marking had never been, because they could hear where understanding stopped rather than inferring it from a script.
The department was honest about the limitation. The approach makes it harder to submit work you do not understand. It does not make it impossible, and they stopped pretending that any assessment ever did.
Common traps
Redesigning assessment without reducing anything. Adding process evidence on top of existing requirements produces overloaded students and exhausted staff. Something has to come out. Usually it is the second essay that was assessing the same thing as the first.
Treating this as an integrity issue rather than a curriculum one. If it sits with an academic conduct committee, it becomes a rules problem. It belongs with the people who design what a degree is meant to evidence.
Assuming students are confident with these tools. Many use them badly and privately, without any sense of where the output is unreliable. Assuming fluency because someone is nineteen is a mistake.
Leaving individual academics to work it out module by module. This produces the inconsistency students complain about and burns out the staff who take it seriously. It needs a programme level decision, even an imperfect one.
Reflection questions
Useful for a programme team to work through together.
- For each assessment on your programme, what capability is it meant to evidence, and does the artefact still evidence it?
- Where on your programme does a student have to defend a choice out loud to another human being?
- If a student submitted excellent work they did not understand, at what point in the year would anyone find out?
- Which of your assessments could be answered by a model that has never attended your seminars, and does that matter for this particular learning outcome?
- What would you remove to make room for process evidence, and who needs to agree to that?
Assessment has always been an inference from evidence to capability. What has changed is that one particular inference stopped working, quickly, across most of the sector at once. Institutions that treat this as a design problem will end up assessing more of what they claimed to care about. The ones that treat it as a policing problem will spend the next decade losing an argument with a tool that keeps getting better.