A teacher in Greece opens a survey link and finds a student’s exercise on screen. Five parts, each already labelled correct or incorrect. Underneath sits a mark somebody else has given: five out of ten. The only thing left to do is enter a mark of their own.
Four of the five parts are ticked correct. By the exercise’s own arithmetic the work is worth eight, so the five sitting there understates it by three points, and everything needed to see that is on the screen.
The promise being tested
The standard reassurance about automated decisions is that a person stays in the loop. A machine drafts, a professional checks, the professional catches what the machine got wrong. Policy on algorithmic grading, algorithmic triage and algorithmic screening tends to rest on that sentence.
Grading is a useful place to test it: the judgement is evaluative, the arithmetic can be made objective, the consequences are real, and the people doing the checking are the ones who shape the outcome. Sofoklis Goulas, Rigissa Megalokonomou and Panagiotis Sotirakopoulos ran that test on 1,339 in-service teachers across Greece, and published it in PNAS Nexus in June 2026.
The vignette
Two things vary. The first is who supplied the wrong mark: a colleague, or an algorithmic grading system. The second is which way it is wrong. Exercises were matched to the teacher’s own subject, and within a subject everyone saw the same one. With four of five answers correct the fair mark is eight and the recommended five is harsh; with one of five correct the fair mark is two and the same five is generous. The recommendation itself never moves off five.
The outcome is what the authors call the grading fairness gap: the distance, in either direction, between the teacher’s own mark and the fair one. Each teacher saw one labelled exercise and gave one mark on it, so each contributes a single number. The trial was preregistered with the AEA registry.
The frame around it is narrow in ordinary ways, and the narrowness is worth carrying through everything that follows. One country. One labelled exercise per teacher. A survey a median teacher finished in 7.2 minutes, not a marking session. Day schools only, with evening, vocational and special-education settings excluded. And the study records where a mark landed, asking about belief only afterwards.
On the harsh version, under the human label, the average gap was 1.384 points. The right answer was one subtraction away from a list of five ticks and crosses, and teachers still landed about 1.4 points from it on average. They anchored hard on the mark they were handed, whoever they were told had set it. Whatever this study is about, it is not about a failure that begins with machines.
Only the harsh marks moved with the label
Under the algorithmic label the same harsh version ran 1.584 against that 1.384. The estimated difference is 0.300 points with controls, p = 0.003, and 0.264 without them, p = 0.022. The authors put that at a 22 per cent increase relative to the human baseline.
On the generous version, no detectable difference. The gap ran 1.930 under the human label and 1.813 under the algorithm — the estimate is −0.108 with a p-value of 0.685, which the study cannot separate from zero. Pooling both versions with an interaction term reproduces the same pattern, as does swapping the absolute gap for a signed one.
The paper’s prose and its Table 2 print slightly different means, which will look like an error to anyone checking. On the harsh version the text gives 1.386 human and 1.586 algorithm against the table’s 1.384 and 1.584; on the generous version it gives 1.818 algorithm and 1.942 human against the table’s 1.813 and 1.930. Which of each pair is larger never changes; the values do. The prose also gives the harsh estimate as 0.302 where the table gives 0.300, and a lenient estimate of −0.120 that matches neither of the table’s −0.116 and −0.108. The figures above are the table’s, and the 22 per cent holds on either set.
Under the human label the generous version carried the bigger gap, 1.930 against 1.384 on the harsh one. The printed paper reports no test of those two against each other, and they are not like for like. The authors report that teachers deviate asymmetrically: under the harsh benchmark they soften grades, under the lenient one they inflate them. By our reading, the same upward pull moves a mark toward the fair grade of eight on the harsh version and away from the fair grade of two on the generous one. The label effect sits on top of a task that was already lopsided, not in place of it.
Harshness read as competence
Why one direction and not the other? After entering their mark, teachers rated the grader on five dimensions: ability, comprehension, fairness, intent and responsibility. Across both versions they rated the algorithmic source well below the human one, the widest gaps falling on the generous version for perceived ability, intent, fairness and above all responsibility. But the deficit in perceived ability was smaller when the mark was harsh. The authors read severity itself as a signal that the grader knows what it is doing.
The mediation analysis follows that thread. On the harsh version the indirect path through perceived ability is 0.218, p < 0.001, and through responsibility 0.143, p = 0.016, while comprehension, fairness and intent are not distinguishable from zero. On the generous version all five paths run negative and significant: every channel that carried deference on the harsh version worked against it there, and the total effect is indistinguishable from zero.
The perception measures were taken after teachers saw who had set the mark, which the authors say makes this suggestive evidence on the channels rather than fully causal mechanisms. They were also taken after the grading decision, an order the authors chose so that the attitude questions could not prime or anchor the mark. It leaves open a reading the authors do not spell out: a teacher who has just let a mark stand has a reason to rate the grader well. The shares the paper attaches to those two paths, roughly 73 per cent for ability and about 47 per cent for responsibility, are each the path’s own coefficient divided by the same 0.300 total. That they sum past a hundred is what happens when overlapping paths are reported one at a time, and that reading is ours rather than the paper’s. The larger figure names the biggest identifiable channel, and nothing more than that.
It reached the confident ones
The deference did not spread evenly. On the harsh version it was significant for teachers under 51 (0.463), for those holding a master’s or a doctorate (0.442), for humanities specialists (0.533), and for those who rated their own technological literacy highly (0.486). For older, bachelor-only, STEM and low-tech-literacy teachers it was small and indistinguishable from zero. On the generous version no subgroup showed a significant difference at all.
The authors attach a caution that belongs beside those figures: the confidence intervals for some of these comparisons overlap, so the differences between the groups are not themselves established. It reached significance among the people most at home with the technology, which is not where a story about wary teachers would predict it.
The sample tilts toward one end of that list and away from the other. Eight per cent of these respondents hold a doctorate against roughly 2 per cent of Greek K-12 teachers, which is the direction the effect ran; their average age is 49 against 40 for the profession, which is not. The authors conclude that the findings may speak most directly to relatively senior teachers in the Greek system.
The end-of-survey questions add a further wrinkle. Asked in general terms, these teachers were not enthusiasts. On a scale from −5 to +5 their average belief that AI can grade fairly was 0.03; their willingness to let it grade, −1.03; their sense that doing so would be ethical, −1.33. Nearly half used generative tools at least weekly for lesson preparation, and more than half rarely or never encouraged colleagues to try them. The survey closed with a neutral open box asking whether there was anything they would like to add, and some used it to explain themselves. Of the comments that voiced reservations, the authors’ coding puts three in four on moral or ethical blind spots, the ill student and the child whose home life belongs in the mark, and one in four on technical limits; the paper does not say how many comments that was. The scepticism was real and it was stated. On our reading it did its work on the generous mark, where all five perception paths ran negative and the label left no net trace, and not on the harsh one, where the total effect ran the other way.
The easiness of the task was deliberate, and it decides what the number means. Making the correct mark deducible from a list of ticks isolates the label. It also means this measures reluctance to overrule a source, not oversight under the conditions that make oversight hard: ambiguity, fatigue, forty scripts and a deadline. The authors say so, and add that real systems arrive with explanations, confidence scores and accuracy records that a vignette does not have.
What this piece can vouch for is one experiment, read end to end: a vignette, a label, and a number. It cannot vouch for what happens in a real marking session, and where it argues past what the paper claims — about the mediation shares, and about why a teacher might rate a grader well after letting its mark stand — it says so in the sentence.
The teachers here had everything they needed on the screen and still landed well off the fair mark, and the label detectably added to that distance only where the machine was being hard on someone. If that is what oversight looks like when the error is one subtraction away and the mark belongs to nobody, what is it worth in a room where the error is buried and the mark decides where a fifteen-year-old goes next?




