Friday, September 4, 2026

Research suggests we file the strangers we deal with under the job independently of the person — in a Yale reading experiment, readers took about 213 milliseconds longer over the closing words when one customer replaced another than when one postal worker did

Must Read

“I need this truck moved, you’re making me late for work.”

Three hundred American adults read that line on Prolific, one chunk at a time, tapping the space bar to reveal the next few words. The people who wrote it were not interested in the complaint. They were interested in who was standing there to receive it.

In one version, the truck blocking the car belonged to a postal worker, and the person stopped ten minutes later was a different postal worker. In another, the truck belonged to another customer, and the person stopped was a different customer. Two labels changed; the grievance and its wording did not. Every figure attached to a named story in this piece, whether a reading time, a recall rate or a split of ratings, is a condition average pooled across that study’s three cover stories rather than a value for the story named. On that basis: over the last few words of the passage, the customer version took readers about 213 milliseconds longer than the postal-worker version. That is the raw gap between condition means, and it is the paper’s own figure; the model estimate behind the significance test is larger, at about 262 milliseconds.

That gap sits at the center of a paper published in Communications Psychology on 19 August by Aaron Baker, Yarrow Dunham and Julian Jara-Ettinger of Yale University. Their argument is that a job title does a lot of quiet work in ordinary life. It lets people anticipate what a stranger will do, expect what that stranger is likely to know, and treat one holder of the job as substitutable for the next, fast enough to register in how long a sentence takes to read.

When the replacement is another postal worker

The third study began an interaction with one person and, in two of its three conditions, finished it with someone else: a pen borrowed from one agent and returned to another at a doctor’s office, a computer borrowed from one hotel employee and its password requested from a second, and the blocked car at the post office.

When a postal worker was replaced by a different postal worker, the final region took 1352.24 milliseconds. When the same customer stayed for the whole interaction, it took 1327.88. When one customer was replaced by a different customer, it climbed to 1565.66. That last figure is over 50 percent above the average elsewhere in the study, where region-sized chunks ran 1024.69 milliseconds; the authors draw that comparison for the customer swap and not for the other two conditions.

Their summary is that readers were “equally unsurprised” by a same-role substitution and by no substitution at all, and surprised when the person taking over had no job to take over with. No significance test is reported for the postal-worker swap against the unchanged-customer condition, and the two differ on both counts at once: the role, and whether anyone was swapped. The paper offers the intuition behind it: a salesperson can step away mid-sale and another can take over, because the role is what the interaction runs on.

An exploratory analysis, flagged as such in the preregistration, found 47 percent of the 293 participants left after five who paused for over ten seconds and two with missing data were set aside had their slowest reading in the swapped-customer condition, against 29 percent for the swapped postal worker and 25 percent for no swap. With three conditions, chance sits at 33 percent, so one cell is above it and two are not.

Reading as a stopwatch

The method is narrow, and the narrowness is the point. In self-paced reading, a passage is broken into regions revealed one at a time, and the interval between presses is treated as processing difficulty. During the reading itself participants are given no task beyond reading, so a slowdown is a by-product rather than a judgement. Only afterwards were they asked recall questions and asked to rate how surprising or confusing each passage had been — the item ran both words together, which matters for an argument resting on the difference between difficulty and judgement.

Those ratings moved with the reading times in all three studies. In the first, the most common answer was “not surprising at all” for a role-consistent story, “pretty surprising” when the role-holder broke the script, and “very surprising” — the top of the scale — when the correct object was taken by someone with no role at all. In the third, the most common answer in the swapped-customer condition was still “not at all”, at 38.3 percent, but nearly matched by “a little” at 35 percent, and the condition differed significantly from both comparisons.

Most of the design work sits in the control conditions. A slowdown when a mechanic takes a cell phone off the counter could just mean an auto shop primes car keys rather than phones, so a third condition kept the object and changed the person: a customer taking your car keys ran to 1068.43 milliseconds, the mechanic taking the phone to 1020.98, the mechanic taking the keys to 849.34. The paper tests the first two against the last and not against each other. A slowdown when a stranger mentions your ulcer could just mean surprise that a non-staff character spoke at all, so a further condition had that character mention their own ulcer instead.

Most of the housekeeping is routine. Six reading times went missing to a software error in each of the second and third studies. The reading-time data break normality assumptions, and the authors report that the results survive a log transform. The model used for the surprise ratings in the first study broke a proportional-odds assumption; they refitted it with the assumption relaxed, found the result held, and still report the preregistered model.

One item is not routine, and it belongs beside the second study’s result. The rule the authors preregistered for outliers was to “exclude all reading times outside +/- 2 standard deviations of that participant’s mean reading time.” They came to regard that as a mistake, and their reasoning is sound: the rule strips each reader’s slowest responses, while the prediction under test is precisely that certain conditions make readers slow down. Their own supplementary figure shows it cutting the most data from exactly the conditions where a slowdown was predicted, “precisely because our experiment worked as predicted.” The main text therefore drops only pauses over ten seconds, 12 of the 2,700 readings across all three studies. Under the preregistered rule, however, the second study’s key contrast disappears: the fellow patient’s remark comes out “no different” from the doctor’s, at p = 0.83. The first study’s one control contrast also falls short of significance, at p = 0.17, though its main contrast holds at p = 0.017. Outlier handling in the third study was never preregistered at all. None of this is hidden — the paper flags the deviation in its methods and sets the whole of it out in a supplementary note. It does mean the second leg of the argument rests on an analysis choice made after the authors had seen their own data.

A doctor mentioning your ulcer

The second study moved from actions to knowledge. A doctor in a waiting room remarking on the ulcer in your stomach took 896.47 milliseconds, which the authors call comparable to the 809.09 that region-sized chunks took elsewhere in that study. A fellow patient making the identical remark pushed it to 1044.55. A patient remarking on the ulcer in their own stomach did not, on average, appear to slow readers down.

That third condition is what lets the authors argue against a duller explanation, though they are careful about the rest. It remains possible, they write, that readers “were not explicitly representing knowledge, and only forming a low-level expectation about the kinds of utterances people produce” — a doctor expected to say doctor-ish things, with no thought formed about what the doctor knows. Even then, they add, roles could still scaffold that thought later, when someone stops to have it.

The errors ran one way

These recall analyses were exploratory: the authors label them so in the first study and describe them as explorations in the other two, against a methods section that says all analyses were preregistered. The abstract nonetheless carries the misreporting result as a headline. Readers were asked two questions about each story, each a choice between two options, so chance sits at 50 percent. No condition fell near it. What varied was the direction of the misses. In the first study, 92 percent correctly named the role-holder who took the object, against 69 percent when the person doing it had no job in the scene, a customer in two stories and a friend in the third. In the second, 97 percent named the doctor as the speaker, against 74 percent when the same line came from another patient.

The paper’s account of that is a hypothesis rather than a result, and the hinge belongs in the quotation. Because readers could not scroll back, the authors suggest, they “might have therefore assumed they misread the passages and reconstructed the events in a role-consistent way. If so, the errors were not failures of memory, but of reconstruction.”

There is a literature running the other way, and the paper cites it: people are elsewhere found to remember unexpected events better, not worse. A reconciliation is proposed, not tested, and this design cannot settle it.

What the labels leave out

This piece was written from the paper rather than from a summary of it, which is the only expertise on offer here; the sharpest cautions in it are the authors’ own. Every story named its roles outright, in words. A physical setting offers uniforms and context instead, and the authors write that whether those cues produce the same rapid expectations is a question for future work. The participants were all in the United States, working with jobs that read as jobs in the United States, and the authors put cultural variation first on their list of limits. They also decline the tidiest available headline: “our work does not imply that role-based reasoning is faster than Theory of Mind.”

The article carries the publisher’s early-access banner. It is an accepted manuscript, and the version of record will replace it.

Two more questions are left standing, and the authors name both. One is how fast these expectations actually arrive and which cue sets them off; the paper raises the possibility that they are pre-loaded as we walk into a familiar space, before any person is seen. The other is how far the library of roles extends. A cashier can be replaced mid-transaction and a friend cannot, though the authors allow that a friend might be replaced over a longer horizon, and they ask for a more precise map of what sits between. The materials, data and analysis code are posted, which makes both cheaper to attempt than to argue about.

 

- Advertisement -spot_img
- Advertisement -spot_img
Latest News

Knowledge Keepers Preview

For millennia, Indigenous communities in North America have practiced a unique form of science. From navigating the featureless Arctic...
- Advertisement -spot_img

More Articles Like This

- Advertisement -spot_img