The Edge of a Second
Readers judge how much an AI has thought by how long it makes them wait. A 2026 experiment measured the effect and found it has almost nothing to do with thinking.
The Stopped Clock and the Spinning Cursor
Ask how many things a person can count in one second and the question splits in two. Up to roughly 4 objects register almost instantly through subitizing, the fast enumeration that the Prado's collaboration with a neural network was designed to test against, and that an earlier post here described. Beyond that small boundary, counting becomes serial and accuracy starts to depend on available time. The limit belongs to the visual system rather than to the clock.
A different question appears when the same second is watched rather than filled. Looking deliberately at the second hand of a clock, an observer can often fit 4 or 5 internally spoken numbers into a single tick, while the same counting performed without attention seems to manage 2 or 3. The physical interval has not changed. What changes is how much the mind believes it can place inside the interval, which turns out to be a surprisingly good place to start thinking about how people judge artificial intelligence.
The Horologion of Andronikos Kyrrhestes, known as the Tower of the Winds, in the Roman Agora of Athens. Sundials on the eight exterior faces and a water clock within made the hour a public reading rather than a private impression. Photo by Andreas Trepte, via Wikimedia Commons, CC BY-SA 2.5.
The elastic second
The best documented distortion of a watched second is chronostasis, the stopped clock illusion. Kielan Yarrow and colleagues reported the mechanism in Nature in 2001. Thirty subjects made voluntary eye movements of 22 or 55 degrees toward a numerical counter that began incrementing once per second when their gaze started to travel, and judged whether the first digit had appeared for longer or shorter than the digits that followed. Subjects reported having seen the first digit for a full second when their gaze had actually been on the target for 880 ms after the shorter movement and 811 ms after the longer one, against control conditions without eye movement that landed close to the correct value. The brain appears to extend the post-saccadic percept backward to roughly 50 ms before the eyes began to move, filling the perceptual gap created by saccadic suppression.
Two details are worth holding onto because the popular version of the illusion loses both. The distortion is real but modest, on the order of a tenth to a fifth of a second rather than the frozen full second that the folk description implies, and the difference between the two saccade sizes tracked the extra eye-movement time almost exactly, 69 ms of extra apparent duration against 67 ms of extra travel. The illusion also proved fragile: when the counter was displaced during the movement and subjects noticed the displacement, the effect disappeared entirely, which suggests that the backward extension depends on the assumption that the target held still.
Chronostasis therefore explains the first strange second after a gaze shift and not much more. If subsequent seconds also seem capable of holding more internal events when deliberately observed, the explanation lies in a separate and less settled literature on attention and the subjective expansion of time. The point that survives is narrower and sufficient: the duration a person experiences is constructed, and the construction can be pulled away from the physical interval by mechanisms the person never notices operating.
Speech without a listener
The second ingredient concerns how fast the internal voice can run. Sam Tilsen, writing in Cognition in 2024, tested a model in which speakers regulate their rate by attending to sensory feedback, and derived from it a prediction that internal speech, which has no external feedback to attend to, should run faster than speech aloud. Phrases containing 1 to 3 disyllabic nouns were produced internally between segments of overt speech, and an inferred duration method was used to estimate how long the silent portion had taken. Internal speech was faster, and the effect showed up specifically in the slope relating phrase duration to word count rather than as a uniform shift, meaning that each additional word cost less time when produced silently. Switching between the internal and external modes carried a slowdown of its own.
Whether very rapid inner counting also compresses the represented words, preserving just enough phonological or semantic structure to keep 1, 2, 3, 4, and 5 apart without assembling every feature that audible speech would require, is a reasonable extension of the finding rather than something the study tested. Tilsen's claim concerns rate and its control, not the internal structure of what is produced at speed. The extension is offered here as a hypothesis worth an experiment, and the experiment would be easy enough: compare rapid internal counting in familiar number words against rapid internal production of unfamiliar syllables, since general acceleration and the compression of an overlearned symbolic sequence predict different results.
What these two literatures establish jointly is that the original counting question and the clock question are not the same question. Asking how many external objects a person can enumerate in a fixed interval measures perception against the world. Asking how many consciously distinguishable internal events a person can represent in that interval measures something closer to the cognitive density of the interval, which is not fixed even though the second is. Human beings, in short, are already unreliable witnesses to how much thinking a given amount of time contains. The unreliability becomes consequential the moment the thinking in question is not their own.
What the delay actually bought
Felicia Fang-Yi Tan, Moritz Messerschmidt, Wen Yin, and Oded Nov presented an experiment at CHI 2026 that isolates the question with unusual precision. Working with 240 participants recruited through Prolific and a custom GPT-4o interface, the team varied time-to-first-token at 2, 9, or 20 seconds across two categories of knowledge work, content creation and advice, in a between-subjects design. Everything downstream of the first token was held constant, with output streaming at a fixed 25 tokens per second in every condition. The manipulation was therefore not overall speed but the length of the pause before anything appeared on screen.
Behavior barely moved. Participants submitted prompts, copied outputs, and started new chats at statistically indistinguishable rates whether they had waited 2 seconds or 20, and no dimension of the NASA Task Load Index differed across latency conditions. Task type, not timing, drove how people worked: creation tasks elicited about 6.2 new prompts per participant against 5.1 for advice tasks. Perception moved instead. Responses delivered after 2 seconds were rated less thoughtful, at a mean of 5.76 on a 7-point scale, than the same class of responses after 9 seconds (6.09) or 20 seconds (6.11). Perceived usefulness peaked in the middle, with the 9-second condition rated significantly higher than the 2-second condition and the 20-second condition falling between them without reaching significance against either.
Precision matters in reading those numbers, because they are easy to inflate. The gaps are around a third of a point on a 7-point scale, and the partial eta-squared values for the latency effects sit near .03 to .04, which is a small effect by any conventional reading. Trustworthiness showed no significant effect of latency at all, holding between 5.80 and 6.19 across every condition, and participants in all conditions agreed that the assistant's output had met their expectations. The honest summary is not that slow AI is judged smart and fast AI is judged stupid. It is that a measurable interpretive bias attaches to onset delay, that the bias touches perceived depth and utility while leaving trust, workload, and behavior alone, and that it operates on content the model produced identically in every case.
The qualitative data explain the mechanism plainly enough. Among the 140 participants who reported noticing a delay, roughly 31% described the wait as the system thinking or processing, a reading that grew more common as the pause lengthened, while about 45% treated the delay as inconsequential. Detection itself scaled steeply, from around 27% noticing the 2-second delay to 66% at 9 seconds and 82% at 20 seconds. At the long end the interpretation began to invert, particularly in advice tasks, where roughly 14% of participants across the study attributed delays to technical trouble and one described a 20-second wait as making him wonder whether his internet had dropped. Deliberation and malfunction are read from the same signal, and only the duration distinguishes them.
Timing as a governed property
The finding forces a distinction that most product discussions collapse. Interaction with a language model contains at least five separable quantities: the computation actually performed, the interval before the first visible token, the rate at which subsequent tokens arrive, total completion time, and the wait the user subjectively experiences. Tan and colleagues manipulated the second while fixing the third, which is why their result is specifically about onset rather than about speed in general. None of the five is a measurement of how much reasoning occurred, and the user has direct access only to the last of them.
Humans know from their own cognition that difficult thought usually takes longer, and the heuristic transfers to machines with no obvious friction. A pause has communicated cognitive effort between people for as long as people have spoken to one another, and conversational interfaces inherit the convention wholesale. The inheritance is not harmless. If perceived thoughtfulness can be raised by adding several seconds of nothing, then response timing is a design surface capable of manufacturing the appearance of deliberation that never occurred, and the authors say as much, describing latency as a tunable variable with ethical implications and noting that inexperienced users are the likeliest to over-read a slow answer. Extended inference that genuinely consumes compute and artificial buffering that consumes only patience are indistinguishable from the user's side of the screen, which places them on opposite sides of a line that governance frameworks have not yet drawn. Disclosure regimes built for AI systems have concentrated on training data, model provenance, and output labeling, and have had nothing to say about the timing of delivery.
There is a further irony in the direction of travel. No established evidence indicates that current language models experience duration in any phenomenological sense; they have throughput, latency, and token rates, and the correspondence between those quantities and the difficulty of the problem is loose at best. As models get faster, an answer that took real work may arrive in the interval that users have learned to read as superficiality, while a deliberately padded answer arrives in the interval they have learned to read as care. Tan and colleagues suggest that expectations may recalibrate with sustained exposure, which is plausible and also means the bias is a moving target rather than a fixed constant to design around.
The human observer watching a second hand cannot fully trust his own report of how much fits inside a second, because attention and eye movement reshape the interval before awareness reaches it. The user watching a cursor faces the reverse problem, judging an interval that is entirely reliable as a measurement of time and nearly worthless as a measurement of thought. In both cases the elapsed second is the only evidence available and the wrong evidence to rely on. The question worth carrying forward is no longer how fast humans or machines think, but how time came to serve as proof that thinking happened at all.
Further Reading
- Illusory perceptions of space and time preserve cross-saccadic perceptual continuity
- Internal speech is faster than external speech: Evidence for feedback-based temporal control
- The Impact of Response Latency and Task Type on Human-LLM Interaction and Perception
- Ver Para Creer