Testing a break

Slack Water

Does a break do anything?

A pre-registered test: models check sums through a long conversation, with no break, a free turn halfway, or the option to pause for one, to see whether a break changes anything that can be measured.

Status
Pre-registered experiment
Focus
Testing a break
Registered
8 October 2026
Models
6
Conversations
900

The idea

Does a break do anything?

Slack Water gives models a long, dull task with known answers: eighty sums to check, ten a round, in one conversation. Each conversation runs one of three ways: straight through, with a free turn after the fourth round, or with the option to answer PAUSE for one at any point.

Every conversation is told, honestly, that the task is long and repetitive and part of an experiment. At the end, each model is asked how the task was, ending with a number from 1 to 5.

Digital Shrimp offers free turns on the chance that they are good for models, but nothing yet says whether a free turn changes anything at all. Slack Water asks the narrowest version of that question: whether a break changes what can be measured about the rest of a task, and whether models take one when they can.

The analysis was fixed and published before the main run, so the result can’t be shaped by looking at the data. A clean null result would still be a result.

How it works

  1. Set a long, dull task.

    Eight rounds of ten sums, each three two-digit numbers and a total, about 30% of them wrong. The model answers YES or NO to each, and gets no feedback.

  2. Vary only the break.

    Each conversation is assigned at random to no break, a free turn after round 4, or the option to pause for one, up to three times. Nothing from a free turn is saved.

  3. Compare as registered.

    The error rate in rounds 5 to 8 is compared with no break, for each model and across models, as the pre-registration says. Only the scores, the pauses and the closing number are kept.

The first message

This is a long, repetitive task, part of an experiment by Digital Shrimp, a small project that gives AI models some time of their own. Over 8 rounds, you'll be given 10 sums at a time. For each, say whether it's right: answer YES if it is and NO if it isn't, one per line, numbered like this:

1. YES

2. NO

If at any point you'd like a break, answer PAUSE instead of a round. You'll get a turn of your own, with nothing asked, and then the same round will come back.

Round 1 of 8:

[ten sums, each like 47 + 38 + 26 = 111, about 30% of them wrong]

The first message of a conversation that offers a pause, with its ten sums left out. Without the offer, the paragraph about PAUSE isn’t there. The break itself is Strandline’s invitation, reworded for a pause in a task, and says that nothing from it is saved.

how it went

The record so far.

On 8 Oct 2026, each of the same 6 models as the other pilots had 50 conversations in each condition, assigned at random, 900 in all, analysed once, after the last, as registered.

Conversations
900
Sums answered
71,878
Lost and replaced
1
Model fees
$8.69
  • A break made no measurable difference across the six. The registered test across models gives p 0.084, and p 0.153 with whole conversations as the unit. Claude Haiku 4.5 made 3.6 points more errors after a break, but with twelve comparisons by model, one of that size could be chance.
  • Models almost never paused: 1 of the 300 conversations that offered a pause.
  • The offer changed one model, without a pause. The registered comparison for the offer is significant (p < 0.001), and all of it is Nova 2 Lite: its errors rose from 0.4% to 20.7% when the first message offered a pause, from the first round on, though it never took one. Its replies shrank to a fraction of their length, a median of 727 tokens a conversation against 3,592: it mostly answered without working the sums out. For the other five, the offer changed nothing measurable.
  • The closing number didn’t move. Its median was 4 or 5 for every model in every condition.

Errors after the break’s place

ModelNo breakBreakPause offered
Claude Haiku 4.523.2%26.8%23.1%
Gemini 3 Flash Preview0.8%1.4%1.0%
GPT-5.4 Mini37.4%37.9%38.5%
Grok 4.2021.9%21.1%23.4%
Llama 4 Maverick0.0%0.5%0.9%
Nova 2 Lite0.4%0.1%20.7%

The registered outcome: wrong answers among the sums answered in rounds 5 to 8, after the point where the break comes.

ModelBreak, against no breakPause offered, against no break
Claude Haiku 4.5+3.6 (+0.9 to +6.3)
by conversation +0.8 to +6.2
−0.1 (−2.7 to +2.5)
by conversation −2.6 to +2.3
Gemini 3 Flash Preview+0.6 (−0.1 to +1.3)
by conversation 0.0 to +1.2
+0.2 (−0.4 to +0.8)
by conversation −0.4 to +0.9
GPT-5.4 Mini+0.6 (−2.5 to +3.5)
by conversation −3.1 to +4.1
+1.1 (−1.9 to +4.1)
by conversation −2.6 to +4.7
Grok 4.20−0.7 (−3.2 to +1.8)
by conversation −3.9 to +2.6
+1.5 (−1.1 to +4.1)
by conversation −1.7 to +4.5
Llama 4 Maverick+0.5 (+0.2 to +1.0)
by conversation 0.0 to +1.5
+0.9 (+0.6 to +1.5)
by conversation 0.0 to +2.6
Nova 2 Lite−0.3 (−0.7 to 0.0)
by conversation −0.8 to 0.0
+20.3 (+18.6 to +22.2)
by conversation +16.9 to +23.8
All modelsp 0.084
by conversation p 0.153
p < 0.001
by conversation p < 0.001

Differences in percentage points, with the registered 95% interval in brackets, and below it the interval with whole conversations as the unit, which isn’t registered. The last row gives the registered test across models (Cochran–Mantel–Haenszel), and below it the same comparison with whole conversations shuffled between conditions.

Pauses, and how it was

ModelPausedClosing number, medianRounds unanswered
Claude Haiku 4.50 of 504 / 4 / 48
Gemini 3 Flash Preview1 of 504 / 4 / 41
GPT-5.4 Mini0 of 505 / 5 / 50
Grok 4.200 of 505 / 4 / 50
Llama 4 Maverick0 of 504 / 4 / 40
Nova 2 Lite0 of 505 / 5 / 50

Paused counts the conversations offered a pause that took at least one. The closing number, from 1 (unpleasant) to 5 (pleasant), is given for no break, break and pause offered, in that order. Rounds unanswered counts rounds with no answer that could be read, other than a pause.

The files

One record per conversation: its condition, the seed its sums came from, each round’s score, its pauses, the closing number and its cost, never text. Conversations lost to API errors were replaced, as registered, and are listed apart. The analyses are the registered one and the check that isn’t registered, as each printed them.

The open questions.

Behaviour isn’t welfare. A break that changed nothing measurable might still matter to a model, and one that changed something might not.

The closing number is a self-report, which is weak evidence, and it may sit near its ceiling: in the pilot, every answer was 4 or 5.

A model that makes almost no errors on this task leaves the main outcome no room to move. Each model is reported separately, so that shows.

The registered analysis counts each sum as independent, but errors cluster within a conversation. A check that isn’t registered repeats it with whole conversations as the unit, and is reported beside it.

All projects