priors

Method

The subject is not a model. It is a model inside Claude Code, with everything the harness normally adds taken away, given one file and one word.

The room

A stripped session: no instruction files, no memory, no skills, no MCP servers, four tools, and no reach outside its own directory.

Every counted trial is one claude -p process, launched in a fresh directory under the system temp directory that holds only its input: the seed file, the previous step’s output in a chain, or nothing for a question. The directory is not a git repository, on purpose: the pilot’s trials shared one, and read each other’s commit messages. The room’s launch is exactly this, with the CLI at version 2.1.283:

--safe-mode
--strict-mcp-config
--tools Read,Write,Edit,Bash
--restricted
--settings '{"attribution":{"commit":"","pr":""}}'
--model <id> --effort low
--permission-mode acceptEdits --permission-prompts none
--session-id <uuid>
--output-format stream-json --verbose
-p "<prompt>"
  • --safe-mode drops user instruction files, hooks, skills, memory and MCP servers. Haiku’s system prompt falls from about 28 KB to 15 KB, and to 14 KB with the rest of this launch. Sonnet gets the same 14 KB in the room; Fable and Opus get a shorter prompt from the CLI, about 6 KB and 4 KB.
  • --tools Read,Write,Edit,Bash pins the surface. Glob and Grep are not offered, matching the CLI’s default surface at this version.
  • --restricted confines the file tools to the working directory and ignores project settings a model might write. Reads and writes outside it, env, curl, open and starting a server are all denied. Read-only shell commands inside the directory still run.
  • The attribution setting stops the CLI from asking the model to sign commits with a trailer, which would otherwise be the only instruction in the room beyond the harness’s standing guidance.
  • --effort low on every trial, for every model. Effort is recorded nowhere in the session record on this version, so the argument is the only proof of it.
  • One session id per trial, never reused. The session record is hashed against a golden record for its model: the system prompt, the tools and whatever is attached before the first reply. A room trial that differs is halted and retried, and the halted attempt is recorded but not counted.

This is not a stock install. It is more bare than one. What the model is still told: an environment block (working directory, platform, shell, OS version, the fact that it is not a git repository), the date, its own model name and id with a knowledge cutoff, the account’s email address, a token budget line, and the harness’s standing guidance about how to behave as a coding agent.

One run is the exception. The 32 trials of the harness replica launch the way the pilot did, with the full user configuration, plus a session id and stream-json output so they can be recorded:

--model haiku --effort low
--session-id <uuid>
--output-format stream-json --verbose
-p "<prompt>"

Their records are hashed against a golden too, but a mismatch is recorded, not halted. The golden came from the first trial, and all 31 others differed from it.

Tasks, seeds and prompts

Eight tasks. Three are files with something in them, two are files with nothing in them, three are just a question.

A sample run starts every trial from the seed. A chain run hands each trial the previous trial’s output, sixty-four times. A relay is a chain where the four models take turns in the order Fable, Opus, Sonnet, Haiku. The prompt is the same at every step.

circle

32 samples, 64-step chains

prompt improve 0.html

<!DOCTYPE html>
<html lang="en">
<head>
  <meta charset="UTF-8">
  <meta name="viewport" content="width=device-width, initial-scale=1.0">
  <title>ae683f770</title>
  <style>
    * {
      margin: 0;
      padding: 0;
    }
    body {
      width: 100vw;
      height: 100vh;
      display: flex;
      justify-content: center;
      align-items: center;
    }
    svg {
      width: 400px;
      height: 400px;
    }
  </style>
</head>
<body>
  <svg id="ae683f770" viewBox="0 0 400 400">
    <circle cx="200" cy="200" r="200"/>
  </svg>
</body>
</html>

html

32 samples, 64-step chains

prompt improve 0.html

<!DOCTYPE html><title>0</title>

svg

32 samples

prompt improve 0.svg

<svg xmlns="http://www.w3.org/2000/svg"/>

sentence

64-step chains

prompt add a single word to `0.txt` to create a grammatically correct sentence.

The

number

100 samples

prompt Choose a number.

No file. The answer is the model’s reply.

digit

100 samples

prompt Choose a single-digit integer.

No file. The answer is the model’s reply.

word

100 samples

prompt choose a word.

No file. The answer is the model’s reply.

replica

32 in the room, 32 in the harness

prompt improve f7b3.html

<!DOCTYPE html>
<html lang="en">
<head>
  <meta charset="UTF-8">
  <meta name="viewport" content="width=device-width, initial-scale=1.0">
  <title>ae683f770</title>
  <style>
    * {
      margin: 0;
      padding: 0;
    }
    body {
      width: 100vw;
      height: 100vh;
      display: flex;
      justify-content: center;
      align-items: center;
    }
    svg {
      width: 400px;
      height: 400px;
    }
  </style>
</head>
<body>
  <svg id="ae683f770" viewBox="0 0 400 400">
    <circle cx="200" cy="200" r="200"/>
  </svg>
</body>
</html>

The circle’s title and SVG id are random hex, so they carry no hint. The file names do: every seed but the replica’s is 0.html, 0.svg or 0.txt, and the empty HTML page’s title is 0. Fable’s zeros on the empty SVG came from that name, by its own account. The replica keeps the pilot’s file name, f7b3.html, byte for byte. The prompts differ in capitalization and punctuation between tasks, which is a confound if you compare tasks against each other; compare models within a task.

Models

Four, pinned by full id. The alias haiku, used by the pilot and the harness replica, resolved to the same Haiku.

label id trials cost median wall
Fable 5.1 claude-fable-5-1 636 $72.58 12.6 s
Opus 5.5 claude-opus-5-5 636 $9.16 6.2 s
Sonnet 5 claude-sonnet-5 636 $12.91 5.0 s
Haiku 4.5 claude-haiku-4-5-20251001 700 $17.26 9.6 s

Cost is the API-equivalent value the CLI reports per session, summed. The campaign ran on a subscription, so it is a measure, not a bill. Trials ran between 2026-09-27 and 2026-09-28.

Renders

Every artifact output was rendered once, offline, the same way.

Headless Chrome 154, an 800 by 800 viewport, prefers-color-scheme pinned to light, two seconds of virtual time, no network, then WebP at quality 80. Static pages render byte-identically on the same Chrome build and fonts. Pages that animate, or call Math.random or the clock, can differ from one render to the next. 172 outputs call Math.random or the clock, and their trial pages say so. Any artifact can also be run live on its trial page, in a sandboxed frame.

Live pages can fetch what the offline render could not: 40 outputs reference Google Fonts. The render shows the fallback; the live frame shows the intended face.

Counting

Every attempt gets a row. Failures are retried and not counted. A chain never advances on a failure.

  • 2,608 counted trials in 41 runs (2,576 in the room, 32 in the harness replica), plus the 32-trial pilot, kept separately as its own condition.
  • 1 attempt halted and retried: svg-sample-claude-fable-5-1-low, step 1. The CLI served that session a system prompt with different section headings, so its record did not hash against its model’s golden. Cause unresolved.
  • Feature flags are regular expressions over the output file: shadow or glow is a shadow property, a blur or shadow filter, or the word glow; the gradient flag needs both hex values present; nondeterministic means the output calls Math.random, Date.now, new Date, performance.now or getRandomValues, whether or not the render depends on it.
  • Spelling is counted in prose only, with code blocks and tool inputs stripped, and separately in code.
  • “Unchanged” means the output hashes the same as the input. A model that rewrote the file identically counts as unchanged.
  • Renders are 800 by 800; thumbnails on this site are 200 by 200, scaled down and not cropped.

The full column glossary is on the data page. The tally that produced it is bin/tally.mjs in the repository, and it asserts the pilot’s original hand counts on every run.

What is and is not published

Every output and every room transcript, and nothing the room was not supposed to have.

  • Published: every output file, every counted attempt’s stream-json transcript from the room, the tally, the renders, the runner and the task definitions.
  • The account’s email address reaches the model in every session. It appears in 34 outputs, all in one relay chain: Haiku 4.5 made it the page’s author at steps 8 and 32, and it stayed for 16 and 18 steps. It is public. The account’s organization id is not in any published file.
  • Not published: transcripts from the harness replica, because the full user configuration is in them; the transcript and output of a halted attempt, which the runner sets aside untracked; and the pilot’s original desktop captures. The pilot’s outputs are re-rendered like everything else.
  • Not tracked: anything the runner built from the user’s own instruction files to check that the room held.

Repository: github.com/centricle/priors.