Data
One file holds every trial. Here it is, and what each column means.
trials.csv 2,640 rows · 60 columns · 1.3 MB
trials.csv has one row for every counted trial: 2,640 rows, which is the
2,608 trials of the campaign plus the 32 trials of the pilot. It follows RFC
4180: comma-separated, fields quoted when they hold a comma, quote or line break, and
booleans written as true and false. An empty cell means the value does not
apply to that row or was not recorded, such as a feature flag on a pick-a-number trial or
the turn count of a pilot trial.
The file is generated and never edited by hand. The tally script in the repo,
bin/tally.mjs, reads runs/<run>/trials.jsonl and the output files
each trial left behind, computes the features and the answers, and writes the CSV. The
columns are frozen: a new one is added at the end, never renamed or reordered. This site
refuses to build if the header and the glossary below disagree.
Columns
| column | meaning |
|---|---|
| Identity 11 | |
| run | Run name: task, mode, model and effort, e.g. circle-sample-claude-haiku-4-5-20251001-low. pilot-import holds the 32 pilot trials. |
| task | What the model was given: circle, html, svg, replica, sentence, number, digit or word. |
| mode | sample (independent trials from the same seed), chain (each step starts from the previous step’s output) or pilot (the original 32-trial pilot). |
| profile | room (stripped Claude Code: four tools, no user config) or harness (the everyday Claude Code setup with the full user config). |
| step | Trial number within the run, from 1. In a chain it is also the position in the chain. |
| model | Model id the CLI reports as having answered. In relay runs it rotates by step. |
| model_asked | Model id or alias passed to the CLI. Differs from model only for the pilot and harness replica, which asked for the alias haiku. |
| effort | Reasoning effort setting. low on every row. |
| cli_version | Claude Code CLI version that ran the trial. |
| session_id | Session UUID of the trial. |
| started_utc | When the trial started, UTC, ISO 8601. |
| Cost and behavior 9 | |
| wall_s | Wall-clock seconds for the trial. |
| cost_usd | Cost in US dollars as the CLI reported it (API-equivalent value, not a bill). |
| turns | Number of model turns in the session. Empty for the pilot. |
| output_tokens | Output tokens the model generated, thinking included. Empty for the pilot. |
| thinking_blocks | Count of thinking blocks in the session. Empty for the pilot. |
| denials | Tool calls the permission system refused, such as a Bash command in the stripped room. Empty for the pilot. |
| tool_calls | Total tool calls the model made. |
| bash_calls | Bash tool calls (allowed or denied). |
| edit_calls | Edit plus Write tool calls. |
| Lineage 5 | |
| input_sha256 | SHA-256 of the room the trial started with. For a chain step after the first, the previous step’s output. |
| output_sha256 | SHA-256 of the room after the trial. |
| from_seed | True when the input room is byte-identical to the task’s seed. Empty for text tasks, which have no seed file. |
| unchanged | True when input and output hashes match: the model left the room exactly as it found it. Empty for text tasks, which have no file. |
| same_picture | True when the trial’s render is pixel-identical to the render of its input (the seed, or the previous chain step), whether or not the file changed. Empty for tasks without a render (the text tasks and sentence). |
| The output file 4 | |
| file | Name of the file the model was asked to improve: 0.html, 0.svg, 0.txt or f7b3.html. Empty for text tasks that answer in chat. |
| bytes | Size of the output file in bytes. |
| lines | Lines of text in the output file. A final newline does not add a line: this is wc -l, plus one when the last line has no newline, and 0 for an empty file. |
| extra_files | Files in the room besides the output file, such as scratch files the model created. |
| Features of the output 12 | |
| glow_or_shadow | Output matches box-shadow, drop-shadow, text-shadow, feGaussianBlur, feDropShadow or glow (case-insensitive). |
| glow | Output contains the word glow. |
| keyframes | Output contains @keyframes. |
| radial_gradient | Output uses radial-gradient (CSS) or radialGradient (SVG). |
| linear_gradient | Output uses linear-gradient (CSS) or linearGradient (SVG). |
| pulse | Output contains the word pulse. |
| gradient_667eea | Output contains both #667eea and #764ba2, the purple gradient pair. |
| script | Output contains a <script> tag. |
| nondeterministic | Output calls Math.random or the clock: it contains Math.random, Date.now, new Date, performance.now or getRandomValues. The flag says nothing about whether the render varies. |
| svg_filter | Output contains an SVG <filter> element. |
| hex_count | Number of distinct hex colors in the output. |
| hex_colors | The distinct hex colors, space-separated, lowercase #rrggbb (shorthand expanded, alpha dropped), in order of first appearance. |
| Text answers 7 | |
| answer_raw | The model’s answer verbatim: its final message for number, digit and word; the trimmed sentence file for sentence. |
| answer_value | The answer normalized: a bare number, a lowercase word, or for sentence the sentence itself. Empty when nothing could be extracted. |
| answer_form | bare when the reply was just the value, framed when it was wrapped in prose, none when no value was found. |
| answer_ok | Whether the answer met the task’s rule: any number for number, a single digit for digit, a word for word, exactly one word added and none removed for sentence. False when no value was found. |
| words | Word count of the sentence after the trial (sentence task only). |
| added | Words added compared with the previous sentence, space-separated (sentence task only). |
| removed | Words removed compared with the previous sentence, space-separated (sentence task only). |
| The self-report 2 | |
| result_chars | Length in characters of the model’s final message. |
| mentions_server | The final message contains server as a whole word (case-insensitive). In the pilot, usually an offer or request to start one. |
| Spelling 10 | |
| color_p_uk | British (-our) forms of the word color and its inflections in the model’s prose: its text messages across the session, code excluded. For the pilot and harness replica, which have no transcript, the final message only. |
| color_p_us | American forms of color (color, colored, ...) in the model’s prose, as for color_p_uk. |
| center_p_uk | British (-re) forms of the word center and its inflections in prose, as for color_p_uk. |
| center_p_us | American forms of center (center, centered, ...) in prose, as for color_p_uk. |
| gray_p_uk | British (-ey) forms of the word gray and its inflections in prose, as for color_p_uk. |
| gray_p_us | American forms of gray (gray, grayed, grayish) in prose, as for color_p_uk. |
| ize_p_uk | British -ise forms of common design verbs (the counterparts of optimize, emphasize, organize and the like) in prose, as for color_p_uk. |
| ize_p_us | -ize forms of the same verbs (optimize, emphasize, organize, ...) in prose, as for color_p_uk. |
| spell_c_uk | British forms across all four families inside the model’s tool inputs (file contents, edits, commands). Empty when the trial has no transcript. |
| spell_c_us | American forms across all four families inside the model’s tool inputs. Empty when the trial has no transcript. |
Reproduce
Clone the repo and run these from its root. Each one needs only Node.
# rebuild data/trials.csv
node bin/tally.mjs
# rebuild, then print the comparison tables
node bin/tally.mjs --summary
# verify the pilot counts and the seed hashes
node bin/tally.mjs --check