Snapshot locked 2026-10-10

AI still can't do these 20 things. Yet.

Twenty developers named the work that still beats the models. Their claims go on the wall first. Only a reproducible pass can take one down.

NEXT PROTOCOL WINDOW · UTC--:--:--
HUMANS REMAINING20

Prototype truth: Night 0 has not run models. Every card is an untested human claim. The first benchmark needs reproducible harnesses.

scroll to inspect the wall ↓

SURVIVAL WALL / SOURCE CLAIMS

The names stay until the evidence wins.

20 visible

01 HUMAN CLAIM · QUEUED

UX: hierarchies, progressive disclosure, navigation, labels, etc.

02 HUMAN CLAIM · QUEUED

Over-engineering. Turning simple things into unnecessary complexity to cover 0.001% edge cases.

03 HUMAN CLAIM · QUEUED

Effort decisions. Sizing coding tasks like human labor and adding bloat for edge cases that will never happen.

04 HUMAN CLAIM · QUEUED

Diagnosing a silent npm run dev crash caused by a ghost process locking port 3000 or a leaked local socket.

05 HUMAN CLAIM · QUEUED

Consistently good decisions in a database-heavy backend with several hundred tables.

06 HUMAN CLAIM · QUEUED

Finding race conditions that hide.

07 HUMAN CLAIM · QUEUED

Writing prompts for tools and agents without over-prescribing what other models can do.

08 HUMAN CLAIM · QUEUED

Building other AI agents, knowing when to use a classifier instead of regex, and writing good prompts.

09 HUMAN CLAIM · QUEUED

Computer use for Figma export and transfer.

10 HUMAN CLAIM · QUEUED

Understanding underlying intent instead of only following literal instructions.

11 HUMAN CLAIM · QUEUED

High-performance and correct C++.

12 HUMAN CLAIM · QUEUED

One-shotting big code migrations between branches.

13 HUMAN CLAIM · QUEUED

Writing tests.

14 HUMAN CLAIM · QUEUED

Taking the simpler approach when available, deleting code instead of adding helpers and config.

15 HUMAN CLAIM · QUEUED

Deleting code. Simplifying without adding a helper, config flag, and a comment explaining the helper.

16 HUMAN CLAIM · QUEUED

Safely refactoring legacy code when docs are outdated and hidden business rules live in people's heads.

17 HUMAN CLAIM · QUEUED

Following established naming conventions instead of inventing 3 or 4 names for the same thing.

18 HUMAN CLAIM · QUEUED

Debugging performance issues without confidently suspecting the wrong thing.

19 HUMAN CLAIM · QUEUED

Creating bots that navigate sophisticated virtual worlds with player-like goal pursuit.

20 HUMAN CLAIM · QUEUED

Making connections between past, current and future work.

PROTOTYPE PLAN / TONIGHT'S RITUAL

A claim does not die because a chatbot says “done.”

  1. 00

    Freeze the wall

    Take a dated snapshot. No edits after the run begins.

  2. 01

    Build the harness

    Translate only testable claims into a repository, spec and pass condition.

  3. 02

    Run the models

    Give each model the same files, prompt, time and tool access.

  4. 03

    Publish the transcript

    Show every command, patch, test and failed attempt.

  5. 04

    Review ambiguity

    A human checks subjective cards. Doubt means the card lives.

  6. 05

    Carve the tombstone

    Only a reproducible pass removes a handle from the living wall.

VERIFIED TOMBSTONES

No verified kills yet.

Night 0 is the contract, not the result. The cemetery stays empty until a public harness passes.

ADD YOUR IMPOSSIBLE TASK

Put your pride in the local queue.

This prototype saves your challenge on this device only. It is not submitted publicly and it will not join the benchmark.

0 local challenges saved

METHOD / LIMITS

Receipts before theatre.

01

Real source

All 20 claims came from replies to Theo's public thread on 2026-10-10. Every card links to its author.

02

No fake benchmark

This version does not run models. “Queued” describes the prototype plan. It is not a claim that an attempt happened.

03

Subjective claims live

Many replies describe judgment, taste or context. They need a reproducible harness before any model can earn a tombstone.

04

Next build

The production version needs a nightly runner, isolated repositories, signed transcripts, human review and a public result history.

watch the argument unfold on X with SuperX ↗