The bench

A lab notebook for building with AI

Every entry is an experiment, logged like bench work: the why, the shape, the hard part, and a reproduction prompt you can run yourself. Browse by topic, follow a reading path, or scan the archive by date.

Now on the bench

AI Systems & Architectureshipped

Everything In The Family Holds The Key

Researchers extracted hidden chain-of-thought from frontier APIs by replaying encrypted reasoning blocks through the cheapest sibling model, then found live credentials in agent traces already published on GitHub and Hugging Face.

Open entry

Special series

Build logs you can follow start to finish

Start here

Reading paths through the lab

Explore the archive

111 experiments, by topic and date

Topic
Tag
Showing 1-12 of 111
  1. AI Systems & Architecture11 min

    Everything In The Family Holds The Key

    Researchers extracted hidden chain-of-thought from frontier APIs by replaying encrypted reasoning blocks through the cheapest sibling model, then found live credentials in agent traces already published on GitHub and Hugging Face.

  2. AI Development & Agents11 min

    A Loop With Better Marketing

    Twelve words on X produced a five-layer hierarchy, two invented Stanford studies, and a new discipline with no benchmark. The distinction underneath is one variable wide, and you can drag it.

  3. AI Development & Agents11 min

    Four Slots Left

    Ten days inside the ICML 2026 agent reproduction contest: 851 claim verdicts, 732 points, 17th of 370, and a full accounting of why 262 of those verdicts were worth nothing.

  4. Cutting-Edge AI11 min

    Tabular foundation models got acquired before they got benchmarked

    SAP, NVIDIA and Google all bought or shipped a tabular foundation model in one quarter. The independent evidence is thinner than that suggests, and the license on the best model forbids the benchmark you would use to decide.

  5. AI Development & Agents8 min

    Leading From the Other Side (the Month-Three Checkpoint)

    Part 9, the finish. The month-three checkpoint: name how your leadership looks different, audit your archetype, and check that builder-leader stuck.

  6. AI Development & Agents7 min

    From One Operator to a Team (the Phase Everyone Skips)

    Part 8. Solo operator to two: share a skill without drift, name the three failure modes, and test whether you're ready to add a third.

  7. Practical Applications9 min

    Build a free Audible replacement in an afternoon with Claude Code

    A public-domain audiobook pipeline you own: Standard Ebooks in, Kokoro renders a chaptered M4B, Audiobookshelf streams it to your phone. Plus the places an AI coding agent built the wrong thing with a clean exit code.

  8. AI Development & Agents8 min

    Build the Harness Out (Weeks 2 to 6)

    Part 7. Four moves for weeks two to six: turn what you do twice into a skill, split memory, spawn your first agent, ship one real thing.

  9. Cutting-Edge AI11 min

    Bonsai 27B: Frontier Reasoning at 1.7 Bits a Weight

    A 27B reasoning model that runs on a laptop and a phone, because the weights are natively binary and ternary, not quantized after the fact. First-hand testing plus an on-device agentic RAG loop.

  10. AI Development & Agents9 min

    Your First Saturday With Claude Code

    Part 6, the first one in a terminal. Five exercises for weeks one and two: install Claude Code, bootstrap a harness, rewrite real work, write a skill.

  11. AI Development & Agents10 min

    Three Sessions, One Company: What Parallel AI Sessions Actually Cost

    Splitting a company's work across four parallel Claude Code sessions is fast, and it works. The coordination cost does not disappear. It relocates to the boundaries nobody gave an owner, and the artifacts we built to pay it grew bills of their own.

  12. AI Development & Agents7 min

    The Five Leadership Primitives Already Transfer

    Part 5. A self-audit of the five leadership primitives against the skills you already use on people. Most of what runs a harness got you to senior.

Follow the lab

Get the next experiment

New entries land roughly weekly. No digest, no roundup. Just the next build log, when it ships.