The bench

A lab notebook for building with AI

Every entry is an experiment, logged like bench work: the why, the shape, the hard part, and a reproduction prompt you can run yourself. Browse by topic, follow a reading path, or scan the archive by date.

Now on the bench

AI Development & Agentsshipped

From One Operator to a Team (the Phase Everyone Skips)

Part 8. Solo operator to two: share a skill without drift, name the three failure modes, and test whether you're ready to add a third.

Open entry

Special series

Build logs you can follow start to finish

Start here

Reading paths through the lab

Explore the archive

106 experiments, by topic and date

Topic
Tag
Showing 1-12 of 106
  1. AI Development & Agents7 min

    From One Operator to a Team (the Phase Everyone Skips)

    Part 8. Solo operator to two: share a skill without drift, name the three failure modes, and test whether you're ready to add a third.

  2. Practical Applications9 min

    Build a free Audible replacement in an afternoon with Claude Code

    A public-domain audiobook pipeline you own: Standard Ebooks in, Kokoro renders a chaptered M4B, Audiobookshelf streams it to your phone. Plus the places an AI coding agent built the wrong thing with a clean exit code.

  3. AI Development & Agents8 min

    Build the Harness Out (Weeks 2 to 6)

    Part 7. Four moves for weeks two to six: turn what you do twice into a skill, split memory, spawn your first agent, ship one real thing.

  4. Cutting-Edge AI11 min

    Bonsai 27B: Frontier Reasoning at 1.7 Bits a Weight

    A 27B reasoning model that runs on a laptop and a phone, because the weights are natively binary and ternary, not quantized after the fact. First-hand testing plus an on-device agentic RAG loop.

  5. AI Development & Agents9 min

    Your First Saturday With Claude Code

    Part 6, the first one in a terminal. Five exercises for weeks one and two: install Claude Code, bootstrap a harness, rewrite real work, write a skill.

  6. AI Development & Agents10 min

    Three Sessions, One Company: What Parallel AI Sessions Actually Cost

    Splitting a company's work across four parallel Claude Code sessions is fast, and it works. The coordination cost does not disappear. It relocates to the boundaries nobody gave an owner, and the artifacts we built to pay it grew bills of their own.

  7. AI Development & Agents7 min

    The Five Leadership Primitives Already Transfer

    Part 5. A self-audit of the five leadership primitives against the skills you already use on people. Most of what runs a harness got you to senior.

  8. AI Development & Agents7 min

    Inventory Your Harness: The Six Components

    Part 4. The six components of a harness, a one-page inventory of which you have versus actually use, and how to pick the first one to build.

  9. AI Development & Agents8 min

    Why Your Org Can't Cross (Run Your Own Pilot Numbers)

    Part 3. Compute your own pilot-to-production ratio against the 88% failure pattern, and find which crossing myth is quietly eating your budget.

  10. Practical Applications9 min

    Maximizing Toward the Local Minimum: How Fast Optimization Drifts, and the Human Habit That Catches It

    An autonomous research engine ran hundreds of experiments in two days and the dashboard stayed green the whole time. Almost all of them were the same experiment. Here is how a system optimizes itself into a groove, why every helper we built pushed it there, and the human-in-the-loop habit that caught it fast.

  11. AI Development & Agents7 min

    Hear the Three Tiers (and Name the Failure Mode)

    Part 2. Three listening drills: hear whether a conversation is about the model, the app, or the harness, and name your org's real failure mode.

  12. AI Development & Agents7 min

    Which Side of the Build Gap Are You On? (Run the Scorecard)

    Part 1 of the Builder-Leader field guide. An eight-question scorecard, run on yourself in fifteen minutes, for which side of the build gap you're really on.

Follow the lab

Get the next experiment

New entries land roughly weekly. No digest, no roundup. Just the next build log, when it ships.