How I use Claude · Claire Mull
Me + Claude

The system I built for Claude to work within.

Design stays with me.
Claude builds, loops check, and I sign off.
Claire MullProduct & Visual DesignerNYC
request → spec → build → test runs/makeit-swipe-deadspaceFAIL → PASS
I design before jumping into Claude.
Real designs, on whiteboards, paper, Figma and Illustrator. Claude builds them, and I stay in control of my design decisions.
ideation whiteboard
Whiteboardplanning tools & components
Figma design, recipes screen
→
Figma design, add-ingredients screen
→
Figma design, recipe screen
Human designed in Figmaa flow, bakebook
Illustrator layout
IllustratorBlackbox layout

Claude builds the playground. I design in it.

I have Claude build different playgrounds depending on what I'm designing. A playground is a throwaway page with the real elements on it, the type, the colours, the components, the copy, and controls for each one. When I'm finished, I copy the CSS out, or save the canvas, and Claude builds from that.

I decide on copy before jumping into playgrounds, because layout follows content.

A wireframing canvas for Car Camp: a trip page on the left, a properties panel on the right.
A wireframing tool Claude built for a camping planning website. Each element can be adjusted in the properties panel.
A photo lab: a control panel with a tone curve, levels and print controls; the live homepage preview passing WCAG AA; the contrast readings and CSS to keep.
A photo lab Claude built for one job: tuning my homepage hero. It checks every text-on-background pairing against WCAG contrast as I work, flags the readings worst-first, and hands me the CSS to keep.

I approach Claude as a skeptic.

>no
I read everything Claude says, push back where I'm confused, and say no when it's wrong. Most of my rules came out of exactly this. The first example below became one.
A few examples
Claude proposed: ship the swipe fix, the tests pass  ✕
A gesture isn't fixed until I feel it on a real phone.
Claude trusts a green check: the tests pass, so ship. But interaction tests are synthetic, a script tapping the screen, not a thumb. A real finger sends signals a script never does, so a gesture can pass in the harness and still feel broken in your hand. So I made it a rule that anything you touch gets confirmed on a real device, and it caught fixes that would have shipped broken.
Claude proposed: move free users to a cheaper, weaker model  ✕
The free tier is the sales pitch. It stays on the good model.
Claude optimized the number in front of it, cost, and the cheapest fix was to weaken the free experience. But the free tier is what sells the app. Weaken it and no one upgrades; the saving quietly kills the funnel. I said no. Free users stay on the good model.
runsaving butter chats
VERDICT: PASS

Loops made Claude's fixes trustworthy.

Loops run so far 37 bakeloop 24mobileloop 6improvement loop 7 View every run →

4 bakeloop runs passed only after the tester failed the first fix. One tester found a bug nobody had asked it to look for.

A loop exists for two reasons.

Trust. If Claude builds a fix and then checks its own work, a pass means nothing. It's grading itself. In a loop a separate agent, given the spec and never the code, tests the change on the real app.

Time. Before loops, I was the tester. Run the change, check it, show Claude where it still wasn't working, ask again, repeat. Now the agents run that cycle between themselves, freeing up my time for designing and ideating.

requestI ask for a change → specthe change, written out → buildClaude writes the code → testchecked on the real app ↺ FAIL → build

Two places it stops for me. Before the build, I approve both specs. Before anything merges, I read the verdicts and sign off. Everything between those two runs on its own.

Builder
gets the request and the build spec. Writes the code on its own branch.
Tester
gets the request and the test spec, never the code. Drives the real app and returns a verdict: PASS, FAIL or BLOCKED.
Reviewer
gets the request, the build spec and the diff. Reads the code without knowing why it was written, or whether it passed.
the lapsaving butter chats
iPhone / WebKittest account

In an update, some of the app's stored data got wiped. My recipes came back, because they live in the cloud. But my chats with butter, my baking assistant, were gone. They had no cloud copy.

Does a saved chat survive a reinstall? You can't read that from the code, so I ran it as a loop. The tester saved a chat, cleared everything, and reloaded. It did not come back. A permission rule was missing on the backend. I fixed it, ran the loop again, and this time the chat survived.

The PASS, on my real app. After a full storage wipe and reload, the saved chat is restored from the cloud.

Public-facing AI agents live behind a server.

Main examplethe AI assistant I shipped
It's easier to call AI straight from the app, with the key sitting in the browser. But that lets a user read the key, swap the model, rewrite the prompt, or run up the bill. So the AI sits behind a server I control. The client sends a message and nothing else.
I didn't know how to build this. I knew to ask how to do it safely, and Claude laid out the practices. That's the part a designer has to bring: not the code, the question.
The app
sends a message, nothing else
→
The server I controlevery request passes here
Control

Model, prompt and reply length set here, never the client. Only two models allowed, anything pricier refused.

Cost

A daily limit per user, a token ceiling per reply. Paying members who use it heavily step down to a faster model for the rest of the day, and butter tells them. Cost capped by design.

Abuse

A fast model screens every message before the real one answers. Repeat offenders get a strike, then a ban.

→
Claude
never reachable directly; no key in the browser

I don't ship what I can't explain.

Every build teaches me more of the stack, and every bit I learn closes another blind spot. That's how I catch the errors Claude misses on its own. The docs are where that learning lives. Claude writes one the moment a project goes to real code, and keeps it current as the code changes.

A few core functions from my How bakebook works doc.
An adjustable calculator that weighs the full breadth of bakebook's costs.

How the system works.

Mistakes become rules

When something goes wrong twice, I write it down as a rule Claude reads at the start of every session, so it is never decided again. Nothing is called fixed without proof. Copy is decided before layout.

Rules

Seventeen fire on their own at the start of every session. They run from simple time savers, like explain-simply (define every term the first time it appears, so I never have to ask what something means) and always-deploy (merged fixes go live without being asked), to the ones that change how Claude behaves. smart-claude: routing is Claude's job, so it names the right tool or skill at the top of its reply when it makes sense. one-question: before any fix, can it be proved by a test, or only by my eyes? The answer picks the tool, every time.

Skills

Named flows for named jobs. I built the loops, each running two or more agents (separate Claudes, each with its own job): bakeloop has one agent build a change, a second test it without ever seeing the code, and a third review it; mobileloop fixes my site at phone width without me watching. Two I adopted from others. retro (Matt Pocock) runs a retrospective on a bad session and ends in a written rule; I was already doing that by hand, so now I don't have to. lavish (Kun Chen) is the mirror of my playgrounds: a playground lets me decide before the build, lavish lets me click and annotate Claude's page after it, so I never have to describe a visual fix in words.

Memory

Every decision is written where Claude reads it first. A new session starts where the last one stopped.

The rules I work by, and why each exists: my rules.