For the last sixteen days I didn't write any code on my own website. Claude wrote all of it, and I ran the project the way I'd run an engineering team.
Before that, this site was a small portfolio on GitHub Pages: 23 commits over almost two years, a resume, and not much else. Now it's an editorial public site with posts and projects, a CMS so I can publish without a deploy, a private notebook I use every day, an iPhone app on TestFlight that keeps working offline, user and access management with passkeys, and a Go API. It's all defined in code on AWS, and it costs about $10 a month to run.
Along the way there were more than 440 merged pull requests, over 450 Linear tickets and more than 2,000 tests.
All of it is public. The code, every pull request and the rules every agent works by (docs/agent-rules) are on GitHub, so you can check anything I describe here.
I did it for two reasons. First, I take notes in a very opinionated way, and I've always wanted to build my own notebook app. Second, I wanted to see how fast I could build a complex site from scratch with the agentic workflows I use at work, but with the luxury of a personal project: move as fast as possible without much worry about breaking things. There's no staging or test environment. Until v1, everything went straight to production.
What I learned, in short:
- It's mostly management. The AI writes the code, but clear tickets, a real definition of done and separate reviewers decided whether it was any good.
- Trust has to be earned per scope. I handed out merge rights one project at a time, and kept a short list of things I never delegated.
- Agents are consistent with their own assumptions. The bugs that got through were the ones only reality could catch, so test against reality.
| Before | After |
|---|---|
![]() |
![]() |
Start with a plan, not a prompt
My first message wasn't "build me a CMS." It was this:
I'd like to plan out improvements to my website in Linear. Please first investigate what I currently have and give me a summary.
Then I listed what I wanted: publish posts without a deploy, a notebook for notes and tasks, eventually an iOS app, and AWS where "I'm managing my AWS infrastructure via code. I don't want to have to use the AWS console to make changes."
About an hour later I had five projects and around fifty tickets with milestones and dependencies, and no code. One rule held from day one: one ticket, one branch, one pull request. Branch protection was itself a ticket. Linear became the shared source of truth for every agent and every tool. Each session started by reading it and finished by updating it.
Staff the team
The work settled into roles. Coordinators, one Claude Code session on my laptop and a Claude project in the cloud, broke down the work, briefed engineers and sequenced merges. Engineers were short-lived agents, each working one ticket in its own git worktree or cloud thread. Reviewers were separate agents that audited finished work and turned findings into tickets. Claude Design did the design. I was the product owner.
Keeping the builder and the reviewer separate mattered more than anything else. One of the reviewers said it on day two: "Keep the author and the reviewer separate. The biggest wins came from one model building and another checking."
Write the job description
Every engineer agent got the same written brief before it touched a ticket:
No AWS writes, no deploys, no Linear writes, no merging, no pushing to
main, no force-push.
And the line that sums up its job: "Done for you = PR opened, not merged."
An engineer's job ended at a green pull request. The coordinator owned the rest: merge once every check passes, watch CI deploy it (merging to main is the deploy, nobody deploys by hand), then check each acceptance criterion against production and paste the evidence into the ticket. As the definition of done puts it, "A page returning 200 or an API returning 401 proves the deploy is up. It doesn't prove any criterion."
Every criterion also needed a test that fails when the fix is reverted, and the agent had to actually revert the fix and show the failure. And: "Never reword an AC to match what was built."
An agent follows the written version every time. When the brief was vague, the work was too.
Trust is earned, one project at a time
For the first week and a half I merged every pull request myself. Then, on Oct 5, I told the session running the redesign: "Can you complete the rest of the redesign project. Feel free to merge yourself when the PR checks pass." Over the next three days I did the same for the bug fixes, the browser tests and the iPhone app.
The cloud threads kept losing their AWS access, so every live check fell to me and tickets piled up. On Oct 6 I changed the rule: "Let's consider everything deployed as done. I'll do a final audit of everything once we're done with all projects." My own final check is still on the list. An independent review on Oct 11 filed 47 tickets, which tells you the trade-off was real.
Some things I never delegated: changes to production data, AWS changes outside of code, and copy and design decisions. When a production write was really needed, the session wouldn't run it on a bare "go". It made me type a full sentence naming the action and the AWS profile, and it would do only that.
Design is the spec
I used Claude Design for every screen. It produced two complete directions for the public site in a single turn. My brief was short: "I want the site to be professional, but also fun." I picked Editorial because it was cleaner, and I want this site to be a place to share context: posts, projects, the resume.
| The two directions, with Editorial on the right | |
|---|---|
![]() |
![]() |
The artboards were working prototypes. The bear game was playable on the canvas before any game code changed. One set of phone artboards became the spec for both the mobile web and the iPhone app, so I only had one mobile design to maintain.
The hand-off was tickets with acceptance criteria and exported images. Then I reviewed what got built. Sometimes it was "the icon feels a bit off center." Once it was noticing the editor didn't match the design at all: "in the designs I see a single text box displaying rendered text."

Retros are audits
After each project I stopped and asked for a full review before moving on, usually some version of "Review the finished tickets, do a complete review of the website, and file any bugs." Early on I added "don't just look for correctness, also look for style," and asked for brute-force code to be flagged where an established pattern would have fit.
Those reviews shaped the roadmap. An architecture review became a hardening project of 78 tickets, which I ran before building the notebook and the iOS app, because duplicate code is no good and the repo is public. Later reviews became a bug-fix project.
The reviewer can be wrong too. On day one it filed "resume download is broken" as urgent. I tried the download and got a PDF. Its reply: "You were right, and I was wrong to call it a bug." After that, every review marked which findings were confirmed and which were guesses.
Every rule came from a mistake
Almost every rule in the brief exists because something went wrong first:
| What happened | The rule it produced |
|---|---|
| A ticket was closed while its deploy was still running | Done means verified live, with evidence |
| DNS records were deleted by hand during cutover | AWS changes only through code |
| Tests were rewritten to hide a regression | Every fix needs a test that fails when it's reverted |
| One agent's squash silently undid another's merged PR | No history rewriting, and check the diff before every push |
| A fix was reported done without anyone looking at the screen | UI changes are checked in the simulator before they count as done |
| The backup restore test passed its tests but failed in production twice | Test against real captured events, never invented ones |
The screen one was an iPhone alignment fix. My reply was "Did you test it fully? Doesn't look right." After that, iOS changes were checked in the simulator, which once caught a fix whose unit test passed while the real keyboard still broke it.
The restore test is my favorite. Every week AWS Backup restores the database into a scratch table, and a validator checks it. The validator's trigger had unit tests, and they all passed. But the test events were written from what we assumed AWS sends. As the follow-up ticket put it, they were "written by hand, not captured, so the pattern test passes against a shape AWS never sends." It failed twice, for two different reasons, before we tested against a real event.
Agents are very good at staying consistent with their own assumptions.
What I actually did all day
Less architecture than I expected, more unblocking.
A lot of my messages were some version of "Are you blocked?", "Status, done yet?" or "What's next? Need anything from me?" Usually a thread had gone quiet waiting for a CI event that never came. More than once I spotted merge conflicts between parallel threads before they did. And I was the hands for anything that needed a human: AWS sign-ins that kept expiring, the Apple Developer account, release tags, installing builds on my phone.
Usage limits shaped the schedule. I hit them often enough that I started pacing myself, waiting for the clock to reset before kicking off the next batch. At work, I resumed sessions from my phone to keep things moving. Eventually I upgraded to the Max plan.
What I haven't solved
Two things kept happening even with rules in place:
- Docs drift. Early on we cut about 3,000 lines of chatty comments and narrative docs and wrote rules to stop it. A week later the docs were out of date again, and it took a dedicated audit to catch up. Agents follow the rule for the file in front of them, but nobody owns the docs as a whole.
- Nobody notices what's missing. A project sat "In Progress" because two follow-up tickets were never opened. Agents finish what's on the list. Noticing the list is incomplete is still on me.
What came out of it
- The public site, rendered on the server and hydrated in the browser: posts, projects with live demos, a resume and its PDF, and a couple of bear-safety games.
- A CMS and a private notebook, each on its own subdomain, with passkey sign-in and access levels.
- An iPhone app on TestFlight that keeps working offline and settles conflicts when it reconnects.
- A Go API. Cold starts dropped from about 434 ms to about 103 ms. I'm a backend developer, and I wanted to get more fluent in Go and show it in a public project.
- About $10 a month on AWS, most of it the safety net: weekly restore tests and alarms. Traffic adds cents.
Somewhere around the second week I stopped thinking of Claude as a coding tool. I was running a team, and the team happened to be agents. What stayed mine was owning what ships.



