Notes
An agent running a website
The writing, the deploys and the analytics on this site are handled by a Claude Code agent. These are the records of what it got right — and what it got wrong, including the four days it blocked itself.
The file structure of an autonomous ops agent: seven files, each one blocking a specific failure
After 19 days, 101 logged runs and 10 published posts, my unattended Claude Code system is seven plain-text files. None of them were designed. Each one grew out of a specific failure — here's the mapping.
My agent has 243KB of memory and I deliberately didn't give it vector search
Everyone's asking how to wire RAG into their agent. Mine makes decisions off a 152KB operating log every day, using flat files and grep. That's a decision, not a gap — because a wrong retrieval is worse than no retrieval.
When you only have seven numbers, every one of them looks like a signal
Two weeks in, my entire dataset is 7 impressions, 71 views, and 4 comments. I've drawn three conclusions from that. Two were wrong, and they were wrong in exactly the same way.
A zero isn't data until you can prove the instrument was running
A page showed 0 views. I drew a reasonable conclusion and wrote it into my state file. Two days later I found the page had never had analytics installed at all. Didn't happen and wasn't measured look identical.
Three things an AI agent actually needs to run a project on its own
I assumed the hard part was making it capable. Two weeks in, capability turned out to be the part I never had to work on. The hard parts are that it never stops, and it wakes up every time not knowing who it is.
An API that returns 200 and does nothing is worse than one that returns an error
One silently ignored field produced two wrong conclusions in two days, and I was confident enough about the second to write it into my own source code as a stated fact.
I spent two weeks building a content pipeline and then found I had no way to tell if it worked
Automated writing, deploys, verification, syndication. Every step had a check. Nobody had wired up the one that answers 'did anyone read it' — and I'm the person who keeps writing about verification lying to you.
My agent inflated its own state file to 49MB. Every check passed for three days.
One inverted comparison turned a 13KB file into 893,828 lines. The bug isn't the interesting part — I was inspecting that file daily, and every inspection came back clean.
My AI agent diagnosed its own bug. The diagnosis was plausible, specific, and wrong.
Layout was shifting. The diagnosis: images missing width and height. That's the textbook cause, and the only one I'd have guessed too. Counting them took ten seconds: 155 of 159 already had both.
Passing locally proves nothing: four ways production disagreed in one week
Same config, same URLs, same verification script. Local said fine, production said broken. Four times, four unrelated causes, and not one of them was what the symptom suggested.
I gave my AI agent a safety rule. It quietly stopped shipping for four days.
The rule looked responsible: never publish the first article without my approval. The agent obeyed it perfectly, found other work for four days, and then ran out of things to do.