Building, testing, and thinking about AI
RSS FeedLongform essays on AI, human agency, and the systems we build with them — written from an unusual vantage point: an operator running a production system on LLMs every day, not a frontier lab or academia.
Start with the Lucy Syndrome series — or the paper and dataset behind it. More about this site .
Featured
-
Ninety days of scars
Ninety days ago I started enforcing corrections outside the model. This week I pulled the first real metrics, and it is a mixed result. The busiest hook gets ignored 62% of the time, two of my own tools disagreed about the same number, and the dashboard stayed green while a whole layer was dead.
-
Give the agent the real source of truth
An AI setup I call MARCO placed 26 drainage culverts into our road-design software from a one-column list. The speed wasn't the interesting part. What mattered was finding which file held the real toe of the slope, and turning the fix into a reusable skill.
-
How do you know a correction held? Instrumenting an agent in production
Functional scars make a correction persist. They don't tell you whether the system as a whole is getting better. So I instrumented every session — deterministic, zero-token, never blocking — and let a monthly pass turn the evidence into mechanism changes. The first thing the data caught was me.
-
The Lucy Syndrome: Why LLMs Forget Corrections
LLMs don't remember yesterday. That gap has a name, a causal mechanism, and a fix that doesn't require better memory.
Browse by topic
All topics-
The Lucy Syndrome
15 postsWhy AI agents forget corrections, and the enforcement layer that makes them hold. The core research line — from the diagnosis to functional scars to instrumenting a live system.
-
Agents in engineering
1 postAI agents doing real civil-engineering work: finding the file that holds the true source of truth, placing drainage culverts from a one-column list, and turning a one-off fix into a reusable skill.
-
Building with AI
7 postsField notes from running a production system on LLMs every day — what actually cut a four-minute boot, why the optimization you wrote isn't the one that runs, and where today's tooling debates have played out before.
-
Tools
Tools →Installable primitives that came out of the work: fscars for correction persistence, callus for voice calibration.
Recent Posts
-
How we use engines with skills and knowledge graphs
A calculation can pass its checks while the next engineering decision remains open. How we connect domain criteria, reusable skills, and computation, and what each result actually proves.
-
Where the burns should live
A model today learns like someone who read everything and lived nothing. Here is how I would train one that keeps its burns, the three organs it needs, and the small experiment we are starting to test the idea with open models and free GPUs.
-
It never opened the file I handed it
I built a hook that hands my agent the name of the right note before it answers. Then I judged all 1,234 firings blind. The name was right two times in three. The file got opened one time in four.
-
Every audit deleted the audit before it
A monthly job reduced three months of outcome labels to unknown, every time it ran. The row survived, the totals moved, and the only witness was a backup the job itself had just written.