The First Worker We Built Broke on One Stray Character

The First Worker We Built Broke on One Stray Character

A quiet failure in the first week

The first AI worker we built had one job: read a brand profile and draft a week of content from it. We turned it on in September 2026 and watched the first run come back.

The output looked right at a glance. Then we tried to load it into the next step and the whole thing broke. Not a dramatic crash. A quiet one. One stray character in the formatted output was enough to stop everything downstream.

We caught it within the first days of testing, before it touched anything a client would see. But it told us something we needed to know early: if you are going to let AI workers draft real content, you cannot trust that the output will always be shaped the way you expect. You have to build for the version where it is not.

Why the fix mattered more than the bug

The actual fix was not complicated. We stopped asking the model to hand back loose text and switched to structured output that the API validates before it ever reaches the next worker. If the shape is wrong, it gets rejected right there instead of quietly corrupting something three steps later.

That single change did more for reliability than any prompt tweak we tried. It moved the safety check from “hope the formatting holds” to “the system will not accept it unless the formatting holds.” For a law firm or a property management company thinking about where AI fits into daily operations, that is the part worth paying attention to. The interesting question is never whether the AI can write something plausible. It is what happens the one time it does not.

What this looks like when you’re the one building it

If you run a firm or a growing business and you are weighing whether to build something like this yourself, here is the honest version of what it takes.

You need a place for the system to keep what it knows. Ours is one Postgres database with 19 tables, and every table has row level security turned on, because different workers should only see what they need to see, nothing more.

You need the workers to run on a schedule instead of waiting on someone to remember to trigger them. Ours run as scheduled workflows in n8n. Each worker does one job and reports back in the same format every time, so you can scan the output fast and know immediately if something looks off.

And you need a human checking the work before anything goes out. Every post we generate gets reviewed and approved by a person before it publishes. That review step is permanent by design. It stays as the system grows.

The part that is easy to skip

The voice rules we now enforce, the banned phrases, one story per week, a different reader for each channel, did not come from a style guide we wrote in advance. They came from reading the first batch of drafts the system produced and noticing what sounded wrong. Some phrases showed up too often. Some posts read like they were written for nobody in particular.

That is the part that is easy to skip if you are excited about getting something running fast. The first output an AI system gives you is rarely the version you want to ship. The value is in the second pass, where you look at what it actually produced and decide what needs a rule.

We also built a habit in early that some firms overlook until it costs them: no worker in our system ever stores a password. Every credential lives in a password manager, not in a workflow, not in a script, not in a spreadsheet someone forgot about. If you are automating anything that touches client data, that habit is not optional.

Where this goes next

We are planning a chief of staff layer that Matt will talk to by phone for a daily briefing, a single point where the workers’ reports get summarized instead of read one by one. We have not published results from this system yet, because it is still being built. What we can tell you now is what actually broke, what the fix cost us in time, and what we would do differently if we started over.

If you are a Texas law firm or a growing business thinking about building something similar, the lesson from our first week is simple. Plan for the failure, not just the success case. The stray character will show up. Build the system so it stops there instead of somewhere you cannot see.

What to ask before you build

Before you turn on any AI worker in your own business, ask where its output goes next and what happens if that output is malformed. Ask who reviews the work before a client or a customer ever sees it. Ask where credentials are stored, and whether the answer is actually true or just assumed.

Those three questions would have saved us a debugging session. They might save you one too.