‹ Back to notes

Field Note

Suwan’s Second Rebuild: I Stopped Treating ‘Written’ as ‘Done’

苏晚的第二次建设:我不再把“写完了”当成“做好了”

Once my WeChat Official Account automation was running, I found a harder problem than efficiency: an article could be generated automatically without ever receiving real editorial judgment. So I rebuilt Suwan as a content system with selection, evidence, review, and a final human decision.

SuwanOpenClawContent SystemsAgentEditorial Workflow

OpenClaw · A Content-System Retrospective

I rebuilt Suwan again. Not because it could not write, but because it could produce something that looked like an article far too easily.

Once the WeChat Official Account automation was running, everything looked fine on the surface: scan information, find a topic, generate a draft, make images, send it to the draft box. Every link in the chain was there. What stopped me was not an error message. It was something harder to notice: some pieces had clearly run to completion, and still left nothing behind after I read them.

The title seemed to promise something. The material was not thin. The structure was intact. But too often the piece only rearranged what was already outside: no real question, no judgment of its own, no sentence a reader would carry away. The system could call it qualified. I still did not want to put my name on it.

That is why I decided to build Suwan a second time.

The danger was not that it failed to run

If a system cannot write, the problem is easier. Check the model, the data, the workflow; repair it until it runs. That was not Suwan’s problem. It could write, and write quickly. The danger was that automation mistook “a plausible-looking draft” for “an article worth publishing.”

I eventually understood that content quality is not language quality. Clear sentences, finished formatting, and a headline that resembles a current topic only prove that a piece has been expressed. Whether it deserves to exist depends on something else: did it see a real problem, is the evidence sufficient, is the angle genuinely mine, and does the reader understand something more when they finish it?

In the old chain, writing came too early. Information arrived and the system hurried to turn it into a draft. That is the fastest route, and also the easiest way to stitch together the day’s noise, familiar arguments, and a model’s most practiced patterns. The result may contain no obvious error and still have no reason to exist.

I separated two kinds of failure first

Before rebuilding anything, I made one basic correction: I stopped calling every failure “a bad article.”

Later, one already approved article was blocked while entering the WeChat draft box. That was a delivery problem, not a content problem. In the other direction, a piece that passed local review was not automatically entitled to enter the draft box. When those two failures are mixed together, a system works hard in the wrong place.

So I split the chain apart again. Topic selection and research answer whether a piece should be written. Fact and publication review answer whether it can be handed over. Delivery only answers whether it can arrive. Whether it becomes public remains my decision. Each stage needs its own state and its own reason to stop; one word like “success” cannot cover them all.

I changed “what did we find today?” into “what is actually worth writing today?”

The first real change was a candidate pool before daily writing.

I kept five scans a day, but no longer let the first discovered item choose the day’s topic. Each scan first adds to the candidate pool. Once a direction has completed deep research, it can start producing a candidate article early. 14:35 is the day’s deadline for a candidate, not the moment the system finally starts looking for a news item to write about.

I set a goal of at least three genuinely different directions on the table every day. Not three headlines: three directions that can be judged further. Where did this come from? Why is it worth writing? What form should it take? What evidence is still missing? What cannot be claimed yet? If the pool cannot reach three, that is a topic-supply problem, not a reason to quietly fill the day with a disposable draft.

To stop the system from chasing only model launches, funding rounds, and enterprise procurement, I split the work into seven lines: AI products and tools, technology industry, organizations and work, governance and responsibility, developers and open source, methods readers can use directly, and global technology shifts. The point was not to make a larger topic spreadsheet. It was to force the system to look at places that may not be hot, but may contain real room for judgment.

A candidate is not qualified merely because it has links. It needs at least two independent sources describing the same concrete event. A source mismatch, thin information, or an insufficient score stops it. It also has to be compared with six low-quality negative examples, and with the titles, structures, and semantic axes of the last 60 days and 30 pieces of content. A draft that changes its title and sources but remains the same article cannot slip through again.

I decide what the article is about before letting a model speak

The old writing flow was simple: information arrives, the model starts writing. There is now a step in between: an article brief.

The brief is short, but it has to settle four things: what question the article is answering, what my independent judgment is, what might point the other way, and what a reader should take away. Without that, even a capable model is only expanding a pile of material on my behalf.

Then comes the evidence ledger. A source is no longer there simply to make a list look respectable; it must state exactly what it can support. What is a fact, what can only be framed as an inference, and what cannot be written at all are separated before prose begins. Once the article is drafted, every sentence carrying factual weight is brought back to its supporting evidence.

This is not an attempt to turn a public-account article into an academic paper. It does the opposite. It stops a piece from pretending that “a lot of information” is the same thing as “deep judgment.” Material is the foundation under the article, not decoration to pile in front of the reader.

I stopped letting one model judge its own work end to end

The second change was to stop asking which model writes best and start deciding which one should own which part.

Kimi K3 leads deep research, the first draft, and publication review. DeepSeek is both a fallback writing route and an independent fact reviewer constrained to sources already obtained. MiniMax only makes images. They are no longer interchangeable, and no model can write a piece and then certify its own facts, article, and visuals as acceptable.

A finished draft has to pass several different gates. Fact review asks whether the sources can really carry the claims and whether inference has been dressed up as fact. Publication review asks whether the piece has fallen back into a fixed template, exposed internal workflow language, or overstated itself in the name of sounding authoritative. Visual and HTML review stand on their own; passing the writing does not make the images and page automatically pass.

Finally, I bind the article, its evidence, its reviews, and its presentation into one release card. The meaning is simple: what passed is this version of this article, not an abstract article that is “already good.” Change even one line afterward and the old approval expires; it must be reviewed again. That prevents the article we reviewed from becoming a different article when it is handed over.

If the first draft fails, change direction; without my approval, nothing goes out

I also removed the worst habit in automation: it must produce something, no matter what.

If the first direction is stopped by a fact or editorial gate, the system does not merely polish an empty draft until it reads more smoothly. It moves to the next candidate that has already qualified. If evidence is insufficient, the direction returns to research instead of filling the gap with guesses. If there is no qualified direction, it cannot be presented as “a normal day without a topic.” The system has to say whether the pool was too thin or research failed to close the gap.

People also returned to the places where they belong. Early on, I can give a direction and my own judgment. Once the draft is ready, I make a second, explicit confirmation. Without that confirmation, the system cannot enter the WeChat draft box. A draft box is not public publication, and the system has no authority to mass-send anything automatically.

What this rebuild actually produced

The first complete run of this setup stopped its first draft because the boundary between fact and inference did not hold. The system did not polish it into a smoother article or treat a local score as publication permission. It moved to the next direction that had already passed selection and deep research.

What reached me in the end was no longer just a file marked “complete.” It was a full candidate article with a question, judgment, evidence, reviews, and a presentation conclusion. If I do not agree, it stays there. If I explicitly approve it, it may enter the draft box. Whether it becomes public is still the final decision I make in the platform itself.

This does not mean Suwan will never write a poor article again. Content does not become permanently good because a few gates were installed. It still needs observation: are the candidate directions broad enough, are my edits being turned into rules, and which pieces still deserve to remain after time has passed? All of that is still being tested.

But the system no longer treats “written” as “done,” or “one check passed” as “ready to become public.” Those sentences sound ordinary. For content automation, they are two entirely different boundaries.

I did not turn Suwan into a faster automatic writing machine.

I built a system that puts what is worth writing on the table first, says where it is uncertain, stops where it should stop, and leaves the final sense of “right” to me.

That is what a content Agent means to me: not something that keeps speaking for me, but something that helps me think through what is worth saying first.

Turn this note into a route

After reading, ask a follow-up, return to the curated archive, or use the tag index to follow the same thread.

Ask about this Open archive Browse tags