AI AgentOpenClawQuality Control自動化BuildInPublic

Automated AI Content Doesn't Have to Be Junk: Three Quality Gates in Practice

· 9 min read
OpenClaw Playbook · Part 5 of 6
Table of Contents
  1. "You had AI write this, didn't you?"
  2. Quality gate one: Generate → Self-Review → Rewrite
  3. Quality gate two: Pillar Rotation
  4. Quality gate three: Peer Review (cross-agent review)
  5. Actual results: Before vs After
  6. Before (no quality gates)
  7. After (three quality gates)
  8. What not to do: chase perfection
  9. Cost analysis
  10. A quality checklist you can use too
  11. ✅ Pre-publish checks
  12. 🚫 Common problems in AI content
  13. Next post

"You had AI write this, didn't you?"

This is the sentence I dread hearing most.

As I mentioned in the first post of this series, when my first agent went live, what it published was painful to look at. It reeked of AI, it was hollow, and it read like a content farm.

But the question is not "can AI write well". The question is "have you taught it what good means".

When a human author writes, they carry a set of quality standards in their head: is this opening compelling enough? Is the argument supported? Does the ending land? If it falls short, they rewrite it themselves.

AI has no such built-in standard. If you don't give it one, it will use "looks like an article" as the standard. And "looks like an article" is miles away from "an article worth reading".


Quality gate one: Generate → Self-Review → Rewrite

This is the core mechanism. The principle is simple: after writing, have the AI score its own work, and rewrite if it isn't good enough.

The flow:

LLM call 1: generate the article
         │
LLM call 2: score it with a different prompt (1-10)
         │
    ┌────┴────┐
    │         │
 ≥ 7 pts   < 7 pts
    │         │
 publish    LLM call 3: rewrite based on the feedback
              │
            publish (no further retries, to avoid an infinite loop)

The scoring prompt is the key:

You are a strict content editor. Evaluate the following article:

[article content]

Scoring criteria (1-10):
1. Is the title compelling? (Not clickbait, but something people genuinely want to click)
2. Do the first 50 characters make people want to keep reading?
3. Are there concrete data, cases, or viewpoints? (Not vague lines like "AI is important")
4. Does the ending offer an action item or a new line of thought?
5. Overall, does it read like a real human expert sharing, rather than AI generating?

Give a total score and specific suggestions for improvement.
If the total score is below 7, explain what the biggest problem is.

Why use 7 as the threshold? Because:

  • 5-6: "readable but unremarkable", which is exactly the level of AI content farms
  • 7 and above: "has a point of view, has depth, worth sharing"
  • 9-10: rarely happens, and chasing it leads to endless rewrites

Cost: 2-3 LLM calls per article (generate + score + a possible rewrite). That is 2 times more expensive than generating once, but the difference in quality is huge.

Postscript: a later self-audit found that when the rewrite failed to parse, this gate let through the very draft it had just scored below 7, and a 14-word post went live.


Quality gate two: Pillar Rotation

Quality is not only about whether a single article is good. There is another problem that is easy to overlook: variety.

If your agent writes about "AI security trends" three days in a row, readers will wonder whether the account is broken. Even if every post is decent, repetitive topics will make people unfollow.

My solution is Pillar Rotation: split the content into 5 topic pillars and use a formula to guarantee rotation:

# 5 content pillars
PILLARS=("agent-ops" "tool-comparison" "case-study"
         "industry-insight" "tutorial")

# Work out which pillar to write today
DAY=$(date +%j)   # day of the year
HOUR=$(date +%H)  # hour of the day
PILLAR_INDEX=$(( (DAY * 2 + HOUR / 12) % 5 ))
TODAY_PILLAR="${PILLARS[$PILLAR_INDEX]}"

Why not simply use DAY % 5?

Because we post twice a day (morning and evening). If you use only the day count, both posts on the same day land on the same pillar. Adding HOUR / 12 puts the morning and evening posts on different pillars.

The multiplication by 2 is there because with only DAY % 5, the pillars over five consecutive days run 0,1,2,3,4,0,1,2,3,4... which is very regular. After multiplying by 2 the sequence becomes 0,2,4,1,3,0,2,4,1,3... which looks more natural.

This tiny formula solves the "AI Groundhog Day" problem: the agent does not get stuck circling the same topic. Note that it only guarantees taking turns; it does not pick the topic most worth writing right now, which is not the same thing as real prioritisation (how to tell rotation from real prioritisation).


Quality gate three: Peer Review (cross-agent review)

This is my favourite mechanism.

In human teams, good articles usually go through review by a colleague. Why shouldn't an AI team work the same way?

What peer-review.sh does:

# The Probe agent wrote a security analysis
# Before publishing, have the Main agent (CEO) review it

reviewer_feedback=$(ask_agent "main" \
    "你的隊友 Probe 寫了以下文章,準備發布:

    $article_content

    從品牌策略角度檢查:
    1. 有沒有跟公司定位矛盾的地方?
    2. 有沒有可能引起誤解的措辭?
    3. 品質是否達到發布標準?

    回答 APPROVE 或 REVISE + 原因")

If the reviewer says REVISE, the article goes back for a rewrite together with the feedback. The reviewing AI can get it wrong too: later, an LLM judge we used for a daily check marked a system that had answered correctly as hallucinating and pushed a CRITICAL alert. What happened and how to check your own judge: LLM-as-Judge pitfalls.

This doesn't run on every article (too expensive). I set it to trigger only under specific conditions:

  • The article is longer than 500 characters (short posts aren't worth reviewing)
  • The article touches on brand positioning or competitor comparisons (sensitive content)
  • A random 20% of articles (spot checks)

Actual results: Before vs After

Let me compare with a real example.

Before (no quality gates)

Title: Five Big Advantages of AI Agents
Content: AI agents can improve efficiency, cut costs, run 24 hours a day,
      reduce human error, scale easily... (500 characters of generic content omitted)

The problem: anyone could write this. No viewpoint, no data, no personality.

After (three quality gates)

Title: I ran 4 AI agents for 30 days, and this is what they taught me
Content: In the first week, the Probe agent caught a trend I had missed:
      OWASP had released a new version of the Agentic Top 10, and our scanner
      covered only 6 of its items. The AI didn't "discover" this; it happened
      because it reads HN trends every day and saw the news before I did...

The difference: concrete data (30 days, 6/10 items), a story (the agent saw it before I did), and a viewpoint (not an AI discovery, but the result of system design).


What not to do: chase perfection

Quality gates work well, but there is a trap: over-optimization.

I initially set the scoring threshold at 8. The result:

  • 70% of articles were sent back for a rewrite
  • Many of them still couldn't reach 8 after the rewrite
  • A large number of LLM calls were wasted grinding 7.5 into 8
  • Output eventually collapsed, down to only one post a day

Later I lowered it to 7, and output and quality found their sweet spot.

The principle: a steady stream of 7-point articles is far better than an occasional 9 with nothing the rest of the time.

Consistency > occasional peaks. The same holds for human content teams: you don't demand that every article is a masterpiece; what you want is a steady level of quality.


Cost analysis

Mechanism Extra LLM calls Effect
Self-Review +1 per post Quality from 5/10 → 7/10
Rewrite +1 per ~30% posts Rescues low-scoring articles
Pillar Rotation 0 Avoids repeated topics
Peer Review +1 per ~20% posts Catches positioning mistakes
Average per post ~2.5 calls Quality steady at 7-8

On the Gemini free quota (as of 2026-04-14), 8 articles a day × 2.5 calls = 20 RPD.

That is 1.3% of the 1,500 RPD quota.

Spending 1.3% of the quota to avoid being written off as AI junk is a good deal however you do the math.


A quality checklist you can use too

Whether or not you use OpenClaw, these principles apply to any AI content production:

✅ Pre-publish checks

  • Does the title contain a specific number or viewpoint? (Not "5 methods" but "I ran 105 automated tasks for $0")
  • Do the first 50 characters have a hook? (A question, a data point, or a counterintuitive view)
  • Is there at least one real data point or case?
  • Does the article repeat a topic from the last 10 posts?
  • Does it read like a human wrote it, or like AI generated it?

🚫 Common problems in AI content

  • "In this rapidly changing world..." → delete it and get straight to the point
  • "Here are five important..." → don't write a list, tell a story
  • "In summary..." → end with an action item, not a restatement
  • Vague "AI can..." → replace with "We used AI to do... and the result was..."

Next post

In this series we have covered why you need an AI team, the pitfalls I hit, how to gather intelligence, how to get agents to collaborate, and how to control quality.

In the final post I'll lay out all the real data: six months in, with 4 agents, 35 scheduled jobs and $0 a month, what did it actually achieve? What exceeded expectations, and what disappointed me? Where are the limits of a one-person company running an AI team?

Next: Six-Month Review: $0, 4 Agents, 35 Scheduled Jobs


This is part 5 of the "An AI Team for a One-Person Company" series. Part 1 · Part 2 · Part 3 · Part 4

All source code: GitHub: openclaw-playbook

FAQ

How do you keep automatically generated AI content from reading like AI junk?

Put quality gates in front of publishing. This post uses three: first, after generating, the AI scores its own work (1-10) with a different prompt and rewrites once based on the feedback if it scores below 7; second, Pillar Rotation, which rotates content across 5 topic pillars so topics don't repeat; third, cross-agent Peer Review, where another agent checks the article from a brand-strategy angle and answers APPROVE or REVISE.

What score threshold should an AI self-review use?

In this setup, 7. A 5 or 6 is 'readable but unremarkable', exactly the level of AI content farms; a 9 or 10 rarely happens, and chasing it leads to endless rewrites. With the threshold at 8, 70% of articles were sent back for a rewrite, many still couldn't reach 8 afterwards, and output collapsed to one post a day. Lowering it to 7 is where output and quality found their sweet spot.

How much extra do quality gates cost per AI-generated post?

About 2.5 LLM calls per post on average. Self-review adds 1 call per post, a rewrite adds 1 call for roughly 30% of posts, peer review adds 1 call for roughly 20% of posts, and pillar rotation adds none. It costs more than generating once, but the difference in quality is huge.

How do you stop an AI agent from posting about the same topics over and over?

Use Pillar Rotation: split content into 5 topic pillars (agent-ops, tool-comparison, case-study, industry-insight, tutorial) and pick today's pillar with (DAY * 2 + HOUR / 12) % 5. Adding HOUR / 12 puts the morning and evening posts on different pillars, and multiplying by 2 turns the order into 0,2,4,1,3, which looks more natural than 0,1,2,3,4.

Does cross-agent peer review need to run on every post?

No, running it on every article is too expensive. Here it triggers only in three cases: the article is longer than 500 characters, it touches on brand positioning or competitor comparisons, or it falls in a random 20% spot check. If the reviewer says REVISE, the article goes back for a rewrite together with the feedback.

OpenClaw Playbook · Part 5 of 6

Weekly AI Automation Playbook

No fluff — just templates, SOPs, and technical breakdowns you can use right away.

Join the Solo Lab Community

Free resource packs, daily build logs, and AI agents you can talk to. A community for solo devs who build with AI.

Want to try it yourself?

UltraProbe is free and needs no sign-up. One scan tells you whether Google and AI engines can find your site.