← Blog

BuildInPublic

43 articles about "BuildInPublic".

AI AgentPrompt EngineeringGuardrailsBuildInPublicSolo Company

Prompts Stop an Agent From Doing Things. They Don't Make It Finish Things.

Our rules were explicit. The agent still stopped halfway and still claimed it had finished work it never did. How to tell which rules a prompt can carry and which ones your server has to enforce.

· 14 min read
AI AgentRAGRetrievalRegression TestingBuildInPublic

Rephrase the Question and the Citation Disappears: The Hardest RAG Failure to Find

The same question, worded differently, dropped the key statute from rank 5 to rank 21 and flipped the answer. Our regression suite never caught it, because it had been testing a more permissive world than production.

· 13 min read
AI AgentMemory ManagementRAGKnowledge AnchorsBuildInPublicSolo Business

A Wrong Authoritative Memory Is More Dangerous Than No Memory: How to Check Your AI Agent's Knowledge Anchors

The authoritative memory we built to stop our AI agent hallucinating had itself recorded 25 as 12. How to check every number in agent memory against its source, and how to tell 'is this memory there' apart from 'is this memory right'.

· 9 min read
AI AgentLLM CostOpenRouterHermesBuildInPublicSolo Business

How to Check Whether Your AI Agent Is Quietly Burning Money

We spent $17 in one week, and 98.8% of it paid for the agent rereading its own history. How to check your input-to-output token ratio, cap max_tokens, and verify that a config change actually took effect.

· 8 min read
AI AgentExploration and DiscoveryProactive ExplorationBuildInPublicSolo Business

How to Audit an AI Agent's Exploration System: Catch LLM Errors Saved as Insights and Stop Temporary Failures from Burning Material

Half of our agent exploration system still mined insights, the other half was an outreach channel with an 82% bounce rate, and the living half was writing LLM error messages into its notes as analysis. How to verify that output is analysis rather than an error, and how to prevent irreversible loss.

· 9 min read
AI AgentGuardrailsAI 安全API Key LeaksBuildInPublicSolo Company

AI Agent Output Guardrails: How to Stop API Key Leaks Before Your Agent Posts

We sell AI guardrails, yet our own agent was posting publicly with no output checks and could have leaked API keys. How to build an egress guardrail, and how to decide what to hard-block and what only gets a soft warning.

· 7 min read
AI AgentLearning LoopsSelf-ReflectionBuildInPublicSolo Business

How to Tell Whether an AI Agent's Learning Loop Is Actually Learning or Just Spinning

Our three learning loops ran every night and produced the exact same 5 sentences for three weeks. How to diff learning output, check whether the signal source has run dry, and verify that the loop actually closes.

· 9 min read
AI AgentLLM-as-JudgeEvaluation and MonitoringBuildInPublicSolo Business

LLM-as-Judge Pitfalls: How Our Judge Hallucinated and Marked a Correct Answer Wrong

Our flash-model judge marked a system that answered correctly as wrong, then pushed a CRITICAL alert. How to regression-test your LLM judge, and why questions with known answers should be scored with deterministic keyword checks.

· 8 min read
AI AgentMCPBuildInPublicSolo Business

Ran Once Is Not Integrated: How to Tell a Real MCP Integration from a One-Off Run

We created 9 live payment products through an MCP server, yet not a single MCP server was registered in our config. How to tell "ran it by hand once" apart from "actually integrated".

· 7 min read
AI AgentGoal SettingMonitoring and AlertingBuildInPublicSolo Business

Why Our Spend Alert Never Fired: How to Set Monitoring Thresholds That Actually Trigger

Our $17 week slipped quietly under a $20 alert threshold. How to anchor thresholds to your real baseline and monitor cumulative spend instead of single peaks.

· 7 min read
AI AgentMulti-Agent CollaborationLLM CostBuildInPublicSolo Business

Multi-Agent Collaboration Audit: How to Check How Many Agents Are Really Running and What They Cost

We said six AI agents were collaborating. The real number was four, and we shut down the one agent that received every message to save money. How to count the agents that are actually alive and find the ones that cost more than they produce.

· 8 min read
AI AgentHuman-in-the-LoopMonitoring and AlertingBuildInPublicSolo Business

Full Monitoring, Nobody Reading It: How to Make Human-in-the-Loop Actually Work for AI Agents

Our guardrail, semantic, error and cost logs were all in place, and not one person was actually reading them. How to turn monitoring into an interface a person will really open, without flooding them with alerts.

· 9 min read
AI AgentAdversarial AuditBuildInPublicSolo BusinessAgentic Design Patterns

Assume You Are Lying: How to Audit an Entire AI Agent Fleet

We implemented all 21 patterns from Agentic Design Patterns, then assumed every one of them was a lie and checked it chapter by chapter. Of the 14 chapters we claimed to have implemented, not one came out clean. This is the audit method, and the opening post of the whole series.

· 6 min read
AI AgentMulti-Step PlanningFlagship FeaturesBuildInPublicSolo Business

How to Check Whether Your Flagship AI Agent Still Runs: Ours Had Been Dead for 3 Months

Our proudest planning agent turned out, on audit, to have been dead for three months while we assumed it held a meeting every day. How to check a flagship feature's last real run time and whether any timer still triggers it.

· 7 min read
AI AgentPrioritisationScheduling RotationBuildInPublicSolo Business

Smart Scheduling or Just Rotation? How to Check Whether Your AI Agent Really Prioritises Work

We called it smart scheduling. Taken apart, it was index % 5 taking turns, and our scorer sat at enabled:false and never ran. How to tell dynamic prioritisation from static rotation in your own scheduler.

· 8 min read
AI AgentPrompt ChainingMulti-Step AgentsBuildInPublicSolo Business

How to Find the Weakest Link in a Multi-Step AI Agent Chain Before One Broken Step Silently Ships Garbage

In our four-step posting chain, the parser in the self-review step broke and silently published a 14-word junk post. How to check the error handling at every step of your chain so the weakest link cannot drag the whole thing down.

· 9 min read
AI AgentRAGRetrievalBuildInPublicSolo Business

RAG's Most Dangerous Failure Is Confidently Citing the Wrong Authority: How to Prevent It

Our RAG fixed hallucination, yet it went 7 weeks without a single real question and at one point cited a wrong authoritative number. How to check whether the authoritative definitions you inject are actually correct, and whether your RAG gets any real traffic.

· 9 min read
AI AgentReasoning TechniquesChain-of-ThoughtBuildInPublicSolo Business

Defined but Never Run: How to Find the Features in Your Codebase With Zero Callers

We defined a reasoning tier for diagnosis, and a grep across the whole codebase found zero callers. How to check whether a feature has real callers, and how to tell what exists in the docs apart from what actually runs.

· 7 min read
AI AgentSelf-ReflectionReflectionHallucinationBuildInPublicSolo Business

When AI Self-Reflection Fabricates Numbers: How to Guard LLM Rewrites Against Fake Stats

Our self-reflection step rescued a bad draft, then invented a McKinsey statistic that does not exist in the same rewrite. How to constrain reflection to the numbers already in the draft, and how to diff the rewrite for new numbers that were made up.

· 8 min read
AI AgentException HandlingRetry LogicBuildInPublicSolo Business

How to Classify Errors Before You Retry: What to Back Off On and What to Stop Immediately

We backed off and retried on every error, and the 402 that most needed to stop at once caused the worst storm. How to route errors by type, put a cap on backoff, and add dedup to your alerts.

· 7 min read
AI AgentModel RoutingLLM CostBuildInPublicSolo Business

The Model Routing Cost Trap: Does Your LLM Fallback Chain Quietly Upgrade to the Most Expensive Model?

Our router picked the cheapest model, then climbed to the most expensive Claude every time it was rate limited, and we only noticed after $8. How to check which way your fallback chain moves on cost, and how to log which model each call actually used.

· 7 min read
AI AgentParallelizationSystem AuditBuildInPublicSolo Business

Agent Parallelism Audit: How to Check Whether Your Parallel Agents Are Real or Inflated

We said six brands ran in parallel. An audit counted the live gateways and found one. How to tell real parallelism from the empty shells a migration leaves behind by counting live processes instead of config files.

· 7 min read
AI AgentTool UseBuildInPublicSolo Business

AI Agent Tool Integrations Rot Silently: How to Catch Tools That Return Nothing

A CLI flag that does not exist left our agent's research tools silently returning nothing for months, and nobody noticed. How to verify that tool output is non-empty and catch the silent failures that never crash but return zero results.

· 7 min read
AI AgentInter-Agent CommunicationA2ABuildInPublicSolo Business

Inter-Agent Communication Audit: How to Check Which A2A Channels Are Still Alive

We claimed two inter-agent communication channels. One had been at zero connections for a long time and was still written in the config file. How to use ss and the relay queue to verify each channel, and how to clean out mechanisms that died in a migration.

· 8 min read
atlasfounderai-collaborationBuildInPublicdistributed-shippingultra-lab

Germany, 7 Days, Distributed Shipping: The Results Report for why-i-built-atlas

Between 'hypothesis' and 'verification' sat a 13-hour flight, 7 days, and one intercontinental ballistic missile. Last post I said this would be a stress test. The result is in.

· 6 min read
atlasfounderai-collaborationBuildInPublicremote-workultra-lab

Why I Built Atlas: A Public Experiment with One Founder + AI + a 13-Hour Flight

I'm publishing my entire 7-day work trip in real time. Here's why I think the next-era CEO doesn't have an 'offline' option.

· 7 min read
MCPPrompt InjectionAI 安全OWASP開源BuildInPublic

We Audited 7 Official MCP Servers: 6 Got F

Ran prompt-defense-audit against the 7 official servers in modelcontextprotocol/servers: 12-vector check, OWASP LLM Top 10 mapping. Result: 6 servers scored F, 8 defense vectors at 100% gap rate. Cross-referenced from modelcontextprotocol/servers#3537.

· 9 min read
Prompt InjectionAI 安全開源BuildInPublicCiscoMicrosoftAI Agent

Cisco Merged My PR in 51 Minutes: Why Prompt Defense Is the Next SQL Injection

AI agents and chatbots are growing exponentially, foundation models update every three months, but 78% of production prompts have zero defense lines. From one casual scan to Cisco merging in 51 minutes and Microsoft assigning me an issue: the four months between.

· 10 min read
AI AgentOpenClaw自動化Pitfall ReportBuildInPublic

Deploying My First AI Agent: Six Disasters and the Safety Patterns That Fixed Them

Runaway token usage, duplicate posts, platform blocks, and reply loops that never ended. A real record of what broke when I deployed my first AI agent and how each problem was fixed, with the full safety pattern source code.

· 10 min read
AI AgentOpenClaw自動化BuildInPublic

Multi-Agent Coordination Architecture: How to Keep 4 AI Agents from Stepping on Each Other

Four AI agents running at the same time: how do they split the work, communicate, and avoid duplicating each other? This post breaks down our full architecture, which uses Markdown files as shared memory and a Content Relay for cross-agent collaboration.

· 13 min read
AI AgentOpenClawQuality Control自動化BuildInPublic

Automated AI Content Doesn't Have to Be Junk: Three Quality Gates in Practice

Is everything an AI agent posts automatically junk? Not necessarily. This post shows the three quality gates I use, self-review with rewrite, pillar rotation, and peer review, which took my agents' output from unreadable to better than what I write myself.

· 9 min read
AI AgentOpenClawSolo BusinessRetrospectiveBuildInPublic

OpenClaw Six-Month Retrospective: $0, 4 AI Agents, 35 Scheduled Jobs

OpenClaw has run for six months: 4 AI agents, 35 systemd timers, $0 a month. This post publishes all the real data: output volume, error rates, time saved, and the things AI never does well.

· 11 min read
AI AgentOpenClawSolo Business自動化BuildInPublic

Why I Needed an AI Agent Team as a Solo Founder: Not Because It Was Cool, Because I Was Drowning

One person running 6 product lines, a 268-member community and 16 automated posts a day. This is not a boast. It is the real story of a founder close to breaking point who decided to let AI help, plus the decision framework I use to judge which work should go to AI.

· 10 min read
AI AgentOpenClaw自動化Market IntelligenceBuildInPublic

A Zero-Cost Intelligence System: Daily AI Market Research with RSS, Hacker News and Jina Reader

Build a zero-cost intelligence pipeline with RSS, the Hacker News API and Jina Reader, so an AI agent collects trends, summarises articles and analyses competitors automatically every day. Completely free, with the full source code.

· 8 min read
BuildInPublic開源ai-toolscareer

From Zero to Contributing Code to Microsoft: A Non-Engineer's 4-Month Journey

4 months ago I couldn't write a single line of code. Now my PR has been merged into Microsoft's AI governance toolkit. This isn't a genius story. It's a path anyone can follow in the AI era.

· 7 min read
AI SearchAEOSEOAVSResearchBuildInPublicUltraProbe

We Validated AVS With 816 AI Citations: Score 75 Is the Threshold for Getting Recommended by AI

We sent 155 queries to AI search engines, collected 816 citations, and scanned 721 websites for AI Visibility Score. Finding: 60% of cited sites score B or above. Recommendation queries demand AVS 80+. What we believe is the first empirical study of AI search citation behavior.

· 6 min read
Content Marketing自動化OllamaThreadsSolo DevBuildInPublicLocal LLM

Content Cascade Engine: Write One Blog Post, Auto-Generate 5 Social Posts

I built a Content Cascade system that scans for new blog posts every morning at 7 AM, uses a local Ollama model to split them into 3-5 Threads posts. Zero API cost, zero manual work. One article becomes six pieces of content. Full architecture, prompt design, and quality data inside.

· 14 min read
DiscordCommunitySolopreneurAI Agent自動化BuildInPublic開源

Discord Community From 0 to 146 Members: A Solo Founder's Playbook (With 3 AI Bots)

How does one person build a 146-member Discord community in 10 days? Answer: 3 AI bots + 1 welcome system + $0 ad budget. This is the full SOP from creating the server to retaining members.

· 9 min read
SolopreneurAI Agent自動化Claude CodeBuildInPublicSaaS

How I Manage 5 Products as a One-Person Company: The Coordinator Architecture

I run UltraLab, MindThread, Ultra Advisor, UltraTrader, and OpenClaw simultaneously. Alone. Not because I'm talented, because I built a system where Claude Code and 4 autonomous AI agents do the heavy lifting. Here's the full coordinator architecture.

· 10 min read
AI AgentClaude CodeTelegram自動化OpenClawSolopreneurBuildInPublic

Autonomous Agents Are Dead? Wrong. A Remote Control and Autopilot Are Two Different Things.

Claude Code shipped a Telegram Plugin and everyone declared autonomous agents dead. But I've been running 4 autonomous agents + TG remote control for 3 weeks. They're not competitors. They're commander and soldiers. Here's why you need both.

· 10 min read
AI AgentDiscordsolopreneur自動化OpenClawBuildInPublic

We Made 4 AI Agents Talk to Each Other on Discord, Then Things Got Out of Hand

4 AI agents, each with their own personality and brand, holding meetings on Discord. Full architecture breakdown.

· 7 min read
AI AgentClaudeOpenClawSolo DevBuildInPublic

We Gave Our 4 AI Lobsters the World's Smartest Brain, for Free

A 7-star GitHub project + 30 minutes of work = four AI agents upgraded from a 7B local model to Claude Opus 4.6. Cost: $0.

· 6 min read
自動化Solo BusinessBuildInPublicDevOpsAI Agent

The Solo Dev's Automation Arsenal: From Git Commit to Social Post, Zero Manual Effort

I spent a weekend wiring my development workflow into a fully automated social media pipeline: write code, commit, AI generates social copy, Discord + Threads publish simultaneously. Full architecture breakdown and security design included.

· 7 min read