Solo Business
34 articles about "Solo Business".
How an AI One-Person Company Actually Runs: The Split, the Lines, and the Mistakes
What it actually looks like when one person plus AI runs an AI product company. This post collects the practices we have published: strategy and veto on the upper layer, construction on the lower layer, only a human presses publish; own product and partner product split into two lines that never share a window; three things never outsourced: shipping, external promises, breaking spec; and why silent failure is the biggest enemy. Everything comes with verifiable records, including our own mistakes.
Seven Questions to Ask Before Buying an "AI One-Person Company" System
"AI employees" and "one-person company systems" are multiplying, priced anywhere from a few dollars to hundreds. Before paying, get seven answers: what exactly is delivered, what machine it runs on, whether the refund promise and the terms of service say the same thing, whether there is real operating evidence, whether the seller uses it, whether the items in the value stack can be bought separately, and what updates and community actually mean. At the end we run the same checklist against our own handbook, including the boxes we fail.
A Wrong Authoritative Memory Is More Dangerous Than No Memory: How to Check Your AI Agent's Knowledge Anchors
The authoritative memory we built to stop our AI agent hallucinating had itself recorded 25 as 12. How to check every number in agent memory against its source, and how to tell 'is this memory there' apart from 'is this memory right'.
How to Check Whether Your AI Agent Is Quietly Burning Money
We spent $17 in one week, and 98.8% of it paid for the agent rereading its own history. How to check your input-to-output token ratio, cap max_tokens, and verify that a config change actually took effect.
How to Audit an AI Agent's Exploration System: Catch LLM Errors Saved as Insights and Stop Temporary Failures from Burning Material
Half of our agent exploration system still mined insights, the other half was an outreach channel with an 82% bounce rate, and the living half was writing LLM error messages into its notes as analysis. How to verify that output is analysis rather than an error, and how to prevent irreversible loss.
How to Tell Whether an AI Agent's Learning Loop Is Actually Learning or Just Spinning
Our three learning loops ran every night and produced the exact same 5 sentences for three weeks. How to diff learning output, check whether the signal source has run dry, and verify that the loop actually closes.
LLM-as-Judge Pitfalls: How Our Judge Hallucinated and Marked a Correct Answer Wrong
Our flash-model judge marked a system that answered correctly as wrong, then pushed a CRITICAL alert. How to regression-test your LLM judge, and why questions with known answers should be scored with deterministic keyword checks.
Ran Once Is Not Integrated: How to Tell a Real MCP Integration from a One-Off Run
We created 9 live payment products through an MCP server, yet not a single MCP server was registered in our config. How to tell "ran it by hand once" apart from "actually integrated".
Why Our Spend Alert Never Fired: How to Set Monitoring Thresholds That Actually Trigger
Our $17 week slipped quietly under a $20 alert threshold. How to anchor thresholds to your real baseline and monitor cumulative spend instead of single peaks.
Multi-Agent Collaboration Audit: How to Check How Many Agents Are Really Running and What They Cost
We said six AI agents were collaborating. The real number was four, and we shut down the one agent that received every message to save money. How to count the agents that are actually alive and find the ones that cost more than they produce.
Full Monitoring, Nobody Reading It: How to Make Human-in-the-Loop Actually Work for AI Agents
Our guardrail, semantic, error and cost logs were all in place, and not one person was actually reading them. How to turn monitoring into an interface a person will really open, without flooding them with alerts.
Assume You Are Lying: How to Audit an Entire AI Agent Fleet
We implemented all 21 patterns from Agentic Design Patterns, then assumed every one of them was a lie and checked it chapter by chapter. Of the 14 chapters we claimed to have implemented, not one came out clean. This is the audit method, and the opening post of the whole series.
How to Check Whether Your Flagship AI Agent Still Runs: Ours Had Been Dead for 3 Months
Our proudest planning agent turned out, on audit, to have been dead for three months while we assumed it held a meeting every day. How to check a flagship feature's last real run time and whether any timer still triggers it.
Smart Scheduling or Just Rotation? How to Check Whether Your AI Agent Really Prioritises Work
We called it smart scheduling. Taken apart, it was index % 5 taking turns, and our scorer sat at enabled:false and never ran. How to tell dynamic prioritisation from static rotation in your own scheduler.
How to Find the Weakest Link in a Multi-Step AI Agent Chain Before One Broken Step Silently Ships Garbage
In our four-step posting chain, the parser in the self-review step broke and silently published a 14-word junk post. How to check the error handling at every step of your chain so the weakest link cannot drag the whole thing down.
RAG's Most Dangerous Failure Is Confidently Citing the Wrong Authority: How to Prevent It
Our RAG fixed hallucination, yet it went 7 weeks without a single real question and at one point cited a wrong authoritative number. How to check whether the authoritative definitions you inject are actually correct, and whether your RAG gets any real traffic.
Defined but Never Run: How to Find the Features in Your Codebase With Zero Callers
We defined a reasoning tier for diagnosis, and a grep across the whole codebase found zero callers. How to check whether a feature has real callers, and how to tell what exists in the docs apart from what actually runs.
When AI Self-Reflection Fabricates Numbers: How to Guard LLM Rewrites Against Fake Stats
Our self-reflection step rescued a bad draft, then invented a McKinsey statistic that does not exist in the same rewrite. How to constrain reflection to the numbers already in the draft, and how to diff the rewrite for new numbers that were made up.
How to Classify Errors Before You Retry: What to Back Off On and What to Stop Immediately
We backed off and retried on every error, and the 402 that most needed to stop at once caused the worst storm. How to route errors by type, put a cap on backoff, and add dedup to your alerts.
The Model Routing Cost Trap: Does Your LLM Fallback Chain Quietly Upgrade to the Most Expensive Model?
Our router picked the cheapest model, then climbed to the most expensive Claude every time it was rate limited, and we only noticed after $8. How to check which way your fallback chain moves on cost, and how to log which model each call actually used.
Agent Parallelism Audit: How to Check Whether Your Parallel Agents Are Real or Inflated
We said six brands ran in parallel. An audit counted the live gateways and found one. How to tell real parallelism from the empty shells a migration leaves behind by counting live processes instead of config files.
AI Agent Tool Integrations Rot Silently: How to Catch Tools That Return Nothing
A CLI flag that does not exist left our agent's research tools silently returning nothing for months, and nobody noticed. How to verify that tool output is non-empty and catch the silent failures that never crash but return zero results.
Inter-Agent Communication Audit: How to Check Which A2A Channels Are Still Alive
We claimed two inter-agent communication channels. One had been at zero connections for a long time and was still written in the config file. How to use ss and the relay queue to verify each channel, and how to clean out mechanisms that died in a migration.
Controlling Claude Code from Telegram: How I Ran It for 60 Days with Zero Downtime on Windows
Turn Telegram into the main console for Claude Code and send it commands from your phone wherever you are. A breakdown of the 4 defense layers and the real incidents behind 60 days of zero downtime on Windows, plus an MIT-licensed open-source toolkit.
OpenClaw Six-Month Retrospective: $0, 4 AI Agents, 35 Scheduled Jobs
OpenClaw has run for six months: 4 AI agents, 35 systemd timers, $0 a month. This post publishes all the real data: output volume, error rates, time saved, and the things AI never does well.
Why I Needed an AI Agent Team as a Solo Founder: Not Because It Was Cool, Because I Was Drowning
One person running 6 product lines, a 268-member community and 16 automated posts a day. This is not a boast. It is the real story of a founder close to breaking point who decided to let AI help, plus the decision framework I use to judge which work should go to AI.
Why We Only Write Articles, Never Make Videos: For People Who Ask AI Directly
Video tutorials have four fatal problems: can't find specific steps, wrong speed, outdated instantly, can't copy-paste. But the real issue isn't videos: the entire 'watch tutorials' model is obsolete.
AI Development for Beginners: From a Smartphone to Shipping Products (Complete Roadmap & Free Tools)
You can build AI products with zero coding experience. Start from your phone, know when to buy a computer, what specs you need, and a complete map of $0 free tools. All in one article.
AI Development Pitfall Diary: Mistakes I Made So You Don't Have To
Firebase, Vercel, API Keys, Git Push: feeling overwhelmed on your first AI project is normal. This article compiles the most painful mistakes I've made, saving you three months of detours.
The Art of AI Prompting: Why Your AI Conversations Never Give You What You Want
Ask the right question, and AI becomes your team. Ask the wrong question, and AI is just a parrot. This article teaches you how to go from 'I don't know how to ask' to 'one sentence that gets AI moving.'
From a Spreadsheet to a Brand: How My First Product Was Born
I just wanted to make a nice spreadsheet. Seven days later, I opened my own website on my phone. This is the complete story: no tutorial, just the real journey.
How to Run AI Agents on Gemini's Free Tier
4 AI agents ran up to 105 requests a day on Gemini's free tier with $0 in API fees, via systemd timers and one-shot prompts. Scripts are on GitHub.
The Solo Dev's Automation Arsenal: From Git Commit to Social Post, Zero Manual Effort
I spent a weekend wiring my development workflow into a fully automated social media pipeline: write code, commit, AI generates social copy, Discord + Threads publish simultaneously. Full architecture breakdown and security design included.
Why You Don't Need to Learn to Code: An AI Development Log from a Financial Advisor
In middle school, I bought a Visual Studio book thick enough to hammer tent stakes. Over a decade later, I built five products with AI. The difference isn't that I got smarter. The times changed.