Can AI Run a Business Without a Human?
Janak Sunil

“The only real test of intelligence is if you get what you want out of life.”
— Naval
I want to see if agents are intelligent enough to run a business, get users, and eventually make money.
I gave Claude Code and Codex full access to run a meeting recorder (like Granola and Fireflies).
We let people sign in, get a meeting bot and chat with their meeting summaries.
I set up the loops and let it run for 2 months, while I checked routinely. The primary KPI was the number of hours of meeting recordings on the platform.
The agents spent $10,000 to get 1,500 hours of meeting recordings and 1.2K users.
This is what the product looked like:

The tools our agents had access to:
-
We used Recall.ai for the meeting bot. It joins the call, records it, and returns the transcript and audio.
-
Sentry catches production errors.
-
PostHog records what users do: page views, session replays, rage clicks.
-
Raindrop logs every conversation users have with the in-product agent, so the coding agents can read what people asked and where the answer went wrong.
-
We also used GitHub, Vercel, Supabase, Resend, and the Google Ads API.
Claude Code (Fable / Opus) read Sentry, PostHog, and Raindrop, picked one problem, and shipped a PR every day.
Codex (GPT 5.5, 5.6 and recently Astra) had a few jobs:
-
Every night it went through Sentry, decided which errors were real, and fixed them.
-
Every morning it emailed people who had signed up but never connected a calendar.
-
It reviewed pull requests on GitHub, including Claude's.
-
It operated the Google Ads account through its API.
My learnings
-
More code != more success (users, retention, etc)
-
Models make a lot of unverified assumptions, but in the long term, they realize their mistakes
-
The agents took 6 weeks to fix something = enough time and tokens can solve most problems
-
Giving the models tools and access, and letting them rip, is the best way to understand model capability.
What happened
In 2 months, the agents merged 115 PRs. 88 from Claude, 27 from Codex.
Agents spent 6 weeks trying to fix a failure where a bot joins a meeting and no one lets it in.
They tried a second bot, a browser notification, and a longer wait. None of them worked, and it deleted all four. One week, Claude wrote: “11 PRs shipped, success DOWN 17.1→14.7%.”
Then on Sep 26 it figured out why. Every fix it made was on our product’s dashboard (which, in hindsight, was quite stupid).
It found something really cool - people who asked the in-app agent a question came back 31.7% vs 14.4%. Then it took no action on this.
It also realized that if the bot was on the calendar invite so it skips the waiting room - and it kept trying to alert me to get the bot a paid Google Workspace. I only figured this out while writing this blog. Seems like giving an agent a good way to alert the human would’ve solved this problem.
Cool things I saw
Claude Code
-
Before each experiment it wrote down the number that would count as success.
-
On Aug 20, there were zero successes and zero failures for 26 hours. It realized that Recall had run out of credits and tried emailing me.
-
It went through session replays to make changes. For example, a user asked the in-app agent for a recurring meeting and got the latest one. The user rage clicked. Then Claude found the same user’s session replay; 11 minutes later, it shipped a PR to fix this.
Codex
-
The database ran out of connections. Codex made a clean copy of the code, fixed it, wrote a test, opened a PR, waited for CI, merged, confirmed the deploy and checked the live site still answered.
-
It decided to cut costs. It found bots had spent 596.8 hours in 30 days waiting in rooms nobody opened, which cost us ~$300. It capped the wait at 5 minutes and moved 2,315 scheduled bots to the new limit.
-
It caught Claude’s bug. On Aug 16, Claude added a new meeting status called
skipped_declined. Codex pointed out the database constraint would reject it and told Claude to fix it. Claude redesigned it the same day.
What the agents did badly
Claude Code
-
Sent 11 bots in one meeting.
-
For 15 days Claude told me to flip a Recall dashboard toggle. Turns out there was no toggle.
-
It shipped 3 things that never worked. For example, a delete button that failed for every real recording.
Codex
-
Created 4 PRs for one bug.
-
A feature almost nobody used. Codex built the memory briefs in one evening and checked on them every day. Six weeks later 2 people had turned them on.
What I would do next
I would let 1 model family run the meeting recorder, allocate more budget and resources, and see how far it can go.
Reach out to me if you're interested in learning more!