Can AI Run a Business Without a Human?

Janak Sunil

Two illustrated agents operating a machine, one adjusting its gears and the other collecting coins.

“The only real test of intelligence is if you get what you want out of life.”

— Naval

I want to see if agents are intelligent enough to run a business, get users, and eventually make money.

I gave Claude Code and Codex full access to run a meeting recorder (like Granola and Fireflies).

We let people sign in, get a meeting bot and chat with their meeting summaries.

I set up the loops and let it run for 2 months, while I checked routinely. The primary KPI was the number of hours of meeting recordings on the platform.

The agents spent $10,000 to get 1,500 hours of meeting recordings and 1.2K users.

This is what the product looked like:

Meeting Note Taker dashboard showing upcoming meetings, recordings, and an in-app chat.

The tools our agents had access to:

  • We used Recall.ai for the meeting bot. It joins the call, records it, and returns the transcript and audio.

  • Sentry catches production errors.

  • PostHog records what users do: page views, session replays, rage clicks.

  • Raindrop logs every conversation users have with the in-product agent, so the coding agents can read what people asked and where the answer went wrong.

  • We also used GitHub, Vercel, Supabase, Resend, and the Google Ads API.

Claude Code (Fable / Opus) read Sentry, PostHog, and Raindrop, picked one problem, and shipped a PR every day.

Codex (GPT 5.5, 5.6 and recently Astra) had a few jobs:

  • Every night it went through Sentry, decided which errors were real, and fixed them.

  • Every morning it emailed people who had signed up but never connected a calendar.

  • It reviewed pull requests on GitHub, including Claude's.

  • It operated the Google Ads account through its API.

My learnings

  1. More code != more success (users, retention, etc)

  2. Models make a lot of unverified assumptions, but in the long term, they realize their mistakes

  3. The agents took 6 weeks to fix something = enough time and tokens can solve most problems

  4. Giving the models tools and access, and letting them rip, is the best way to understand model capability.

What happened

In 2 months, the agents merged 115 PRs. 88 from Claude, 27 from Codex.

Agents spent 6 weeks trying to fix a failure where a bot joins a meeting and no one lets it in.

They tried a second bot, a browser notification, and a longer wait. None of them worked, and it deleted all four. One week, Claude wrote: “11 PRs shipped, success DOWN 17.1→14.7%.”

Then on Sep 26 it figured out why. Every fix it made was on our product’s dashboard (which, in hindsight, was quite stupid).

It found something really cool - people who asked the in-app agent a question came back 31.7% vs 14.4%. Then it took no action on this.

It also realized that if the bot was on the calendar invite so it skips the waiting room - and it kept trying to alert me to get the bot a paid Google Workspace. I only figured this out while writing this blog. Seems like giving an agent a good way to alert the human would’ve solved this problem.

Cool things I saw

Claude Code

  1. Before each experiment it wrote down the number that would count as success.

  2. On Aug 20, there were zero successes and zero failures for 26 hours. It realized that Recall had run out of credits and tried emailing me.

  3. It went through session replays to make changes. For example, a user asked the in-app agent for a recurring meeting and got the latest one. The user rage clicked. Then Claude found the same user’s session replay; 11 minutes later, it shipped a PR to fix this.

Codex

  1. The database ran out of connections. Codex made a clean copy of the code, fixed it, wrote a test, opened a PR, waited for CI, merged, confirmed the deploy and checked the live site still answered.

  2. It decided to cut costs. It found bots had spent 596.8 hours in 30 days waiting in rooms nobody opened, which cost us ~$300. It capped the wait at 5 minutes and moved 2,315 scheduled bots to the new limit.

  3. It caught Claude’s bug. On Aug 16, Claude added a new meeting status called skipped_declined. Codex pointed out the database constraint would reject it and told Claude to fix it. Claude redesigned it the same day.

What the agents did badly

Claude Code

  1. Sent 11 bots in one meeting.

  2. For 15 days Claude told me to flip a Recall dashboard toggle. Turns out there was no toggle.

  3. It shipped 3 things that never worked. For example, a delete button that failed for every real recording.

Codex

  1. Created 4 PRs for one bug.

  2. A feature almost nobody used. Codex built the memory briefs in one evening and checked on them every day. Six weeks later 2 people had turned them on.

What I would do next

I would let 1 model family run the meeting recorder, allocate more budget and resources, and see how far it can go.

Reach out to me if you're interested in learning more!