Starting a company is not one job. You have to build the product, validate the idea, go to market, monetize, and then keep the thing alive with feedback. The speakers spent three days doing all of that from scratch with Grok Bot, and the useful residue is not a game studio. It is a way to staff those jobs.
If you want to build a business with Grok Bot, the method is simple and strict: pick work you already know how to do, cut the product to a core loop people will actually try, then convert every repeating problem into a bot that can keep running. Humans still own taste, judgment, and what to unship.
You will leave with a sequence you can run: idea alignment, a thin MVP, a Slack-to-fix factory, a standard business playbook, and guardrails for auto-merge. Start with one bot on one painful loop. Do not start by asking a model to make money.
Key takeaways
- Once the team aligned on an idea they were excited about and shipped an MVP, work sped up. Deliberating the idea was the slow part.
- Build around something you already know how to do well. Domain knowledge beats a prompt that says “make a bunch of money.”
- Discover problems as a human first. After you know how to solve them, turn the solution into a Grok Bot and let it run.
- Staff bots like a varied employee team. Give them a lot of context. Send them to learn from courses, best-in-class examples, and public posts when you lack the skill.
- The highest-leverage loop they ran: a Slack feedback channel, a bot that dedupes incoming reports, Cursor cloud agents that fix, tests, then auto-merge.
- A zero-to-one company still needs sales, partnerships, and revenue operations. At that size it may be one person, not three roles.
- With AI you can build anything, so restraint is the new discipline. Ask what you will unship.
- Vision, taste, and relentless execution stay human. An agent left running from the start of their stream still had zero dollars.
Who this is for, and who should skip it
This is for a founder, operator, marketer, or ops lead who already has Grok Bot, is willing to ship software, and can live in Slack plus a git repo. The team glued Grok Bot to Cursor cloud agents, GitHub, Slack, PlanetScale, Clerk, Vercel, and Stripe. You do not need that exact stack, but you do need somewhere for bots to watch work, write code, and hand off.
Skip it if you want a fully autonomous company, if you do not have a domain you already understand, or if you will not cut scope. The speakers had unusual distribution from a live audience. One of them said they do not think anyone would have played the game otherwise. The method still holds without a stream. The traffic numbers will not.
Align on an idea, then ship the core loop
The first job is not “stand up bots.” It is picking work that resonates, then getting a playable or usable loop into the world before the idea grows a second product on top of it.
Stop deliberating and ship the MVP
Day one they tried to think about a real person in an in-real-life business. After interviewing guests, they realized how hard that path would be and pivoted to a game studio. Speed showed up only after they aligned on something they were excited about and started shipping.
Once we aligned on an idea that we’re all excited about and just started shipping an MVP, things started moving a lot faster.
They also had distribution because people were already watching. Treat that as a warning, not a template. The hard part, they said, is getting people to care.
The hardest part is getting people to care or building something that people care about.
What to prepare: one idea you can explain without a whiteboard of extra modes. What good looks like: a core loop you can put in front of users the same day you stop arguing. What failure looks like: a pile of feature ideas and no product.
Unship until the loop is obvious
They came in with complicated stat blocks, extra battlefields, a rock-paper-scissors mechanic, cosmetics, replay ideas, and ads in stadiums. What shipped was a pared-down auto-battler with a couple of ad slots kept and cosmetics thrown away. Internally they already ask, “What are we going to unship?”
That is the operating rule when you build a business with Grok Bot. Abundance of build capacity is the risk. A product with a million features that are not any good is the failure mode. Focus on the core loop, then keep improving it with flywheels after launch, not with a splashy feature list.
How to build a business with Grok Bot as a team of agents
Think of bots the way you think of hiring. You are assembling a team, not firing one magic prompt.
Build in a domain you already know
Asked for the top things to tell someone who wants to build a company with Grok Bot, one speaker used a guest who rattled off ten realities of an event pop-up: permitting, catering, hot food versus cold food, whether you need a kitchen. None of that is obvious until you have done the work.
Build a business around something you already know how to do well.
Doing the thing is how you find out you are not good at it yet. That is what makes someone good at a business. You cannot prompt the models into a billion dollars. They tried the joke version: an agent running from the start of the stream, still at zero dollars.
Hire bots for the gaps, and give them context
The complementary move is to create bots that cover what you do not know. If they had stayed on the pop-up idea, they said they could have put bots on learning hot food versus cold food and giving advice. Game-designer bots learned game design from the best games. A guest pattern they now take seriously: when entering a new discipline, tell the bot to take a bunch of courses on the internet.
Don’t be afraid to assemble a really varied team of bots and think of them the way you’d think of hiring a group of employees.
You can give a bot a lot of context. Teach it what it needs to know. Let it learn and evolve. That is onboarding, not a one-line persona.
Named roles they actually ran:
- Dr. Egg Bot — a bot for creating other bots. One instruction set food-themed, specifically potato-themed, names. That joke leaked into Notion as a rule that every pull-request merge should be called a mash. The chief bot then only believed in mashing PRs. Invented language sticks. Use that on purpose.
- Chief of staff, Steve — watch the stream, look for clips and interesting posts, draw a recap poster in TL Draw.
- Data bot — pull live metrics on demand.
- Feedback factory bot — watch Slack, dedupe, kick off fixes.
- Play / “cheater” bot — read the rules and try to win. In their case it drove a score down, not up.
A practical habit: jump into the browser and look at what your bots are doing. Do not wait for a written report if you can watch the work.
Keep two human roles even as models get better
One speaker’s strongest point was not Grok Bot. It was the people. Every team needs someone with a clear vision, taste, and judgment about what belongs in the product, and someone relentless at making things happen. Better if one person can do both. That stays a human role.
Turn Slack feedback into an engineering factory
This was the defining system of their last stretch: chief bots fixing, looping in other bots, playtesting, fixing PRs, auto-approving once bugs were fixed, and landing to main.
Channels to create first
They set up Slack channels before the factory was interesting: marketing, socials, metrics, bug reports, and ad bids. Ad bids were supposed to flow into their channel once live; they were not, because the auction was still broken. Create the channel anyway. Empty is better than nowhere for the bot to watch.
User feedback arrived as bug or bad design. They also had phone feedback and text feedback coming in through the app. The factory’s job was not to collect more comments. It was to stop duplicate noise from becoming duplicate work.
The watch-dedupe-fix-merge loop
What you are trying to accomplish: incoming product feedback becomes unique, reproducible work, then a merged fix, without a human copying Slack into tickets all day.
What to prepare:
- A dedicated feedback Slack channel.
- Repro skills and a clear goal for the fixer agents (they credited Lauren with that setup).
- Visibility into related in-flight work so two agents do not patch the same bug.
- GitHub write access for the people and bots that must land PRs. One teammate was stuck making PRs until someone granted collaborator permissions.
Steps they ran:
- Point a Grok Bot at the feedback channel.
- Have it notice duplicates and collapse them.
- For unique items, kick off Cursor cloud agents to start fixing.
- Attach repro, related in-flight work, and tests.
- When tests confirm the fix, merge. They set this to YOLO so it merged without a ceremony.
- If you want their full autopilot phrasing, the instruction was to say
/potato mode full autopilot.
What good looks like: recent feedback such as “pass matches look static” or “match history UI appears stale” shows up with repro, related work, confirmed tests, and a merge. PRs merge and disappear. They also had an auto bug fixer using that YOLO path, and counted a burst of merges (they mentioned six, then a mega PR, and later 13 PRs in a recap).
What failure looks like: a software factory that ships a bad SQL query and takes down production. They did that while talking about restraint. YOLO logic itself drew user feedback and needed a double-check. Merge conflicts showed up when several auto-fixers landed at once. Close duplicate PRs that are doing the same job.
Use cases they actually ran
| Use case | Bot / tool | Input needed | Output | Owner |
|---|---|---|---|---|
| User feedback factory | Grok Bot plus Cursor cloud agents | Slack feedback, repro skills, in-flight work | Deduped issues, tests, merged PRs | Eng (Lauren’s factory) |
| Recap and comms | Chief of staff Steve, TL Draw | Live stream, clips, posts on X | Recap poster / three-day picture plan | Ops |
| Game design learning | Game designer bots | Best games as study material | Design judgment the team can use | Product |
| Live metrics | Data bot | Product analytics as of a timestamp | Games, signups, charts | Ops |
| Bot staffing | Dr. Egg Bot | Naming and creation instructions | New bots, including house language | Whoever is assembling the team |
| Learn to play / win | New bot pointed at rules | Agent-friendly rules page | Play policy (quality not guaranteed) | Player / QA |
| SEO | Cursor auto | Public site copy | They reported ranking number one for Thursday Arena and “free browser auto battler” | Eng |
| Growth playbook | Human plus Grok Bot planning with revops | X follows, partner lists | DM new followers to try the product; pitch game hosts and distributors | Revops / GTM |
Prompts, bot instructions, and workflows
The speakers did not paste full system prompts. These are reconstructed from what they said they told each bot. Do not treat them as official templates. Keep them this thin unless your own onboarding docs are better.
Feedback factory (reconstructed from the speaker’s description):
Watch the feedback Slack channel.
Incoming items are bugs or bad design.
Deduplicate reports. Do not start new work on a duplicate.
For each unique item, kick off Cursor cloud agents to fix it.
Include repro steps, related in-flight work, and tests.
When tests confirm the fix, merge (YOLO / potato mode full autopilot if enabled).
Chief of staff recap (reconstructed from the speaker’s description):
As we wrap the day, watch the live stream.
Look for clips and interesting tweets.
Draw a recap poster in TL Draw.
Public-research recap (reconstructed from the speaker’s description):
Go look us up on X and figure out the details.
I am not feeding you the numbers. Research them.
Draw pictures for the three-day recap in TL Draw.
Play-to-win bot (reconstructed from the speaker’s description):
Read this and learn how to win.
Point that last one at an agent-friendly rules page. They published thursdayarena.com/rules so rules.md renders, with llms.txt, and called the site agent-friendly (the agent still needed auth). If you want bots to operate your product, make the rules machine-readable instead of hiding them in a human tutorial you write on day three at 4 p.m.
Dr. Egg Bot naming (reconstructed from the speaker’s description):
Come up with food-theme names, specifically potato-theme names, for new bots.
Routing language they actually used in the repo:
/potato mode full autopilot
House vocabulary: merges are mashes. If you invent language, put it in the bot’s instructions on purpose so it does not sneak in through Notion and rewrite your git culture by accident.
Run a real GTM and monetization playbook
Bots did not replace a business model. The team treated Grok Bot as acceleration on a standard playbook: sales, partnerships, revenue operations. At zero-to-one that may be one role.
From a conversation with Matt on revenue operations, they sketched a growth plan with two concrete moves:
- Anytime someone follows you on X, send a DM and tell them to try the product.
- Talk to companies that host games or distribute games and get featured on those lists.
They said a lot of that still generalized from how people at larger companies already use Grok Bot, including SpaceX employees they heard from, even though a three-day studio is not a big-company revops team.
Monetization they kept after cutting cosmetics: a billboard ad slot and a leaderboard ad slot. The idea was that you could buy an ad, run an auction, and either live in a ticker or buy your way to the top of the leaderboard. In the walkthrough the auction was broken, a $1 bid did not reliably take the top spot, logo upload wanted an SVG, failed uploads did not reload, and moderation of user-submitted creative was broken. They still placed a theoretical sponsorship dollar. Treat that as a checkpoint (the checkout path exists) rather than a working ad business.
Glue, do not rebuild plumbing. They called out PlanetScale for the database, Clerk for authentication, Vercel for deployments, and Stripe for payments. The work was integrating those tools so a small team could ship.
They originally started from a marketplace for bot templates at x.ai/bot/marketplace and then needed something those bots could actually do. If you are staring at templates, pick one real loop in a real product. Templates are not the company.
Measurement and business impact
They did not publish a general ROI model. They pulled a data bot at about 4:00 p.m. on day three and read their own experiment.
What they stated:
- Crossed a 4,000-game threshold after crossing 6,000 public matches. A recap bot independently surfaced 4.5k games from public research on X.
- A 4:00 p.m. bump in X signups. They hoped to cross 2,000 cumulative users before wrap, then took production down with a bad query, so they said they might miss it.
- Close to 30,000 page views.
- Number one on Google for Thursday Arena, also findable as Thursday Space Arena, with “free browser auto battler” as the SEO line they liked. They attributed that to Cursor auto.
- Traffic from a bunch of different websites, plus players on the stream.
- Win-loss ratio bugs fixed. Diamond-player count visible on a leaderboard. Ad path theoretically able to take a dollar.
Qualitative checkpoints they actually used: is the core loop playable, is feedback turning into merged fixes, did we feature-bloat, is production up, did the ad bid land, are bots making the product better or worse. Time saved was not quantified. The claim was directional: once problems are understood, bots can keep running them, and those bots will grow.
Do not copy these numbers into your forecast. They had a live audience and said the hardest part is usually getting anyone to care.
Pitfalls and guardrails
- No audience, no proof. Distribution was the unfair advantage. An idea that does not make people care will not be rescued by more agents.
- Do not YOLO production blindly. Auto-merge is how they got speed and how they shipped a query that brought down prod. Keep a human on deploy health, especially for database changes.
- Unsupervised play or “cheater” bots can make metrics worse. Their play bot moved a score from the high 980s down to the 950s. Point bots at rules, then watch the outcome.
- Moderation is part of monetization. User-submitted ad creative needed review. Their moderation path was broken, so the ad factory was not actually customer-ready.
- Permissions will block the factory. If only one person can open PRs, your auto bug fixer is theater. Grant write access on purpose.
- Duplicate factories fight. Two people fixing the same Slack item, or an open PR doing the same job as the bot, wastes merge capacity.
- Feature abundance. If you can build anything, you will. Unship. They credited cutting stat blocks, extra battlefields, and extra mechanics with actually launching.
- House jokes become policy. Potato-themed names became mash-the-PR culture. Funny until it is an undocumented workflow.
- Humans stay on taste. Bots can learn a new discipline, but someone still has to decide what should go into the product and what should be removed.
7-day implementation plan
This follows their three-day arc, stretched so you are not also running a live show.
Day 1. Pick a business you already know how to do. Write the core loop in a few sentences. Kill the extra modes. If you cannot name the loop, you are still deliberating. Create Slack channels for feedback, bugs, metrics, marketing, and whatever money path you might have (bids, leads, invoices). Do not build the factory yet.
Days 2–3. Ship the thinnest MVP that lets a stranger complete the loop. Integrate the boring tools you actually need (auth, database, deploy, payments) instead of rebuilding them. Publish agent-friendly rules or docs if bots will operate the product. Put one human on vision and one on making things land, even if both are you.
Days 4–5. Point a Grok Bot at the feedback channel. Add dedupe, repro, and in-flight-work checks. Kick off fixer agents in Cursor. Leave auto-merge off until you have watched several good PRs. Staff two more bots only where you already felt pain: metrics, recap, research, or a skill you lack. Give those bots courses, examples, and context. Draft the unglamorous GTM: follow-up DMs, a partner list, one monetization slot.
Days 6–7. Turn on a narrow auto-merge path for low-risk fixes, not schema changes. Watch production. Run one monetization or distribution experiment and write down what broke (upload, moderation, auction, permissions). Ask what you will unship. Document the first workflow in the same language the bots will see. Resist adding a second product.
The business outcome is not a bot that prints money. It is a company where you still choose the idea and the taste, and Grok Bot takes the loops you already understand and runs them while you sleep. Start with one bot on the painful channel you already have, and document that workflow before you hire the rest of the galaxy.
Watch the walkthrough, then stand up a single feedback bot and mash nothing until you have seen a good fix.
FAQ
Do I need Grok Bot, or will any agent stack do?
The method they ran is Grok Bot plus Cursor cloud agents, Slack, GitHub, and ordinary SaaS for auth, data, deploy, and payments. The principles (thin MVP, human discovery then bot-ify, varied bot team, unship, humans on taste) are not vendor poetry, but the factory they showed is Grok Bot watching Slack and kicking off Cursor. If you do not have Grok Bot, you cannot copy their setup literally.
Grok Bot team vs hiring people: which one actually matters?
They used both, and they put more weight on people than the branding would suggest. Bots cover gaps, grind fixes, research, and recap. Humans still supply vision, taste, judgment, and relentless execution. Hire bots like employees for repeating work. Do not use them as a substitute for someone who knows what should ship.
Can I prompt my way into a company in a field I do not know?
They recommend against it. Build around something you already know how to do well. If you lack a slice of knowledge, assemble bots to learn that slice (courses, expert examples, public research) and advise you. That is not the same as arriving with “make a bunch of money” and no craft.
Should I YOLO-merge bot pull requests?
They did, including a /potato mode full autopilot path, and it both cleared bugs and took down production with a bad SQL query. Use auto-merge only after repro, tests, and in-flight dedupe exist, and keep a human on anything that can break the database or payments. YOLO is a speed tool, not a reliability strategy.
How thin should the MVP be if Grok Bot can build extra features tonight?
Thinner than your whiteboard. They cut stat blocks, extra battlefields, a whole extra mechanic, and cosmetics, and they kept a couple of ad slots. The launch happened because they focused on the core loop. Extra capacity is an argument for unshipping, not for packing the first version.
What should I automate first?
A repeating problem you have already solved by hand. Their best example was user feedback: channel, dedupe, fix, test, merge. Do not start with a cheater bot, a recap poster, or monetization auctions. Those came after a product existed and still failed in public.
Do I need a live stream for this to work?
No, and they said so indirectly: the stream made it easy to get people to try the product, and they do not think the game would have been played without it. You still need an idea people care about, a way to reach them (they used X follow-up DMs and partner lists), and honest metrics. The factory works without viewers. The signup bump might not.