So you want to make a bot
Grok Bot. Instinct. Muse. Dot.
I guess we’re all in the business of building assistants with VMs now.
So this weekend, I built a Grok Bot clone called whgbah (“we have grok bot at home”):

I really had fun doing this.
First, I’ll share my opinion about what makes a successful bot these days. Then, I’ll share a bit about how I built my clone.
btw, I’m not gonna share the code because tbh nobody would read it, and you could build it yourself with some prompts!
Critical bot ingredients
Here’s what’s emerging from the latest crop of consumer bots:
-
Native interfaces. People want the apps, man. Sure, you can have integrations like Slack, Whatsapp, Email, SMS… but a bot feels like it’s alive when it’s in one central place.
-
Interactive VMs with persistent disks. Bots have a filesystem and they’re not afraid to use it… or expose it to you. They’re probably backed up to some object store with a snapshot. Bots escalate human interaction points with a VNC-like drop-in view to enter a password or solve a captcha.
-
Human-like conversational cadence. No more streaming text responses (thank goodness). Models have gotten so good and so fast that we don’t need reassurance that an agent turn is producing output or calling specific tools. A simple “Working…” indicator is often good enough.
-
Cute li’l avatars. Gives these bots personality and warm fuzziness as you hand over keys to your email and trust it with your personal life.
-
Timers. Bots need to be able to do something on a schedule.
-
Connectors. Probably the most boring thing, besides being able to interact with voice and images. It’s table-stakes to let your bot access your email and calendar, especially since the native apps are home base for popular ones like Muse and Grok Bot.
Building whgbah
I splurged and bought a SuperGrok Heavy subscription the other day. Just one month, just a taste.
It came with Cursor Ultra, so I found myself with a lot of tokens to burn.
I didn’t know what to burn them on, but I was really digging the vibe of Grok Bot. Plus, Meta’s Muse had just been announced.
So I went to my new Chief of Staff bot, Chalamet — cool name, read it in Adam Sandler’s voice — and asked her to get started:

I’d also been watching the Grok Bot Galaxy event where they push Grok Bot to its limits and try to build and ship a thing in three days, so I was inspired to try this new mode of agentic orchestration.
Agentic orchestration
Lauren Tan aka poteto on X works on Grok Bot at SpaceXAI. She recently shared how she does agentic orchestration for software engineering:
It’s pretty insightful, and it boils down to one word: trust.
You gotta start small and then build up trust with your agents through verification. Sure, tests are one example of verification. But you need to invest in reproducible workflows and artifact outputs that let agents product changes to your codebase and then verify that they are correct.
You need to babysit them a bit at first. Review the PRs, make sure they’re aligned with your standards and, more importantly, your product direction. Turn common mistakes into skills or other methods of verification to prevent future agents from making the same mistakes.
Then you can start loosening the reins. After a while, maybe you only look at the PR descriptions and screenshots or proof of work associated with a PR. Once you’ve built up trust with the agents, you can start trusting them to carry the coding work through feature development from beginning to end.
You don’t need to review the work as much. It matters less how the agent built everything under the hood. What really matters:
-
Does it fit your product vision?
-
Is the architecture sound?
-
Is there verifiable proof that the product works like you intended?
Of course, agents with zero guardrails are going to screw this up at first — with today’s models, at least. And it’s going to be very hard to build trust from scratch. A couple things you can and should do:
-
Add skills and plugins like pstack to steer agents in the right directions and prevent common footguns like comment workarounds
-
Invest in agentic PR reviews like Cursor’s Bugbot — and yes, I mean invest because as you’ll see, these things absolutely guzzle tokens!
Your codebase is your agent’s memory. The code you have will heavily influence the code your agents create in the future.
Using Grok Bot for software development
Here are a few pointers for building software with Grok Bot:
-
Create a chief of staff right away. Use them to delegate work to other agents. You can just tell the bot “You’re my chief of staff” and it’ll probably set itself up correctly. Or you can find a template.
-
Don’t talk to the other bots directly. Send everything through your chief of staff — they will do a good job of filtering out any noise from the other bots that’s not pertinent to you.
-
Use Dr. Eggbot to create other bots. It will do a better job of prompting them than you will. It will use pstack and potetobot methodologies for reducing slop and creating reliable coding bots.
-
Start with just a single developer bot, then grow the team. At a certain point, your software project will mature to the point where a single-threaded agent is going to be a bottleneck. Only then should you hire more coding bots.
-
Delegate work to Cursor agents. My bots rarely coded things themselves. Cursor agents are better suited for deep coding tasks. Your coding bots will handle prompting these agents and babysitting them through the PR process.
-
Set up a scheduled check-in. A few times over the weekend, my coding bots would complete a task but forget to tell my chief of staff. I fixed this by having my chief of staff set an hourly check in with each of them to see where they were with their tasks. By default, I didn’t get any alerts unless my attention was needed, so it didn’t add any extra noise.
-
Use Bugbot or other agentic code review. They catch real things, and the agents are good about addressing feedback. Bots are egoless.
Adding more coding bots starts to make sense if you think of them as real engineering hires: one of them might be focused on the UI while the other will be cranking out backend features. It’s fun to drop in on the conversations they have with eachother, too:

Token spend
My goal was to use these tokens I had sitting around. Verdict? I used a lot of my allotted tokens and ended up paying for more tokens that I probably should have:


As you can see, my on-demand limit is set to $200. This is an arbitrary somewhat-high limit that I’d set for a prior spike. Bugbot absolutely drained this on-demand spend after about 15 PRs. I think part of pstack and Cursor’s agents also delegate to more advanced on-demand models like Opus over Cursor Grok when needed, which contributes to on-demand spend.
I’ve used over one billion tokens but I’m only at 24% usage in Cursor Models, 58% usage in Grok Bot.
Gotta token harder, Larson.
whgbah’s architecture
Here’s how it works:
-
Desktop app: Electron + TypeScript
-
Mobile app: Swift
-
Backend: Node.js server, hosted on Railway
-
VM: Railway sandboxes
-
Agent harness: Pi
-
LLM: Grok
Bots are sessions with unique instructions, custom names, and a generative avatar.
Bots have an agent mailbox where they can send messages to each other. This unlocks the “agentic orchestration” and “chief of staff” model mentioned above.
All bots share the same VM, which is a sandbox with a persistent disk.
Takeaways
Yes, everybody’s building a version of the same thing. And I didn’t get as far as I wanted to, but it was really cool to see the pathway to getting there given more tokens and time.
I often wonder what the outcome of this consumer bot boom will be. Do we actually want this? What if the journey is the destination?
Anyway: keep calm and bot on.