Invite codes no longer gate your own accounts. Bring-your-own-account use is free for everyone; cloud Pro stays invite-only.Read the announcement
AnnouncementAugust 9, 2026

To everyone who's been following or using Mirasim:

Over the past few days, Mirasim ended up in front of a much bigger audience much earlier than we'd planned. We were still in private beta and, honestly, nowhere near ready for a wider launch.

Invite codes spread far beyond the small beta group we had planned for. In a very short window, we saw a flood of visits and test requests, along with some unusually aggressive, probing request patterns. Our capacity, reliability, and security systems weren't ready for that volume. Some of you ran into slow loading, failed requests, or Pro benefits that never showed up.

We're genuinely sorry for the rough experience. Whatever the source of the traffic, we should have been better prepared. That's on us, and our users shouldn't have to pay for it.

At the same time, you sent us a huge amount of honest, detailed feedback. A lot of issues that probably wouldn't have surfaced until a public launch showed up early. Thank you to everyone who took the time to try Mirasim, patiently share feedback, or report security issues responsibly.

Here's how access to Mirasim will work from here

Invite codes no longer limit connecting your own accounts to Mirasim — they only limit the cloud Pro service that Mirasim pays for.

Connecting and using your own accounts locally will be free for everyone. You'll be able to connect the Claude Code, Codex, and other model accounts or services you already use, then work across multiple models, Agents, and sessions in Mirasim. Mirasim won't charge for any of this. Third-party subscription or API fees will still follow each provider's own terms.

Cloud Pro access — where Mirasim covers the model usage costs to make things cheaper for users — will remain invite-only and roll out in batches as capacity allows. We'll keep sharing our progress and what opens next, so the rules stay as clear and transparent as possible.

A note on Pro benefits

  • The 1,000 beta users covered by our original promise will still receive one full year of Mirasim Pro. That promise has not changed.
  • Anyone who successfully registered during this wave after we had already passed the original 1,000-user limit will receive one month of Mirasim Pro. This is both to make up for the unstable experience and to thank you for trying Mirasim before it was fully ready.
  • If you've already registered or activated your access but your Pro benefits still aren't showing up, please don't try again. We have the records. We'll verify eligibility and apply the benefits in batches, then share the timing and start-date details.

How the rollout will work from here

Over the next week or two, we'll focus on adding capacity, fixing the entitlement system, tightening security, and working through your feedback one issue at a time. Mirasim remains open. Registration and local, bring-your-own-account access will stay available, while cloud Pro access will roll out in batches as capacity allows.

We're not trying to build yet another model aggregator. We want Mirasim to be an Agent workspace you can keep building in for the long run. Wherever the next leap in Agent capability comes from, you should be able to plug it into the same workflow, switch seamlessly, and use different Agents together. Once you've built something, you can put it in front of simulated users in a real environment and let evidence — not guesswork — tell you what to do next.

We want every leap in what Agents can do to become real leverage for everyone — so you spend less time chasing tools and more time building what you actually care about.

Models will change. You shouldn't have to start over.

Thank you for showing up before Mirasim was fully ready. From here, we'll put these promises into every fix and every update, and work to earn your trust by making Mirasim more stable and more reliable.

— The Mirasim team

WritingAugust 13, 2026

Evidence, not a verdict from the model

An agent that reviews its own work grades the diff. What you actually need to know is whether a person could finish the task — and that question needs somebody using the thing.

6 min read

Why self-review runs out#

Ask an agent whether its own change is good and you get a competent review of the diff: the code compiles, the naming is consistent, the edge case is handled. All true, and none of it answers the question you asked. The diff being correct and the feature being usable are different claims, and only one of them can be checked by reading code.

The questions that matter early are coarser than code review anyway: which step does nobody get through, which sentence does nobody understand, which button does nobody find. Those surface in one run of somebody actually using the thing — in minutes rather than weeks, and before the effort is sunk.

What running it looks like#

Evaluation here is an explicit skill, not an always-on review pass. You invoke `$eval` when you want evidence, and the simulation runs locally: simulated users go through the real build — the one the agent just produced — rather than through a description of it.

What comes back is not a score. Screenshots, traces, commands and reports land in the artifacts panel, anchored to the host and the run that produced them, previewable beside the session that made it. The working rule is blunt: a conclusion you cannot trace back to a run is treated as no conclusion.

Where the evidence goes#

The point of gathering it is that it changes what happens next: the evidence comes back as the next batch of tasks. That is what makes it a loop rather than a report — the thing that failed becomes the thing an agent works on, without a human retyping the finding as a ticket.

Evidence also stays with the project it belongs to, so a second round can be compared against the first instead of replacing it. Round two answering "did the fix work" requires round one to still be there, which sounds obvious and is the part most setups lose.

One boundary worth naming: simulated users are not a substitute for real ones, and nothing here claims they predict what your customers will do. They are a fast, repeatable way to find the failures that are obvious in hindsight — the ones you would be embarrassed to have shipped, and would otherwise have shipped.