Stress-tested MiroFish with 3 live backtests, writeup inside #537
Replies: 1 comment
|
Note 🤖 Automated maintainer response — Generated by the MiroFish triage agent and checked against Thanks for sharing these detailed community observations. The variation across domains and the resource figures are useful leads for further testing. To avoid community results being mistaken for an official benchmark, the maintainers have not reproduced these runs and do not validate the accuracy of their outputs. If possible, add durable, sanitized artifacts such as the seed inputs, exact MiroFish commit, model and provider, generated configuration, round count, logs, outputs, and hardware details so others can compare and reproduce the results. We will keep this Discussion open under Show and tell. |
Uh oh!
There was an error while loading. Please reload this page.
Hey, wanted to share a real writeup of testing MiroFish on three different scenarios from the same week.
The three backtests:
57 agents total across the three sims, 4,748 agent-generated posts and comments. All three completed cleanly on a 16GB VM.
The ontology generation, agent persona system, and document-to-simulation pipeline are all genuinely novel. The political and cultural sims built rich knowledge graphs with meaningful relationships. The economic sim came out much sparser (9 nodes, 1 edge) which was the most interesting thing I found. Looks like domain-specific tuning could make a big difference there.
Full writeup with specifics on what impressed me and where it broke: https://x.com/Metroxe/status/2044496303938539744?s=20
Happy to share any of the raw simulation data or dig into specific results if anyone's curious. Excited to see where this goes.
All reactions