English · 中文
Four independent model slots (see Configuration for the env vars). This doc covers which models to put in which slot.
One benchmarked preset ships in .env.example. Copy it and set your API key.
Each slot controls a different quality axis:
| Slot | Controls | Key finding |
|---|---|---|
| Default | Persona richness, sim density | Mercury 2 — a diffusion LLM picked for speed; off the report's quality path |
| Smart | Report quality (#1 lever) | Gemini 3 Flash holds up on ReACT report loops with reasoning disabled |
| NER | Extraction reliability | Needs deterministic JSON — pick a model that doesn't silently emit CoT |
| Wonderwall | Cost (biggest consumer) | 850+ calls, 7M+ tokens. Verbosity matters more than $/M |
Mercury 2 personas + Gemini 3 Flash smart/NER + DeepSeek V4 Flash for the sim loop and web search. Reasoning is disabled on every slot (LLM_DISABLE_REASONING=true sends reasoning: {enabled: false} in extra_body), which is the difference between a ~45s scenario-suggest call and a ~3s one.
| Slot | Model | Notes |
|---|---|---|
| Default | inception/mercury-2:nitro |
Persona generation, sim config, memory compaction |
| Smart | google/gemini-3-flash-preview |
Report ReACT loop — only ~19 calls/run |
| NER | google/gemini-3-flash-preview |
Stable JSON with reasoning off |
| Wonderwall | deepseek/deepseek-v4-flash:nitro |
850+ agent-action calls/run; keep verbosity low |
Embeddings use openai/text-embedding-3-large (truncated to 768 dims via Matryoshka). Web enrichment uses deepseek/deepseek-v4-flash:online.
Latency note — every OpenRouter call goes through
LLMClient, which injectsreasoning: {enabled: false}intoextra_bodyby default. Turn it off withLLM_DISABLE_REASONING=falseonly if a specific slot benefits from chain-of-thought (rare for MiroShark's structured prompts).
The Wonderwall slot accepts a per-slot endpoint override so you can run a self-hosted or fine-tuned model alongside the OpenRouter-backed Default/Smart/NER slots:
WONDERWALL_BASE_URL=https://your-endpoint.example.com/v1
WONDERWALL_API_KEY=not-checked # any string for open endpoints
WONDERWALL_MODEL_NAME=your-model-idEither field can be omitted — a blank WONDERWALL_BASE_URL reuses LLM_BASE_URL, a blank WONDERWALL_API_KEY reuses LLM_API_KEY. Useful for routing the 850+ agent-action calls per run to a vLLM / Modal / Ollama-on-a-server deployment while keeping the report and graph-build slots on a hosted provider.
OrcaRouter is an OpenAI-compatible gateway with namespaced model IDs — one key covers every slot plus embeddings, and you can mix vendors per slot (e.g. Anthropic for the report slot, OpenAI for the high-volume sim loop). A ready-made block ships in .env.example and docs/INSTALL.md → Option A.4. All models below were verified live against the OrcaRouter API.
| Slot | Model | Notes |
|---|---|---|
| Default | openai/gpt-5.5 |
Persona generation, sim config, memory compaction |
| Smart | anthropic/claude-sonnet-5 |
Report ReACT loop; OrcaRouter accepts Anthropic cache_control blocks |
| NER | google/gemini-3.5-flash |
Deterministic JSON; reasoning is not force-disabled on OrcaRouter, so LLMClient strips any <think> blocks client-side |
| Wonderwall | openai/gpt-4o-mini |
850+ agent-action calls/run; keep verbosity low |
Embeddings use openai/text-embedding-3-large at https://api.orcarouter.ai (truncated to 768 dims via Matryoshka — OrcaRouter honors the dimensions param). OrcaRouter has no :online web-search variants — leave WEB_SEARCH_MODEL= blank and use MIROSHARK_SEARXNG_BASE_URL for web enrichment, or the default model is used as a fallback.
Latency note — the OpenRouter-only
reasoning: {enabled: false}injection does not apply to OrcaRouter base URLs, so keep the high-volume slots on the fast picks above (Wonderwall →openai/gpt-4o-mini, Default →openai/gpt-5.5).
Context override required. Ollama defaults to 4096 tokens, but MiroShark prompts need 10–30k. Create a custom Modelfile:
printf 'FROM qwen3:14b\nPARAMETER num_ctx 32768' > Modelfile ollama create mirosharkai -f Modelfile
| Model | VRAM | Speed | Notes |
|---|---|---|---|
qwen2.5:32b |
20GB+ | ~40 t/s | Default in .env.example — solid all-rounder |
qwen3:30b-a3b (MoE) |
18GB | ~110 t/s | Fastest — MoE activates only 3B params per token |
qwen3:14b |
12GB | ~60 t/s | Good balance for mid-range GPUs |
qwen3:8b |
8GB | ~42 t/s | Minimum viable; drop Wonderwall rounds if context is tight |
| Setup | Model |
|---|---|
| RTX 3090/4090 or M2 Pro 32GB+ | qwen2.5:32b |
| RTX 4080 / M2 Pro 16GB | qwen3:30b-a3b |
| RTX 4070 / M1 Pro | qwen3:14b |
| 8GB VRAM / laptop | qwen3:8b |
Embeddings locally: ollama pull nomic-embed-text — 768 dimensions, matches the Neo4j default.
Most users land here naturally: run local for the high-volume simulation rounds, route to Claude for reports.
LLM_MODEL_NAME=qwen2.5:32b
SMART_PROVIDER=claude-code
SMART_MODEL_NAME=claude-sonnet-4-20250514