This is the local launch checklist for DM-Code-Agent v2. It is a plan, not a record of completed external posting.
- Tag
v2.0.0after CI passes. - Release title:
DM-Code-Agent v2: auditable local code agent with Reflexion and critic review. - Attach links to:
README.mddocs/research-log/06-final-writeup.mdbench_reports/swebench_lite_baseline.mdbench_reports/economics.md
- State clearly that the SWE-bench Lite number is Tier-1 host verifier, not leaderboard comparable.
- Publish the final write-up from
06-final-writeup.md. - Keep the headline focused on the engineering artifact, not on unverified score improvements.
- Include the frozen baseline caveat near the first SWE-bench mention.
- X / Twitter: one thread with architecture, trace screenshot/GIF, and the frozen baseline note.
- Jike / Weibo: short Chinese summary with the research-log index.
- Reddit / Hacker News: post only after a clean install smoke test from a fresh clone.
awesome-llmawesome-ai-agentsawesome-mcp
Suggested description:
Local-first Python code agent with JSONL trace/replay, MCP tools, hidden-test benchmarks, SWE-bench Lite Tier-1 harness, and default-off Reflexion/Critic/Self-Consistency modules.
The intended 90-second demo flow:
- Fresh checkout and
pip install -e ".[dev]". dm-agent-bench --suite maintenance --list.- Show one deterministic eval or benchmark manifest.
- Show trace replay.
- Show
dm-agent-economicsregeneratingbench_reports/economics.md.
Do not present a new SWE-bench score unless a permitted real evaluation has been run.