GLM-5: From Vibe Coding to Agentic Engineering
-
Updated
Jul 15, 2026
GLM-5: From Vibe Coding to Agentic Engineering
The long-horizon computer-use harness. Run AI agents across desktop apps and the CLI for extended periods while preserving task state and making reliable progress on complex workflows. Features fresh-context execution, durable verified state, independent auditing, recoverable progress, and native Claude Code / Codex / OpenClaw integration.
Programmatic memory for long-horizon LLM agents: the harness appends everything to one log, and the agent searches it with code. 97.4% on ARC-AGI-3 (arXiv:2607.20064)
Fixing GRPO training collapse in long-horizon multi-tool agents. A lightweight PRM-Lite + LATA joint approach achieves +37% over vanilla GRPO on τ-bench airline (50-task, multi-turn).
Simple Long Horizon Agent - A simple yet effective AI agent for learning, experimentation, and long horizon work.
SWE-Marathon: an ultra long-horizon SWE benchmark
Code that accompanies the paper release for "LLMs Corrupt Your Documents When You Delegate"
🧠 Awesome Memory-VLA: A curated list of Visual-Language-Action models with memory
MerchantBench is a 365-day, order-level benchmark for evaluating the long-term coherence of LLM agents in seller-side e-commerce operations.
Long-horizon agent skill for Claude Code / Cursor / Codex / Grok Build — multi-task ledger loop, host-portable, clean-context supervisor, verified gates. Markdown library (loop-graph), not a framework.
Official Repository for our paper: PlanBench-XL: Evaluating Long-Horizon Planning of LLM Tool-Use Agents in Large-Scale Tool Ecosystems
Code for Scalable Offline Model-Based RL with Action chunking
Code for Tackling Long-Horizon Tasks with Model-based Offline Reinforcement Learning
The simplest way to build long-horizon environments
VLM-RL Hierarchical Loco-Manupilation For Long-Horizon Tasks With G1 robot in Isaac Lab/Sim
Local-first, eval-first memory for long-horizon AI agents — no LLM at ingest. Python SDK + MCP server with source-traceable recall, belief revision, selective forgetting, and reproducible benchmarks.
LLM Agent Harnesses: A Survey of Human-Machine Interaction, Self-Evolution, and Long-Horizon Execution
A family of long-horizon software-engineering environments for OpenEnv, adapted from https://github.com/Proximal-Labs/frontier-swe
MobileMem: On-Device Memory for Continually Evolving Agents
Long-horizon agent execution harness — reliable autonomous runs for Claude Code, Codex, OpenHands, and custom agents. Goal graphs, spin detection, HITL gates, fork/merge, 8 strategies, 6 validators.
Add a description, image, and links to the long-horizon topic page so that developers can more easily learn about it.
To associate your repository with the long-horizon topic, visit your repo's landing page and select "manage topics."