What advances agents next: five research directions we are watching
Published 8 July 2026 · Editorial update 6 September 2026 · 2 min read · Enternovate

One: long-horizon reliability. Longer workflows create more opportunities for error. Checkpoints and intermediate verification can make failures easier to detect and recover from. Measure completion on real tasks rather than inferring it from conversational quality.
Two: useful memory. Persistent information can help an agent, but stale or untrusted memory can mislead it. Nyarhi provides a knowledge-graph layer. Retention, provenance, access control and deletion need deliberate design around that layer.
Three: tool governance. MCP standardises a way to expose tools; it does not automatically authorise their use. Enforce permissions at the service boundary, scope credentials and keep records appropriate to the risk of each action.
Four: evaluation. Averages can hide serious failures. Measure recovery, budget use, data leakage and respect for approval boundaries. Do not claim published evaluation results unless readers can inspect the cases, setup and measured outputs.
Five: isolation. Sandboxes, network restrictions and human approval reduce different risks. Confirm what the selected execution backend actually isolates. Xavani coordinates work, Gavaza supports compliance evidence and Mhangani audits web security; none is a universal safety certificate.