The best AI models for agent work: August 2026
Published 27 August 2026 · Editorial update 6 September 2026 · 3 min read · Enternovate

This is an archived comparison dated 27 August 2026. The chart retains the previously recorded Artificial Analysis snapshot and links to its model pages. We have not revalidated those figures in this editorial update. Availability, prices and rankings can change; check the linked sources and provider documentation before buying or configuring a model.
The original comparison highlighted GLM-5.3-Flash for its recorded balance of quality and cost. That is a historical editorial choice, not a universal default or a promise that it will win on your workload. A provider name and model name alone do not identify the full configuration.
Choose a small evaluation set from work you actually need done. Include a code change with tests, a research question with source checks, a document task and an action that requires approval. Define success before running the agent.
Record the model build, provider, reasoning setting, context limit and tool permissions. Compare completed tasks, retries, latency and cost per successful task. A low token price can still produce an expensive workflow if the agent repeatedly fails.
The chart separates benchmark intelligence from output-token pricing. Neither measures deployment privacy or suitability for your business. Missing prices remain Not listed. Do not infer a hosting cost for open weights from a managed API price.
Archived model comparison
Intelligence against output price
Recorded Artificial Analysis snapshot: 27 August 2026. Not revalidated in this editorial update. Not a live price list or current recommendation. Prices are USD per 1M tokens. The horizontal scale is logarithmic.
Qwen3.8-Flash-Next is included with “Not listed” pricing because Artificial Analysis had no public API price at this snapshot. We do not substitute a guessed hosting price.
Xavani lets you change the model while retaining your working environment. Re-test tool calling and permissions after switching. For local inference, also measure memory use and performance on the actual machine. For hosted inference, review retention, data location and contractual terms.
Pin the configuration that meets your acceptance criteria and keep a rollback option. This archive explains a selection method; it does not replace a fresh evaluation before production use.