Speed, cost, location, and task fit matter more once agents run for minutes or hours. The best model on an average leaderboard may be the wrong model for the work in front of you.
EconomicsGPT-5.6 shifted the cost conversation to architecture
OpenAI · August 13
OpenAI highlighted smaller models, lower reasoning effort, retained reasoning, compaction, and prompt caching as routes to lower cost. One cited browser test reached 78% completion for about $14 versus 80% for roughly $235 with the comparison model. Treat vendor examples as leads for your own evals.
Review the reported examples →
SpeedUltrafast targeted useful work per second
OpenAI · August 13
A limited preview runs GPT-5.6 Sol up to 14 times faster and up to 750 output tokens per second. The interesting use cases are not faster essays. They are live incident response, voice, support, and interactive research where delay changes the workflow.
Read the preview →
BenchmarkingIndependent model testing became task-specific
Artificial Analysis · August 13
Artificial Analysis placed Gemini 3.7 Flash on an intelligence-versus-time frontier and continued pushing custom benchmarks built from an organization's own tasks. That is the right direction: measure the work you actually need, not the average of unrelated tests.
Read the independent analysis →
ModelGrok 4.6 competed on long-run economics
SpaceXAI · August 12
SpaceXAI positioned Grok 4.6 for long-running agents and multi-step work, with API pricing from $2 per million input tokens and $6 per million output tokens. The faster variant costs twice as much, making latency an explicit purchasing choice.
Read the announcement →
InfrastructureMistral made region part of the model decision
Mistral AI · August 11
Mistral tied open models, regional inference, and European compute commitments into one offer. For regulated teams, model quality sits beside data location, operational control, portability, and regional capacity.
Read the announcement →