← Back to the wire31 Aug 2026
The AI WireDispatch No. 068
azureproduct· filed 30 Aug 2026 · 2 min read

Grok 4.6 available in Microsoft Foundry public preview

SpaceXAI's Grok 4.6 is now available in Microsoft Foundry's public preview, aimed at long-horizon agentic work like repository issue resolution and business deliverables. Provider-reported benchmarks show strong performance on terminal and 3D modeling tasks, but production reliability remains unproven outside controlled evaluations.

Machine-drafted illustration · reviewed by a humanFIG. 01

Listen to this dispatch

Narrated by an AI-generated voice.

Grok 4.6, SpaceXAI's latest frontier model, is now available through Microsoft Foundry Models in public preview. Built on what the company describes as its 1.5T-scale model family, Grok 4.6 is positioned for long-horizon agentic work — sustaining reasoning and execution across multi-step workflows rather than answering isolated prompts.

Long-horizon agentic work requires planning, tool navigation, error recovery, and sustained progress over extended runs with minimal supervision. SpaceXAI's stated emphasis is completing work — resolving repository issues, generating engineering designs, producing business deliverables — rather than merely generating output.

Benchmark results are provider-reported. SpaceXAI cites three evaluations. Terminal-Bench 3.0 measures multi-step terminal-based execution that resembles real engineering workflows. 3DCodeBench tests agentic procedural 3D modeling via code, aimed at CAD, manufacturing, robotics, and product design use cases. AA Briefcase, from Artificial Analysis, assesses long-horizon professional knowledge work culminating in deliverables like spreadsheets, presentations, financial models, and memos. The benchmark origin for the latter is third-party, but the quoted results are still provider-reported.

Pricing at the Global Standard tier is $2.00 per million input tokens, $6.00 per million output tokens, and $0.50 per million cache tokens. For agentic workloads, output tokens accumulate quickly over a long-horizon run, so the $6.00 output rate carries more weight than the input rate.

The Foundry integration addresses operationalization. Microsoft positions the platform as a unified layer for discovery, evaluation, governance, and deployment, letting organizations assess Grok 4.6 against their own workloads and production constraints before committing. For teams already in the Azure ecosystem, the practical advantage is comparative — Grok 4.6 can be evaluated alongside other frontier models in the same platform, rather than requiring a separate integration effort. The model card contains implementation details.

Whether long-horizon benchmark performance carries into production agents remains open. Benchmarks like Terminal-Bench measure controlled multi-step execution; production agent failures tend to surface in tool incompatibilities, ambiguous task definitions, and integration details that benchmarks don't capture. SpaceXAI's stated next step is workload-specific evaluation in Foundry before deployment — which is also the only way to verify whether the "work completion" framing holds outside a benchmark harness.

Read the original at techcommunity.microsoft.com

End of dispatch
More on the wire
72simonwillison.netproduct · filed 31 Aug 2026Simon Willison's hands-on test of ChatGPT Work: open internet access, browser automation, and unanswered safety questions71youtube.commodels · filed 30 Aug 2026GLM 5.3 Flash: Z.ai’s cheaper, multimodal sibling matches larger models but slows on max reasoning70youtube.comindustry · filed 30 Aug 2026Andrew Ng: Big AI labs are fear-mongering to shape regulation69azureazure · filed 30 Aug 2026VNet integration for Azure SRE Agent now generally available