Grok 4.6 comes to Microsoft Foundry Models: Built for long-horizon reasoning and complex workflows

The next wave of AI applications won’t be defined by how quickly models answer questions. They’ll be defined by how much work they can complete. Whether resolving repository issues, navigating codebases, generating engineering designs, or executing multi-step enterprise workflows, developers are increasingly building systems that need models capable of staying on task across complex, long-running objectives.

Today, Grok 4.6 from SpaceXAI is available in Microsoft Foundry Models through public preview, bringing SpaceXAI’s latest frontier model to developers through a unified platform for model discovery, evaluation, deployment, and governance.

Built on SpaceXAI’s 1.5T-scale model family, Grok 4.6 focuses on strong performance across coding, engineering, office productivity, research enablement, and inference optimization tasks. Rather than optimizing for a single benchmark category, Grok 4.6 is designed for sustained reasoning and execution across software engineering, agentic workflows, and technical problem solving.

Long-Running Agents That Stay On Task

Many AI systems perform well on isolated prompts. Real-world agents are different.

They need to plan, recover from errors, navigate tools, and continue making progress over extended workflows with minimal supervision.

Grok 4.6 is designed for this type of long-horizon agentic work, delivering strong performance on benchmarks that measure sustained software engineering and agent execution.

Benchmark highlights

Terminal-Bench 3.0

* Benchmarks are from model provider, see details.

Terminal-Bench 3.0 evaluates a model’s ability to complete multi-step, terminal-based execution that closely resemble real-world engineering workflows. For developers building coding agents, software engineering copilots, or autonomous workflows, Grok 4.6 is optimized not only for generating output, but for completing work.

Engineering-Grade Reasoning Beyond Software

Grok 4.6 extends beyond traditional coding into technical and engineering workflows. From computer-aided design (CAD) and procedural design, to complex engineering analysis, it is designed to reason through structured technical challenges and support multi-step problem solving.

Benchmark highlights

3DCodeBench

* Benchmarks are from model provider, see details.

3DCodeBench evaluates agentic procedural 3D modeling via code, testing how effectively models can reason about and generate complex engineering designs through code. The result highlights Grok 4.6’s potential beyond traditional software development, supporting technical workloads across manufacturing, robotics, product design, and industrial AI.

Enterprise Agents That Deliver Work, Not Just Answers

The most valuable AI workflows often don’t end with an answer. They produce business-ready deliverables, such as reports, presentations, spreadsheets, recommendations, and other business artifacts that can be acted on immediately.

Grok 4.6 demonstrates strong performance on evaluations designed to measure knowledge work and enterprise productivity.

Benchmark highlights

AA Briefcase

* Benchmarks are from model provider, see details.

AA Briefcase (Artificial Analysis) evaluates agents on long-horizon, complex professional knowledge-work projects that culminate in deliverables such as spreadsheets, presentations, memos, financial models, and PDFs. The results highlights Grok 4.6’s ability to support enterprise assistants, research agents, and business workflow automation that transform information into actionable business outcomes.

Pricing
Model Deployment Input/1M Tokens Output/1M Tokens Cache/1M Tokens
Grok 4.6Global Standard$2.00$6.00$0.50
Why build with Grok 4.6 in Foundry?

As organizations adopt more frontier models, the challenge is no longer accessing models. It’s operationalizing them.

Foundry provides a unified platform to discover, evaluate, and deploy models while applying enterprise-grade governance, security, and operational controls throughout the AI lifecycle. Organizations can assess Grok 4.6 against their own workloads, business requirements, and production constraints, enabling confident model selection and faster deployment.

With Grok 4.6 in Foundry, developers can:

  • Evaluate Grok 4.6 alongside other leading frontier models
  • Run workload-specific evaluations before deployment
  • Deploy through managed endpoints
  • Integrate the model into agentic applications and workflows
  • Operate with enterprise governance and security controls

The combination of Grok 4.6’s strengths in long-horizon execution, engineering reasoning, and enterprise productivity with Foundry’s model platform capabilities gives developers another powerful option for building the next generation of AI applications.

Get started

Grok 4.6 is now available in Microsoft Foundry Models. Explore the model card to learn more and get started.

If you’re building coding agents, engineering copilots, research assistants, or enterprise automation solutions, Grok 4.6 offers a compelling combination of reasoning, execution, and technical depth, all through the unified Foundry platform experience.

  This article was originally published by Microsoft's Azure AI Foundry Blog. You can find the original article here.