Microsoft unveils Maia 200 AI inference chip for Azure, aiming faster and cheaper model deployment
Microsoft has announced Maia 200, its new AI inference accelerator built on advanced 3-nanometre technology and designed for large-scale deployment in Azure. The company says the chip focuses on improving inference economics for real-world apps, with early deployments in US data-centre regions and a software stack to help developers optimise models.
What Microsoft announced
Microsoft has introduced Maia 200, a new AI accelerator aimed specifically at inference—the stage where trained models generate outputs inside real applications. The company positioned Maia 200 as a chip designed to improve both speed and cost efficiency for running large models across its Azure cloud infrastructure.

According to Microsoft, Maia 200 is built on cutting-edge 3-nanometre technology and is designed for compute-intensive AI inference workloads. The company also highlighted a redesigned memory and scalable networking approach intended to keep large models highly utilised during production deployments.
Performance claims, memory and networking
Microsoft said the chip delivers performance exceeding 10 petaFLOPS at 4-bit precision (FP4) and more than 5 petaFLOPS at 8-bit precision (FP8). It also emphasised that the platform can scale across standard Ethernet to large clusters—useful for serving bigger models and higher traffic without re-architecting systems.
Alongside the silicon, Microsoft has pointed to an enabling software layer—tools and SDK components to help developers build, port, and optimise workloads across heterogeneous accelerators inside Azure, which remains important for enterprise adoption where compatibility and operational stability matter.
Where Maia 200 will be used first
Microsoft indicated Maia 200 will initially be deployed in US Azure regions and used for high-end AI workloads, including internal teams working on advanced models and deployment scenarios. The company also framed Maia 200 as part of its broader strategy to strengthen first-party infrastructure for AI, reducing dependence on any single external accelerator supplier and balancing performance with availability.
Why it matters for India’s AI builders
While the first deployments are in the US, the announcement matters for Indian developers and enterprises using Azure: better inference economics can translate into lower unit costs for chatbots, copilots, vision systems and analytics in production. Over time, if expanded globally, it could improve latency and cost efficiency for India-based deployments too, especially for high-volume applications.