Skip to content

Microsoft Large Language Models — MAI Family

🪟 Microsoft Large Language Models — MAI Family

Section titled “🪟 Microsoft Large Language Models — MAI Family”

Microsoft has shifted from purely consuming third-party frontier models to building and running its own. The MAI (Microsoft AI) family is that bet. At Microsoft Build 2026 (June 2–3, 2026), Microsoft launched seven new models in a single day — spanning reasoning, coding, image generation, speech transcription, and voice synthesis.

CEO Satya Nadella described the moment as a shift from “consuming a frontier model to fully participating at the frontier.”


🚀 Seven Models in One Day — Microsoft Build 2026

Section titled “🚀 Seven Models in One Day — Microsoft Build 2026”

All seven MAI models were announced on 2–3 June 2026 at Microsoft Build. They cover every major modality:

🤖 Model🗂️ Category⚙️ Key Spec🎯 Primary Use
🧠 MAI-Thinking-1Reasoning LLM35B active params · 256K contextComplex reasoning, agent planning
💻 MAI-Code-1-FlashCoding LLM5B active / 137B total paramsGitHub Copilot, VS Code, code generation
🖼️ MAI-Image-2.5Image generationText-to-image + image-to-imageCreative visuals, image editing
🖼️ MAI-Image-2.5 FlashImage generation (efficient)Optimised for production throughputHigh-volume image workloads
🎙️ MAI-Transcribe-1.5Speech-to-text43 languages · streaming supportMeeting transcripts, audio search
🔊 MAI-Voice-2Text-to-speech15+ new languages · voice cloningVoice agents, narration
🔊 MAI-Voice-2 FlashText-to-speech (efficient)Low-latency optimised variantLatency-sensitive voice agents

🧠 MAI-Thinking-1 — Reasoning Flagship

Section titled “🧠 MAI-Thinking-1 — Reasoning Flagship”

MAI-Thinking-1 is Microsoft’s first published reasoning LLM. It uses a mixture-of-experts (MoE) architecture, activating only 35 billion parameters out of a much larger total — giving it strong reasoning at a low inference cost.

🔧 Property📋 Detail
🏷️ Model typeReasoning LLM (MoE architecture)
⚙️ Active parameters35 billion
📄 Context window256,000 tokens
🌐 AvailabilityPrivate preview on Azure AI Foundry
📍 Also available onFireworks AI · Baseten · OpenRouter
  • 🧩 Multi-step reasoning and chain-of-thought analysis
  • 🏗️ Architecture decisions and system design
  • 🚨 Production debugging and root-cause analysis
  • 🤖 Agent planning and orchestration
  • 💻 Complex coding tasks that require reasoning beyond code generation
  • On blind tests, independent raters prefer it to Sonnet 4.6
  • Matches Opus 4.6 on coding benchmarks (SWE-Bench Pro)
  • Positioned as a cost-efficient alternative to the largest frontier models

“We have a distributed system with 12 microservices. Our payment service has 99.2% uptime but drops to 97.1% under Black Friday load. Analyse the likely failure points and propose a mitigation plan.”


💻 MAI-Code-1-Flash — Coding for GitHub and VS Code

Section titled “💻 MAI-Code-1-Flash — Coding for GitHub and VS Code”

MAI-Code-1-Flash is purpose-built for software engineering workflows. It is already live inside GitHub Copilot and VS Code, making it one of the most immediately accessible MAI models.

🔧 Property📋 Detail
🏷️ Model typeCoding LLM (MoE architecture)
⚙️ Active parameters5 billion active / 137 billion total
🌐 AvailabilityLive in GitHub Copilot and VS Code
⚡ Design goalHigh performance at low inference cost
  • ✏️ Code completion and generation from natural language
  • 🐛 Bug detection and fix suggestions
  • 🔄 Refactoring and code transformation
  • 🧪 Test generation
  • 📝 Inline code explanation and documentation
  • 🔀 Transpilation (e.g. Python → TypeScript)

“Add pagination to this Express.js route. Use cursor-based pagination, return a nextCursor field in the response, and add a Jest test for the edge case where the last page has fewer items than the page size.”


🖼️ MAI-Image-2.5 — Image Generation and Editing

Section titled “🖼️ MAI-Image-2.5 — Image Generation and Editing”

MAI-Image-2.5 is Microsoft’s flagship image model. It is the first Microsoft model to support both text-to-image and image-to-image editing in a single model.

🔧 Property📋 Detail
🏷️ Model typeImage generation + image editing
🏆 Text-to-image rank#3 on Arena AI leaderboard
🏆 Image-to-image rank#2 on Arena AI leaderboard (surpassing Nano Banana 2)
🌐 AvailabilityAzure AI Foundry
🤖 Variant⚡ Speed🎯 Use Case
🖼️ MAI-Image-2.5StandardHighest quality text-to-image and editing
⚡ MAI-Image-2.5 FlashFasterHigh-volume production image workloads
  • 🎨 Photorealistic image generation from text prompts
  • ✏️ Image editing and style transfer
  • 🖼️ UI mockups, blog thumbnails, and marketing visuals
  • 🏷️ Text-based image generation (logos, diagrams with embedded text)

“Generate a photorealistic hero image for a developer conference landing page. The scene should show a futuristic developer workspace with blue ambient lighting, multiple monitors showing code, and a subtle Microsoft logo reflection.”


🎙️ MAI-Transcribe-1.5 — Speech-to-Text at Scale

Section titled “🎙️ MAI-Transcribe-1.5 — Speech-to-Text at Scale”

MAI-Transcribe-1.5 is Microsoft’s latest speech recognition model, designed for enterprise-grade transcription across many languages.

🔧 Property📋 Detail
🏷️ Model typeSpeech-to-text
🌍 Languages supported43 languages
⚡ SpeedClaimed 5× faster than competing models
🔄 StreamingComing soon
🏢 Domain supportBuilt-in domain-specific terminology recognition
🌐 AvailabilityAzure AI Foundry
  • 📋 Meeting and call transcription at enterprise scale
  • 🌍 Multilingual transcription in 43 languages
  • 🏥 Domain-specific accuracy (medical, legal, technical vocabulary)
  • 🎬 Subtitle and caption generation
  • 🔍 Audio content indexing for search

🎙️ Input: 60-minute board meeting recording in English and Spanish → 📄 Output: timestamped bilingual transcript with speaker diarisation


🔊 MAI-Voice-2 — Text-to-Speech and Voice Synthesis

Section titled “🔊 MAI-Voice-2 — Text-to-Speech and Voice Synthesis”

MAI-Voice-2 is Microsoft’s latest voice generation model. It supports 15+ new languages, voice cloning from short audio samples, and anti-abuse protections.

🔧 Property📋 Detail
🏷️ Model typeText-to-speech / voice synthesis
🌍 New languages15+ additional languages
🎭 Voice cloning✅ Supported — adapts from short audio samples
🛡️ SafetyBuilt-in anti-abuse protection
🌐 AvailabilityAzure AI Foundry
🤖 Variant⚡ Speed🎯 Use Case
🔊 MAI-Voice-2StandardHighest quality narration and voice agents
⚡ MAI-Voice-2 FlashLow-latencyReal-time voice agents and interactive conversations
  • 🤖 Conversational voice agents
  • 📖 Narration and audiobook generation
  • ♿ Accessibility — screen readers and assistive tech
  • 📞 IVR and enterprise phone systems
  • 🌍 Multilingual customer support automation

A customer service voice bot that detects the caller’s language, transcribes via MAI-Transcribe-1.5, reasons with MAI-Thinking-1, and responds via MAI-Voice-2 in the caller’s language — all on Azure.


Before Build 2026, Microsoft quietly developed MAI-1, a much larger foundational model.

🔧 Property📋 Detail
🏷️ NameMAI-1
📅 Announced~May 2024
⚙️ Parameters~500 billion
🎯 PurposeFoundational general-purpose LLM for Microsoft services
🌐 AccessInternal / limited Azure preview

MAI-1 was Microsoft’s first major step toward training frontier-scale models in-house, and it laid the groundwork for the more focused and efficient MAI models announced at Build 2026.


🏷️ Platform🤖 Models Available🔗 Access Type
🔷 Azure AI FoundryAll MAI modelsAPI — pay per token
🐙 GitHub CopilotMAI-Code-1-FlashBuilt-in to Copilot
🟦 VS CodeMAI-Code-1-FlashBuilt-in extension
🔥 Fireworks AIMAI-Thinking-1, othersThird-party API
🚀 BasetenMAI-Thinking-1, othersThird-party deployment
🔀 OpenRouterMAI-Thinking-1, othersUnified API gateway

🎯 Task✅ Best MAI Model💡 Why
🏗️ Architecture decisions or migration planningMAI-Thinking-1Requires deep multi-step reasoning
🚨 Production incident debuggingMAI-Thinking-1High-stakes reasoning with large context
💻 Code completion in VS Code or CopilotMAI-Code-1-FlashAlready integrated, low latency
🔄 Refactoring or test generationMAI-Code-1-FlashFast and purpose-built for software tasks
🖼️ Creating a blog image or marketing visualMAI-Image-2.5Highest quality output
📸 High-volume image batch generationMAI-Image-2.5 FlashEfficient variant for production scale
📋 Transcribing a meeting or callMAI-Transcribe-1.543 languages, enterprise accuracy
🔊 Building a voice agentMAI-Voice-2 FlashLow-latency for real-time conversation
📖 Narrating long-form contentMAI-Voice-2Highest quality voice output

  • MAI-Thinking-1 is in private preview — request access through Azure AI Foundry
  • MAI-Code-1-Flash is available now via GitHub Copilot and VS Code
  • Pricing for all models is token-based through Azure; third-party platforms (Fireworks, Baseten, OpenRouter) have their own pricing
  • All models are designed to run on Azure infrastructure, reducing cost compared to paying OpenAI API rates
  • MAI models are separate from OpenAI models on Azure — they are Microsoft’s own, not GPT variants