Microsoft Large Language Models — MAI Family
🪟 Microsoft Large Language Models — MAI Family
Section titled “🪟 Microsoft Large Language Models — MAI Family”Microsoft has shifted from purely consuming third-party frontier models to building and running its own. The MAI (Microsoft AI) family is that bet. At Microsoft Build 2026 (June 2–3, 2026), Microsoft launched seven new models in a single day — spanning reasoning, coding, image generation, speech transcription, and voice synthesis.
CEO Satya Nadella described the moment as a shift from “consuming a frontier model to fully participating at the frontier.”
🚀 Seven Models in One Day — Microsoft Build 2026
Section titled “🚀 Seven Models in One Day — Microsoft Build 2026”All seven MAI models were announced on 2–3 June 2026 at Microsoft Build. They cover every major modality:
| 🤖 Model | 🗂️ Category | ⚙️ Key Spec | 🎯 Primary Use |
|---|---|---|---|
| 🧠 MAI-Thinking-1 | Reasoning LLM | 35B active params · 256K context | Complex reasoning, agent planning |
| 💻 MAI-Code-1-Flash | Coding LLM | 5B active / 137B total params | GitHub Copilot, VS Code, code generation |
| 🖼️ MAI-Image-2.5 | Image generation | Text-to-image + image-to-image | Creative visuals, image editing |
| 🖼️ MAI-Image-2.5 Flash | Image generation (efficient) | Optimised for production throughput | High-volume image workloads |
| 🎙️ MAI-Transcribe-1.5 | Speech-to-text | 43 languages · streaming support | Meeting transcripts, audio search |
| 🔊 MAI-Voice-2 | Text-to-speech | 15+ new languages · voice cloning | Voice agents, narration |
| 🔊 MAI-Voice-2 Flash | Text-to-speech (efficient) | Low-latency optimised variant | Latency-sensitive voice agents |
🧠 MAI-Thinking-1 — Reasoning Flagship
Section titled “🧠 MAI-Thinking-1 — Reasoning Flagship”MAI-Thinking-1 is Microsoft’s first published reasoning LLM. It uses a mixture-of-experts (MoE) architecture, activating only 35 billion parameters out of a much larger total — giving it strong reasoning at a low inference cost.
📊 Specifications
Section titled “📊 Specifications”| 🔧 Property | 📋 Detail |
|---|---|
| 🏷️ Model type | Reasoning LLM (MoE architecture) |
| ⚙️ Active parameters | 35 billion |
| 📄 Context window | 256,000 tokens |
| 🌐 Availability | Private preview on Azure AI Foundry |
| 📍 Also available on | Fireworks AI · Baseten · OpenRouter |
✅ What It Does Well
Section titled “✅ What It Does Well”- 🧩 Multi-step reasoning and chain-of-thought analysis
- 🏗️ Architecture decisions and system design
- 🚨 Production debugging and root-cause analysis
- 🤖 Agent planning and orchestration
- 💻 Complex coding tasks that require reasoning beyond code generation
🏆 How It Compares
Section titled “🏆 How It Compares”- On blind tests, independent raters prefer it to Sonnet 4.6
- Matches Opus 4.6 on coding benchmarks (SWE-Bench Pro)
- Positioned as a cost-efficient alternative to the largest frontier models
💡 Example Prompt
Section titled “💡 Example Prompt”“We have a distributed system with 12 microservices. Our payment service has 99.2% uptime but drops to 97.1% under Black Friday load. Analyse the likely failure points and propose a mitigation plan.”
💻 MAI-Code-1-Flash — Coding for GitHub and VS Code
Section titled “💻 MAI-Code-1-Flash — Coding for GitHub and VS Code”MAI-Code-1-Flash is purpose-built for software engineering workflows. It is already live inside GitHub Copilot and VS Code, making it one of the most immediately accessible MAI models.
📊 Specifications
Section titled “📊 Specifications”| 🔧 Property | 📋 Detail |
|---|---|
| 🏷️ Model type | Coding LLM (MoE architecture) |
| ⚙️ Active parameters | 5 billion active / 137 billion total |
| 🌐 Availability | Live in GitHub Copilot and VS Code |
| ⚡ Design goal | High performance at low inference cost |
✅ What It Does Well
Section titled “✅ What It Does Well”- ✏️ Code completion and generation from natural language
- 🐛 Bug detection and fix suggestions
- 🔄 Refactoring and code transformation
- 🧪 Test generation
- 📝 Inline code explanation and documentation
- 🔀 Transpilation (e.g. Python → TypeScript)
💡 Example Prompt
Section titled “💡 Example Prompt”“Add pagination to this Express.js route. Use cursor-based pagination, return a
nextCursorfield in the response, and add a Jest test for the edge case where the last page has fewer items than the page size.”
🖼️ MAI-Image-2.5 — Image Generation and Editing
Section titled “🖼️ MAI-Image-2.5 — Image Generation and Editing”MAI-Image-2.5 is Microsoft’s flagship image model. It is the first Microsoft model to support both text-to-image and image-to-image editing in a single model.
📊 Specifications
Section titled “📊 Specifications”| 🔧 Property | 📋 Detail |
|---|---|
| 🏷️ Model type | Image generation + image editing |
| 🏆 Text-to-image rank | #3 on Arena AI leaderboard |
| 🏆 Image-to-image rank | #2 on Arena AI leaderboard (surpassing Nano Banana 2) |
| 🌐 Availability | Azure AI Foundry |
📦 Two Variants
Section titled “📦 Two Variants”| 🤖 Variant | ⚡ Speed | 🎯 Use Case |
|---|---|---|
| 🖼️ MAI-Image-2.5 | Standard | Highest quality text-to-image and editing |
| ⚡ MAI-Image-2.5 Flash | Faster | High-volume production image workloads |
✅ What It Does Well
Section titled “✅ What It Does Well”- 🎨 Photorealistic image generation from text prompts
- ✏️ Image editing and style transfer
- 🖼️ UI mockups, blog thumbnails, and marketing visuals
- 🏷️ Text-based image generation (logos, diagrams with embedded text)
💡 Example Prompt
Section titled “💡 Example Prompt”“Generate a photorealistic hero image for a developer conference landing page. The scene should show a futuristic developer workspace with blue ambient lighting, multiple monitors showing code, and a subtle Microsoft logo reflection.”
🎙️ MAI-Transcribe-1.5 — Speech-to-Text at Scale
Section titled “🎙️ MAI-Transcribe-1.5 — Speech-to-Text at Scale”MAI-Transcribe-1.5 is Microsoft’s latest speech recognition model, designed for enterprise-grade transcription across many languages.
📊 Specifications
Section titled “📊 Specifications”| 🔧 Property | 📋 Detail |
|---|---|
| 🏷️ Model type | Speech-to-text |
| 🌍 Languages supported | 43 languages |
| ⚡ Speed | Claimed 5× faster than competing models |
| 🔄 Streaming | Coming soon |
| 🏢 Domain support | Built-in domain-specific terminology recognition |
| 🌐 Availability | Azure AI Foundry |
✅ What It Does Well
Section titled “✅ What It Does Well”- 📋 Meeting and call transcription at enterprise scale
- 🌍 Multilingual transcription in 43 languages
- 🏥 Domain-specific accuracy (medical, legal, technical vocabulary)
- 🎬 Subtitle and caption generation
- 🔍 Audio content indexing for search
💡 Example Prompt
Section titled “💡 Example Prompt”🎙️ Input: 60-minute board meeting recording in English and Spanish → 📄 Output: timestamped bilingual transcript with speaker diarisation
🔊 MAI-Voice-2 — Text-to-Speech and Voice Synthesis
Section titled “🔊 MAI-Voice-2 — Text-to-Speech and Voice Synthesis”MAI-Voice-2 is Microsoft’s latest voice generation model. It supports 15+ new languages, voice cloning from short audio samples, and anti-abuse protections.
📊 Specifications
Section titled “📊 Specifications”| 🔧 Property | 📋 Detail |
|---|---|
| 🏷️ Model type | Text-to-speech / voice synthesis |
| 🌍 New languages | 15+ additional languages |
| 🎭 Voice cloning | ✅ Supported — adapts from short audio samples |
| 🛡️ Safety | Built-in anti-abuse protection |
| 🌐 Availability | Azure AI Foundry |
📦 Two Variants
Section titled “📦 Two Variants”| 🤖 Variant | ⚡ Speed | 🎯 Use Case |
|---|---|---|
| 🔊 MAI-Voice-2 | Standard | Highest quality narration and voice agents |
| ⚡ MAI-Voice-2 Flash | Low-latency | Real-time voice agents and interactive conversations |
✅ What It Does Well
Section titled “✅ What It Does Well”- 🤖 Conversational voice agents
- 📖 Narration and audiobook generation
- ♿ Accessibility — screen readers and assistive tech
- 📞 IVR and enterprise phone systems
- 🌍 Multilingual customer support automation
💡 Example Use Case
Section titled “💡 Example Use Case”A customer service voice bot that detects the caller’s language, transcribes via MAI-Transcribe-1.5, reasons with MAI-Thinking-1, and responds via MAI-Voice-2 in the caller’s language — all on Azure.
🗓️ Earlier Microsoft LLM — MAI-1
Section titled “🗓️ Earlier Microsoft LLM — MAI-1”Before Build 2026, Microsoft quietly developed MAI-1, a much larger foundational model.
| 🔧 Property | 📋 Detail |
|---|---|
| 🏷️ Name | MAI-1 |
| 📅 Announced | ~May 2024 |
| ⚙️ Parameters | ~500 billion |
| 🎯 Purpose | Foundational general-purpose LLM for Microsoft services |
| 🌐 Access | Internal / limited Azure preview |
MAI-1 was Microsoft’s first major step toward training frontier-scale models in-house, and it laid the groundwork for the more focused and efficient MAI models announced at Build 2026.
🌐 Where to Access MAI Models
Section titled “🌐 Where to Access MAI Models”| 🏷️ Platform | 🤖 Models Available | 🔗 Access Type |
|---|---|---|
| 🔷 Azure AI Foundry | All MAI models | API — pay per token |
| 🐙 GitHub Copilot | MAI-Code-1-Flash | Built-in to Copilot |
| 🟦 VS Code | MAI-Code-1-Flash | Built-in extension |
| 🔥 Fireworks AI | MAI-Thinking-1, others | Third-party API |
| 🚀 Baseten | MAI-Thinking-1, others | Third-party deployment |
| 🔀 OpenRouter | MAI-Thinking-1, others | Unified API gateway |
🎯 When to Use Which MAI Model
Section titled “🎯 When to Use Which MAI Model”| 🎯 Task | ✅ Best MAI Model | 💡 Why |
|---|---|---|
| 🏗️ Architecture decisions or migration planning | MAI-Thinking-1 | Requires deep multi-step reasoning |
| 🚨 Production incident debugging | MAI-Thinking-1 | High-stakes reasoning with large context |
| 💻 Code completion in VS Code or Copilot | MAI-Code-1-Flash | Already integrated, low latency |
| 🔄 Refactoring or test generation | MAI-Code-1-Flash | Fast and purpose-built for software tasks |
| 🖼️ Creating a blog image or marketing visual | MAI-Image-2.5 | Highest quality output |
| 📸 High-volume image batch generation | MAI-Image-2.5 Flash | Efficient variant for production scale |
| 📋 Transcribing a meeting or call | MAI-Transcribe-1.5 | 43 languages, enterprise accuracy |
| 🔊 Building a voice agent | MAI-Voice-2 Flash | Low-latency for real-time conversation |
| 📖 Narrating long-form content | MAI-Voice-2 | Highest quality voice output |
⚠️ Things to Know
Section titled “⚠️ Things to Know”- MAI-Thinking-1 is in private preview — request access through Azure AI Foundry
- MAI-Code-1-Flash is available now via GitHub Copilot and VS Code
- Pricing for all models is token-based through Azure; third-party platforms (Fireworks, Baseten, OpenRouter) have their own pricing
- All models are designed to run on Azure infrastructure, reducing cost compared to paying OpenAI API rates
- MAI models are separate from OpenAI models on Azure — they are Microsoft’s own, not GPT variants