Aurora 1.5
Earth-system Foundation Model
Aurora 1.5 is Microsoft’s open Earth-system foundation model that combines 22 new weather variables, hourly forecasting, and ensemble forecasting to deliver more accurate weather and climate intelligence.
Our Experiments
Browse, filter, and explore every experiment live in Foundry Labs. Built from frontier research. Available to you today.
Earth-system Foundation Model
Aurora 1.5 is Microsoft’s open Earth-system foundation model that combines 22 new weather variables, hourly forecasting, and ensemble forecasting to deliver more accurate weather and climate intelligence.
Autoregressive Vector Map Extraction
A feature-extraction model that converts satellite imagery into attributed, vectorized GIS data. MARS uses an autoregressive transformer to generate native GeoJSON geometries — polygons for buildings, polylines for roads and railways — each with category labels and confidence scores
Controllable Image Generation and Editing
Microsoft AI's flagship image-generation model adds image-to-image editing with identity preservation, style control, and PPT-ready outputs. Debuted at #3 on Arena.ai with an average +74.5 ELO over MAI-Image-2, +104 on text rendering.
Mid-Sized Sparse MoE Reasoning Model
Microsoft AI's first large language model — 35B-active, ~1T-total sparse MoE trained from scratch on clean data. Matches Claude Opus 4.6 on SWE-Bench Pro and earns gold at IMO 2025, at a fraction of frontier cost.
High-Speed Multilingual Speech Recognition
Microsoft AI's updated speech-to-text model retains #1 on FLEURS with a 3.9% → 3.7% WER improvement and adds content biasing for domain-specific vocabulary. Transcribes one hour of audio in under 10 seconds — over 5× faster than competing systems.
Multilingual TTS with Voice Cloning
Microsoft AI's updated multilingual text-to-speech model adds zero-shot voice cloning from 10–60 seconds of reference audio, voice prompting, and granular emotion control. Preferred over MAI-Voice-1 by 72% in blind tests; often indistinguishable from real speech.
Next-Gen Computer-Use Model Family
A family of three small computer-use models — 4B, 9B, and 27B — built for browser automation. The 9B flagship sets SOTA among small computer-use models on Online-Mind2Web, nearly doubling Fara-7B.
Orchestration Model for Agentic Systems
An 8B-parameter orchestration model from Microsoft Research that plans, codes, and delegates within the MagenticLite agent stack. Turns natural-language requests into concrete plans, picks tools or sub-agents, writes code, and recovers from errors mid-task.
Open-Source Agentic App for Small Models
An agentic application from Microsoft Research that runs on small models — the next generation of Magentic-UI. Refreshed chat and browser view with a harness rebuilt for MagenticBrain and Fara1.5, all sandboxed in Quicksand.
Long-Form Multi-Speaker Transcription
A unified speech-to-text model designed to transcribe up to 60 minutes of continuous audio in a single pass, producing structured output capturing who said what and when. Jointly performs ASR, speaker diarization, and timestamping; supports 50+ languages and hotwords.
Earth Observation & Overhead Sensing
Microsoft first-party model that identifies and localizes objects in satellite and aerial imagery, returning bounding-box detections optimized for batch processing of large image archives. Built by the Spectre team behind Planetary Computer; part of the new GeoAI category in Microsoft Foundry.
Agentic SLM for Computer Use
Microsoft's first agentic small language model designed for computer use. With 7B parameters, Fara-7B achieves SOTA performance within its size class on WebVoyager / Online-Mind2Web / DeepShop / WebTailBench. Built on Qwen2.5-VL; perceives via screenshots and predicts coordinates directly.
Open-Source Agentic Market Simulation
An open-source simulation environment for studying agentic markets at scale. Models large numbers of agents simultaneously searching, communicating, and transacting; explores welfare, efficiency, fairness, manipulation resistance, and bias in agentic economies.
High-Speed Text-to-Image
Microsoft's latest text-to-image model engineered for high-volume production. Built on the MAI-Image-2 architecture (debuted at #3 on Arena.ai), Image-2e is up to 22% faster with 4× efficiency vs. MAI-Image-2 and outpaces leading T2I models by ~40% on average.
Synthetic Bug Generation for SWE Agents
A pipeline where LLM agents create new features and unintentionally break tests — yielding realistic bugs more similar to human-generated ones. Produced FrogBoss (32B) and FrogMini (14B), Qwen3-based coding agents specialized in bug fixing on SWE-Bench Verified.
Open-Source Multilingual Text Embeddings
A family of open-source multilingual text embedding models from Microsoft, delivering SOTA retrieval and semantic understanding across 94 languages. Decoder-only architecture with last-token pooling and L2 normalization; instruction-tunable. Sizes: 270M, 0.6B, 27B.
State-of-the-Art Text-to-Image
State-of-the-art text-to-image model from Microsoft AI, debuted at #3 on the Arena.ai leaderboard. Photorealistic generation with reliable text rendering and expressive range, designed in collaboration with photographers, designers, and visual storytellers.
Multilingual Speech Recognition
Speech recognition model that supports up to 25 languages, delivering enterprise-grade transcription accuracy at close to half the GPU cost of leading systems. Engineered for accessibility tools, captioning, content workflows, and voice agents.
Open-Source Engine for Agentic AI Apps
The open-source framework for developers and AI engineers building production agentic applications — the direct successor to Semantic Kernel and AutoGen. Combines simple agent abstractions with enterprise-grade state management, type safety, middleware, telemetry, and multi-agent orchestration.
Compact Multimodal Reasoning Model
A compact open-weight multimodal reasoning model built on Phi-4-Reasoning + SigLIP-2 with a mid-fusion architecture. Hybrid reasoning design (THINK / NOTHINK modes), high-resolution perception with up to 3,600 visual tokens — competitive with models 10× its size.
Benchmark for Agent Social Reasoning
An open-source benchmark from Microsoft Research AI Frontiers that measures whether AI agents can negotiate competently and advocate for their user's best interest in multi-party settings. Scores agents on Calendar Coordination and Marketplace Negotiation across Outcome Optimality and Due Diligence.
H&E to Virtual mIF Translation
A multimodal AI model for translating routine H&E pathology slides into virtual multiplex immunofluorescence (mIF) images across 21 protein channels. Trained on a Providence dataset of 40M cells; applied to 14,256 patients across 24 cancer types.
NL-to-Optimization SLM
A small language model that converts business problems described in natural language into the mathematical formulations needed by optimization software. Trained on expert-aligned data with domain-specific hints and inference-time self-checks.
Neural Text-to-Speech
A lightning-fast next-generation neural TTS model that generates a full minute of audio in under 1 second on a single GPU. Supports SSML-based emotion/style control, real-time synthesis via Azure Speech SDK, and stable voice persona across long-form content.
Geologic Map Understanding for MLLMs
Enhances multimodal LLMs with geologic expertise (emPowering gEologic mAp holistiC undErstanding). Combines Hierarchical Information Extraction, Domain Knowledge Injection, and Prompt-enhanced QA. Includes GeoMap-Bench, the first benchmark for geologic-map MLLMs.
Code Editing via SeleKT Adaptation
A model series adapted from QwenCoder-2.5 for code editing. Uses a synthetic data pipeline plus the SeleKT adaptation algorithm — a dense gradient step to identify edit-critical weights and a sparse projection back onto the base model — outperforming peers on five code-editing benchmarks.
Deep-Learning DFT Functional
A deep-learning-based exchange-correlation (XC) functional for density functional theory that achieves experimental-level accuracy on atomization energies. Competitive with the best hybrid functionals while retaining semi-local DFT computational efficiency.
Copilot Tools for Lab Models
An MCP Server that equips GitHub Copilot with tools for model discovery, integration guidance, and rapid prototyping across 46+ Foundry Labs models — reducing the idea-to-prototype cycle to under 10 minutes.
AI Chatbots for Cognitive Training
A web-based framework that helps researchers build AI chatbots for personalized memory and cognitive training. Combines a puzzle engine, life-logging module, and multimodal training interface (text, image, voice) for digital health research.
Unified Biomolecular Structure Prediction
A unified biomolecular modeling system that predicts 3D structures of proteins, nucleic acids, and small molecules within a single framework. Combines multimodal transformers and generative diffusion; supports protein–ligand, protein–DNA, protein–RNA interactions with atom-level conditioning.
Human-in-the-Loop Agent Platform
A research platform for advancing human-AI collaboration: co-planning, co-tasking, action guards, task learning, and Sentinel Steps for long-running monitoring tasks. Sandboxed for safe operation of browsers and code executors.
Real-Time LLM Routing
A trained language model in Microsoft Foundry that routes each prompt in real time to the most suitable underlying LLM. Supports 18 models across OpenAI, Anthropic, DeepSeek, Meta, and xAI under one deployment, with Balanced/Cost/Quality routing modes, failover, prompt caching, and tool-use.
Dynamic Prompt UI Middleware
Dynamically generates UI elements (radio buttons, checkboxes, toggles) to help users steer AI responses. Renders inline prompt options refined in real time, grounded in user studies showing strong preference vs. static controls.
Native 1-bit LLM at 2B Scale
The first open-source, native 1-bit LLM at 2B parameter scale, trained from scratch with W1.58A8 quantization on 4T tokens. Achieves performance comparable to leading full-precision models while dramatically reducing memory, energy, and latency. Optimized inference via bitnet.cpp.
Interactive Debugging Environment for LLM Agents
An open-source research environment for teaching AI coding agents to debug interactively with tools like Python's pdb. Includes benchmarks (Aider, Mini-nightmare, SWE-bench) and integration with swe-smith.
Protein Structural Ensembles
A diffusion-based generative model from Microsoft Research that predicts the structural ensembles of proteins — not a single fold, but the full conformational landscape — generated efficiently at scale on a single GPU.
Multimodal Foundation Model for AI Agents
A multimodal foundation model that perceives text and visuals and generates actions in both digital and physical environments — navigating UIs and manipulating real-world tools. Innovations: Set-of-Mark and Trace-of-Mark trained on unlabeled video at scale.
World and Human Action Model (WHAM)
A generative AI model of a video game (Bleeding Edge) that can produce visuals, controller actions, or both. Trained on 1B+ images and controller actions (~7 years of gameplay). The WHAM demonstrator on Foundry enables creators to explore and tweak gameplay continuations.
Pure-Vision GUI Screen Parser
A general screen parsing module that converts UI screenshots into structured, actionable elements. Combines a fine-tuned YOLOv8 icon detector with a Florence-2-based caption model. V2 delivers 60% lower latency and 39.6 average accuracy on ScreenSpot Pro.
Generative Model for Inorganic Materials
A diffusion-based generative model for inorganic materials design. Jointly predicts atomic coordinates, elements, and lattice vectors; supports property-guided generation (bulk modulus, band gap, chemical system, magnetic density). Published in Nature.
Atomistic Materials Simulator
A deep learning atomistic model spanning 0–5,000 K and pressures up to 10⁷ atm. Handles metals, oxides, sulfides, halides and crystalline/amorphous/liquid phases. Supports customization with user data for in silico materials design.
Phi-4 Family (Multimodal, Mini)
Phi-4 expands Microsoft's SLM family with Phi-4-multimodal (5.6B, integrates speech/vision/text in a unified architecture) and Phi-4-mini (3.8B, excels at reasoning, coding, and long-context tasks with 128K context).
Self-Evolving Prompt Optimization
A self-evolving framework that automates prompt optimization. Iteratively refines instructions and in-context examples using LLM feedback, jointly optimizes prompts + examples, and synthesizes chain-of-thought reasoning — delivering high-quality prompts in minutes.
Image-to-3D Asset Generation
Generates high-quality 3D assets from a single image or text. Built on Structured LATent (SLat) representation; outputs meshes, radiance fields, and 3D Gaussians. Trained on 500K 3D objects with rectified flow transformers up to 2B params. Adopted by NVIDIA AI Blueprints (Sept 2025).
Distilling LLMs into Logical Personal Agents
Research sample code exploring how to build a single personal agent with natural-language interfaces. Distills LLMs into logical structures for Actions, Memory, and Plans; uses Structured RAG for agent memory with superior recall over Classic RAG.
Generalist Multi-Agent System
A generalist multi-agent system for complex web and file-based tasks. A lead Orchestrator coordinates specialized agents for web navigation, code execution, and file management, using a Task Ledger and Progress Ledger for adaptive planning.
Reflective MCTS for AI Agents
An approach for teaching AI agents to explore more effectively via Reflective MCTS and Exploratory Learning. Balances exploitation vs. exploration, adapts to changing environments, and improves decision-making via test-time compute scaling.
Multimodal Critical Thinking Agent
A multimodal critical thinking agent framework for complex visual reasoning on images and long-form videos. Decomposes questions, adapts strategies, and uses a built-in critic to self-verify answers. Outperforms MLLMs and tool-augmented pipelines across multiple benchmarks.
Retrosynthesis Reaction Prediction
Takes a target molecule (SMILES) and produces several potential chemical reactions to synthesize it. Combines predictions from diverse models via a learning-based ensembling strategy; PhD-level chemists preferred its predictions over the reactions it was trained on.
Target-Aware Drug Design
A transformer-based chemical language model for target-specific drug design. Generates novel compounds or optimizes existing molecules via target-aware fragments — accelerating discovery for infectious diseases such as tuberculosis.
Foundation Model of the Atmosphere
A large-scale foundation model for atmospheric forecasting trained on 1M+ hours of weather and climate simulations. Forecasts wind, temperature, and air quality at 0.1° resolution (~11 km), ~5,000× faster than conventional numerical weather prediction. Published in Nature.
Agentic Data Exploration & Visualization
Blends natural language and visual interfaces to help analysts explore and visualize data with AI agents. Agent mode for fully delegated exploration; data-threads for branching workflows; multi-modal chart builder combining drag-and-drop with NL.
Diffusion for Controllable Protein Generation
A diffusion framework for controllable protein generation in sequence space. Trained on evolutionary-scale data, EvoDiff-Seq and EvoDiff-MSA produce high-fidelity, diverse, and structurally plausible proteins — including disordered regions inaccessible to structure-based models.
Robotics VLA+ Model from Phi
The first robotics model derived from Microsoft's Phi series of vision-language models. Translates natural language into control signals for bimanual manipulation; adds tactile sensing and learns continually from human teleoperation feedback during deployment.