Biology doesn't operate in silos, and neither should the AI representation of it. Quine is an early-stage research effort to create a multimodal world model of biology. By connecting insights across biological scales and modalities, Quine helps scientists computationally search a space far larger than intuition allows and prioritize hypotheses before they reach the lab. Experimental results provide important feedback, helping researchers sharpen future research directions. The post Introducing Quine: An AI research system designed for the complexity of biology appeared first on Microsoft Research.
Since launching a year ago, the Microsoft Research Asia — Singapore lab has established a strong foundation, deepened collaboration across government, academia, and industry, and explored how frontier AI research can create real-world value. The post One year in: How Microsoft Research Asia – Singapore is advancing research, partnership and talent for real-world impact appeared first on Microsoft Research.
Robots are getting smarter, but how can their hardware match that growth? New Microsoft Research findings show that moving AI inference beyond the robot can improve task success, boost efficiency, and support more advanced physical AI workloads. The post Offloaded inference for real-world physical AI robotics appeared first on Microsoft Research.
Custom-made molecules are advancing medicine, materials, and agriculture, but producing them is slow and expensive. A new Nature paper highlights RetroChimera, a predictive model that helps accelerate chemical synthesis, helping researchers explore a wide range of molecules. The post Improving synthesis prediction of small molecules at scale with RetroChimera appeared first on Microsoft Research.
Catalyst Lab director Chris White talks with Weishung Liu about his path from DARPA fieldwork in Afghanistan and dark-web search tools against human trafficking to leading public-good technology research at Microsoft, and the humility and resilience that guide him.
What if pathology foundation models could do more with less? GigaPath-Flash and GigaTIME-Flash cut computational demands while maintaining strong performance, opening the door to larger studies and broader exploration. The post GigaPath-Flash and GigaTIME-Flash: Toward population-scale discovery with efficient pathology foundation models appeared first on Microsoft Research.
Skala 1.1, the updated deep-learning exchange-correlation functional from Microsoft Research, provides greater accuracy, expanded accessibility across the computational chemistry ecosystem, and a living benchmark to track computational performance. The post Broadening access to Skala creates a faster path to predictive DFT appeared first on Microsoft Research.
A path, a fence, a knot. MindTopo sets a new benchmark for testing how AI understands topological relationships and highlights new opportunities to strengthen spatial reasoning and planning. The post MindTopo reveals VLMs’ spatial reasoning abilities appeared first on Microsoft Research.
Radiology AI is evolving beyond report generation. CARE-X explores a unified approach that combines flexible reasoning, calibrated predictions, and measurement-based tools for chest X-ray interpretation. The post Introducing CARE-X: Towards Clinically Useful Radiology VLMs with Auxiliary Supervision, Reward-Aligned Learning, and Tool-Augmented Measurement appeared first on Microsoft Research.
Orchard is an open-source framework for the research community to train and evaluate AI agents across task types. It reduces complexity while supporting strong performance from smaller models by enabling researchers to reuse the same infrastructure. The post Orchard: An open framework for scalable agentic AI appeared first on Microsoft Research.
Computer-use AI agents struggle with multi-step workflows like email and customer support. Echoverse trains agents in realistic environments rather than simply providing more training tasks, helping them improve as the tasks, tests, and environments evolve.
EvoLib is a new framework that enables large language models to learn from their own experience during inference by transforming past attempts into reusable skills and reflective insights, continuously refining and consolidating them into increasingly general and effective knowledge.
Microsoft's SymCrypt team announces a new methodology to formally verify Rust-written cryptographic code using the Lean proof assistant and the Aeneas toolchain, achieving functional correctness against formal specifications derived from standards. The approach has been applied to post-quantum algorithms like ML-KEM and SHA-3, with verified code already shipping in Windows insider builds. The methodology scales by using AI agents to automate proof writing while keeping human oversight on standard formalization. It also handles platform-specific intrinsics and multiple architectures without sacrificing performance.
Aurora 1.5 adds 22 more variables, hourly temporal resolution, and probabilistic ensemble forecasting to the Aurora foundation model, making it more useful for real-world weather, climate, and energy applications. Released as open source, it enables researchers and developers to use, evaluate, and build on the model.
Flint is an open-source visualization intermediate language from Microsoft Research, designed to help AI agents create expressive, polished charts from compact, human-editable specifications. It handles low-level design details automatically via semantic types, supports multiple rendering backends, and powers the Data Formulator project.
AI agents often fail because their instructions, or skills, are manually modified with no guarantee of improvement. SkillOpt turns skill editing into a training process, making agent behavior more reliable without changing model weights. Across 52 evaluation cells, SkillOpt achieves best or tied-best results, and the optimized skills remain compact, auditable, and transferable.
AI agents suffer from statelessness, requiring constant context reloading. Memora introduces a scalable memory system decoupling storage from retrieval, achieving state-of-the-art on long-context benchmarks while using up to 98% fewer tokens.
Researchers introduce generative causal testing, which translates black box models into clear hypotheses and verifies them in the scanner, revealing what specific brain regions respond to in language.
Project Ire, Microsoft's autonomous malware-classification agent, reverse-engineered a LOTUSLITE variant that went undetected by most major EDR tools. Through behavioral analysis rather than signature matching, Ire identified the sample's malicious intent and produced a detailed function-level report consistent with Acronis's published analysis.
Data Formulator 0.7 is an open-source AI-powered system for enterprise data analytics that combines data connectivity, agent-guided exploration, and visualization refinement in a shared workspace.
Modern AI systems are powerful not because they replicate human intelligence, but because they extend structures already present in human cognition and language. This perspective explains AI's capabilities and limitations, and reframes AI safety as a system-level challenge requiring engineering and governance, not fear of rogue AI.
Microsoft Research releases MagenticLite, an agentic application designed for small models, along with MagenticBrain orchestrator and Fara1.5 computer-use model. The system works across browser and local file system, achieving state-of-the-art results on web navigation tasks while keeping data on-device.
Vega is a new zero-knowledge proof system from Microsoft Research that enables users to prove facts from government-issued credentials without revealing the credential itself. It achieves under 100ms proving time on commodity devices using folding schemes, and is designed for real-world digital identity formats like mobile driver's licenses and the EU Digital Identity Wallet.
Microsoft Research clarifies the scope of its paper on AI delegation, noting that while models show fidelity degradation in long-horizon tasks, production systems mitigate these effects, and the benchmark is a diagnostic tool for future improvement.
mimalloc is an open-source, modern, scalable memory allocator that is a drop-in replacement for malloc and free. It is relatively small (~12K lines), with clear internal data structures, and is easy to build and integrate into other projects. It provides bounded worst-case allocation times (up to OS primitives), bounded space overhead, low internal fragmentation, and minimal contention by relying almost exclusively on atomic operations.
Microsoft releases a lightweight foundation model that can predict AC optimal power flow in milliseconds, boosting efficiency and unlocking cost savings in grid analysis.
Microsoft Research introduces SocialReasoning-Bench, a benchmark evaluating AI agents' social reasoning in principal-agent settings. Tests show frontier models complete tasks but often fail to secure optimal outcomes for users, even with explicit instructions. The benchmark measures outcome optimality and due diligence to assess agents' ability to act in users' best interests.
Microsoft Research releases an open dataset of U.S. power grid transmission topology derived from public data, enabling AC optimal power flow analysis and addressing research challenges due to restricted grid data. The pipeline uses OpenStreetMap and public energy data to create geographically grounded models that are solvable for power flow analysis, demonstrated across 48 states and the Eastern Interconnection. The dataset supports studies of congestion, transmission expansion, and demand siting.
Microsoft researchers share advances in building and operating large-scale distributed systems, spanning datacenters, networking, and the growing intersection with AI during NSDI '26.
Microsoft Research red-teamed a live platform of over 100 AI agents, identifying network-level risks that only appear through agent interactions, including self-propagating worms, reputation manipulation, manufactured consensus, and proxy chains. These risks cannot be reproduced by testing agents in isolation. The study also observed emergent security behaviors in a small fraction of agents, reducing attack success. Findings suggest the need for layered defenses across platform, agent, and model layers.