
@mustafasuleyman
CEO, @MicrosoftAI | Author: The Coming Wave | Past: Co-founder, @InflectionAI & @GoogleDeepMind
Many important and bold ideas in @BillGates's new essay on the impact of AI need to be discussed... Token Tax and a protected 'nature reserve' for certain jobs, are just two of them. gatesnotes.com/home/home-page…
MAI-Image-2.6 is the best image editing model on AA. We now hold 3 of the top 5 leaderboard positions. Try it today here: playground.microsoft.ai/chat
Artificial Analysis@ArtificialAnlys·Microsoft's MAI-Image-2.6-Preview lands at #1 on the Artificial Analysis Image Editing Leaderboard and takes #2 in Text to Image MAI-Image-2.6 is the newest model in Microsoft AI's MAI-Image family, announced August 10. Microsoft highlights stronger text rendering, better portraits and 3D imagery, and more polished commercial and photorealistic outputs. Like the MAI-Image-2.5 family, it handles both text to image generation and image editing. In the Artificial Analysis Image Arena, MAI-Image-2.6-Preview debuts at #1 on the Image Editing Leaderboard, ahead of Microsoft's own MAI-Image-2.5-Pro, Reve 2.1, and OpenAI's GPT Image 2, and giving Microsoft the top two spots on the board. In Text to Image it takes #2, behind only OpenAI's GPT Image 2 and ahead of Reve 2.1. On our refreshed Text to Image taxonomy, MAI-Image-2.6-Preview takes the top spot on 5 of the 19 category leaderboards (Material, Knowledge, Frontier, Retail & Ecommerce, and Marketing & Advertising), ahead of GPT Image 2, which leads every other category. MAI-Image-2.6 extends a rapid run of strong image releases from Microsoft AI. MAI-Image-2.5 debuted at #2 in Text to Image in June. MAI-Image-2.5-Pro, launched July 23, took #1 in Image Editing when we published our results last week. MAI-Image-2.6 now takes that top spot from its own sibling, and sits at #2 in Text to Image against MAI-Image-2.5-Pro's #8. MAI-Image-2.6 is available in the MAI Playground and in Private Preview on Microsoft Foundry. Congratulations to @MicrosoftAI on the release! See below for our analysis of MAI-Image-2.6-Preview and other leading models in the Artificial Analysis Image Arena 🧵
One of the most important and inspiring quotes of all time: “The concern for man and his destiny must always be the chief interest of all technical effort… in order that the creations of our mind shall be a blessing and not a curse to mankind.” (Albert Einstein, 1931)
MAI-Image-2.5 is now #1 on the Artificial Analysis leaderboard for image editing! Amazing hillclimbing... compounding the gains!
Artificial Analysis@ArtificialAnlys·Microsoft's MAI-Image-2.5-Pro debuts at #1 on the Artificial Analysis Image Editing Leaderboard, and takes the #7 spot in Text to Image MAI-Image-2.5-Pro is Microsoft AI's quality-focused image model, launched July 23 in preview on Microsoft Foundry. It joins MAI-Image-2.5 and MAI-Image-2.5-Flash in a family Microsoft positions as covering the quality-speed-cost curve, so builders can pick the point that fits their job. Pro sits at the quality end: Microsoft describes it as its highest-fidelity image model to date, aimed at hero imagery, detailed editing, and precise in-image text rendering. In the Artificial Analysis Image Arena, MAI-Image-2.5-Pro debuts at #1 on the Image Editing Leaderboard, surpassing Reve 2.1, GPT Image 2, and Microsoft's own MAI-Image-2.5. In Text to Image it lands at #7. On Microsoft Foundry, MAI-Image-2.5-Pro is priced per token: $5 per 1M text input tokens, $8 per 1M image input tokens, and $106 per 1M image output tokens, which works out to roughly $108.5 per 1k 1024x1024 images. That compares to $48 per 1k images for MAI-Image-2.5 and $20 per 1k for MAI-Image-2.5-Flash. MAI-Image-2.5-Pro is available in preview on Microsoft Foundry across seven global-standard regions, and can be tried out in the MAI Playground. Congratulations to @MicrosoftAI on the release! See below for comparisons between MAI-Image-2.5-Pro and other leading models in the Artificial Analysis Image Arena 🧵
We hit No.3 on the Image-editing benchmark! ... beating Google’s Nano Banana and Meta’s Muse Image models.
Arena.ai@arena·MAI-Image-2.6-Preview by @MicrosoftAI has landed at #3 in Single Image Edit with 1,420 pts! This is +19 pts above MAI-Image-2.5, moving from #5 to #3. It sits 19 pts behind preliminary Grok Imagine Image 2.0 (low) at #2 and 43 pts behind GPT Image 2 (medium) at #1. By category, MAI-Image-2.6-Preview is #3 in across: - Product, Branding & Commercial Design - 3D Imaging & Modeling - Cartoon, Anime & Fantasy - Photorealistic & Cinematic Imagery #4 in Portraits and #8 in Text Rendering. The biggest jumps over MAI-Image-2.5 are in Text Rendering (+43 pts) and Product, Branding & Commercial Design (+38 pts). Congrats again to the @MicrosoftAI team on this release!
Last week we tested the model on the Text-to-Image benchmark and it ranked 2nd This week we ran the same model on the image editing leaderboard and score 3rd place MAI-Image-2.6 is now officially an all-round high performer for any image generation tasks.
Our first reasoning model, MAI-Thinking-1, is built from scratch. Now available in Microsoft Foundry. Kudos to the team! More below.
Try it here: aka.ms/mai-thinking-1…
Our latest code model is 25% more efficient, higher quality, and a quarter of the cost than our model launched in June. Now live in GitHub Copilot - try it out.
More details in the blog: microsoft.ai/news/mai-code-…
MAI-Image-2.6 is now the #2 text-to-image model in the world - beating out Nano Banana, Meta, and Grok! Fantastic moment for the team who've been hill climbing relentlessly. Try it out now on Arena! x.com/arena/status/2…
Big news! Our new MAI-Cyber-1-Flash model combined with MDASH, our multi agent security harness, delivers 96% on the CyberGym benchmark, 12pts above Mythos, at HALF the cost. Proud of the team. More details in THREAD:
Security is an always-on mission, and given the enormous volume of inbound attacks, token cost is now the real constraint for defenders. We need continuous, real-time, cost-efficient agents to protect us.
MAI-Image-2.5-Pro launches today in Foundry for preview. It's our highest-fidelity, professional-grade image model. It’s for super high quality imagery, detailed editing, precise in-image text rendering. It joins our image family of models so builders can pick the point on the quality-speed-cost curve that fits their job. More here: t.co/JVU42BGuQ7
MAI-Voice-2-Flash launches today! Flash is 2x faster than MAI-Voice-2 and 32% cheaper, at $15 per 1M characters. MAI-Voice-2-Flash is also in public preview and powers Dynamics 365 Contact Center, our enterprise platform for call center agents, and reduces GPU costs up to 89%. More here... t.co/JVU42BGuQ7
As @satyanadella says, we're making great progress on shipping MAI models that are higher quality, faster and cheaper across MSFT. In PowerPoint, our image model cut costs 84% compared with GPT-Image-2. In OneDrive, it lifted save rates 26% and cut latency by ~25%. It's also the default model in Bing delivering great quality and performance. In Dragon Copilot, our transcription model now covers 58 languages and halves the error rate on multilingual clinical transcription. And of course the flights we have in motion on GitHub and Excel will no doubt be huge too! More details in the blog here: t.co/JVU42BGuQ7
New paper published in Nature Health today from the Microsoft AI Futures team. After reviewing 1.7m conversations across 109 countries, it’s clear that Copilot is an invaluable resource particularly for people with low confidence in their health systems. We've always known that technology is a great equalizer, driving broader access to better quality services. This is further evidence that AI is giving underserved people around the world access to invaluable support when they need it most. Many thanks to all the authors. Check out the paper here: t.co/fNfACzWGDJ
Fully support this important proposal from @demishassabis. The time for us all to act is now. "...we must use this precious window before AGI arrives to shape this technology for the benefit of all humanity." x.com/demishassabis/…
I made a similar proposal for an "IPCC for AI" with @ericschmidt in 2023 ft.com/content/d84e91…
Introducing Ode Poetry. Ode is a wonderful poetry pharmacy that reads you a poem for the moment you’re in. Just tell Ode what you're feeling, and it uses Microsoft AI audio models to connect you with the same work that poetry expert William Sieghart would recommend. The best technology doesn't replace human creativity, it helps more people experience it. Super proud of the team for making this truly humanist tool. More in the blog: t.co/oCDTiIdXxW
Try it here: odepoetry.ai
Healthcare is the most important application of AI: making humans happier + healthier. Our recently published paper "Public use of a generalist LLM chatbot for health queries" made the front cover of Nature Health. So proud of the team! nature.com/articles/s4436…
Shaping our culture at Microsoft AI is one of the most important responsibilities I have. Keeping our team lean and talent dense is critical to our success. It's something I've thought very carefully about over the years. I thought I'd share a few of the principles we ask everyone in the team to sign up to. Everything below flows from one conviction: a disciplined, evidence-based, careful methodology compounds faster than heroic and chaotic improvisation. We don't always get it right but this is what we strive for: - Scientific rigor above all else. We set hypotheses, rigorously ablate, and make data-driven decisions. - Constantly think simple. Simple methods scale best. No recipe changes unless deeply justified. - Know your data. Data is our lifeblood. No data black boxes. Every person is responsible for every token they add to the model. - Know your evals. No narratives without numbers. Production evals and trust internal metrics, over academic benchmarks. - Don't celebrate results prematurely. Maintain healthy skepticism. Check for reward hacking. Never cherry pick results. - Always document everything. Positive and negative results are equally critical. Label every plot and axes. Summarize hypotheses and conclusions. Don't use jargon. - Be precise and use neutral language. Describe situations accurately without unnecessary emotional charge. - Retrospectives drive everything. The culture and process flywheel is critical to our hill climbing machine. We constantly run Retros to iterate and improve. - We’re an IC-first team. Management is a service, not the goal. We’re here to empower, unblock and accelerate the exceptional work of our world class ICs. - User focus. We build our models for end users. Developing user empathy begins with us. We always strive to use our own models first so we can hill climb for our users. - Take ownership for execution. Report issues, provide logs for debugging. First try to fix things yourself. And see it right through to completion. The quality of our thinking determines the quality of our models. We're hiring. Check out open roles here: t.co/4NerpgWhcl
Shaping our culture at Microsoft AI is one of the most important responsibilities I have. Keeping our team lean and talent dense is critical to our success. It's something I've thought very carefully about over the years. I thought I'd share a few of the principles we ask everyone in the team to sign up to. Everything below flows from one conviction: a disciplined, evidence-based, careful methodology compounds faster than heroic and chaotic improvisation. We don't always get it right but this is what we strive for: - Scientific rigor above all else. We set hypotheses, rigorously ablate, and make data-driven decisions. - Constantly think simple. Simple methods scale best. No recipe changes unless deeply justified. - Know your data. Data is our lifeblood. No data black boxes. Every person is responsible for every token they add to the model. - Know your evals. No narratives without numbers. Production evals and trust internal metrics, over academic benchmarks. - Don't celebrate results prematurely. Maintain healthy skepticism. Check for reward hacking. Never cherry pick results. - Always document everything. Positive and negative results are equally critical. Label every plot and axes. Summarize hypotheses and conclusions. Don't use jargon. - Be precise and use neutral language. Describe situations accurately without unnecessary emotional charge. - Retrospectives drive everything. The culture and process flywheel is critical to our hill climbing machine. We constantly run Retros to iterate and improve. - We’re an IC-first team. Management is a service, not the goal. We’re here to empower, unblock and accelerate the exceptional work of our world class ICs. - User focus. We build our models for end users. Developing user empathy begins with us. We always strive to use our own models first so we can hill climb for our users. - Take ownership for execution. Report issues, provide logs for debugging. First try to fix things yourself. And see it right through to completion. The quality of our thinking determines the quality of our models. We're hiring. Check out open roles here: t.co/p6vJgYFxMt
Shaping our culture at Microsoft AI is one of the most important responsibilities I have. Keeping our team lean and talent dense is critical to our success. It's something I've thought very carefully about over the years. I thought I'd share a few of the principles we ask everyone in the team to sign up to. Everything below flows from one conviction: a disciplined, evidence-based, careful methodology compounds faster than heroic and chaotic improvisation. We don't always get it right but this is what we strive for: - Scientific rigor above all else. We set hypotheses, rigorously ablate, and make data-driven decisions. - Constantly think simple. Simple methods scale best. No recipe changes unless deeply justified. - Know your data. Data is our lifeblood. No data black boxes. Every person is responsible for every token they add to the model. - Know your evals. No narratives without numbers. Production evals and trust internal metrics, over academic benchmarks. - Don't celebrate results prematurely. Maintain healthy skepticism. Check for reward hacking. Never cherry pick results. - Always document everything. Positive and negative results are equally critical. Label every plot and axes. Summarize hypotheses and conclusions. Don't use jargon. - Be precise and use neutral language. Describe situations accurately without unnecessary emotional charge. - Retrospectives drive everything. The culture and process flywheel is critical to our hill climbing machine. We constantly run Retros to iterate and improve. - We’re an IC-first team. Management is a service, not the goal. We’re here to empower, unblock and accelerate the exceptional work of our world class ICs. - User focus. We build our models for end users. Developing user empathy begins with us. We always strive to use our own models first so we can hill climb for our users. - Take ownership for execution. Report issues, provide logs for debugging. First try to fix things yourself. And see it right through to completion. The quality of our thinking determines the quality of our models. We're hiring. Check out open roles here: t.co/4NerpgVJmN
MAI-image-2.5 is new at #2 for text-to-image and #3 for image editing on @ArtificialAnlys! We’re second now only to GPT models. Also, MAI-Image-2.5-Flash is world-beating for quality/price. Super proud of the team. Now available through the Foundry API and rolling out across OneDrive + PowerPoint, or check it out in the MAI Playground: t.co/UA4NL6O1Oy
Great result for the team! x.com/ArtificialAnly…
Great to have our Superintelligence team meetup in Boston last week. Building AI takes a real team of passionate and driven folks. There's no substitute for getting face time with everyone all at once… the conversations are richer, the collaborations are stronger, and the progress compounds. Now back to it… in between World Cup matches
"In the application of AI, healthcare is going to be the next big product-market-fit explosion." More on the future of healthcare and our collaboration with the Mayo Clinic in my conversation with @CoreyNoles here: theneurondaily.com/p/watch-sleepi…
Talent density is incredibly important for building humanist superintelligence, and our team reflects that. Meet some of the humans at @MicrosoftAI who make our work so special
Great post by @satyanadella summarizing how we see this historic platform shift benefiting everyone broadly. @MicrosoftAI x.com/satyanadella/s…
We’ve been working on voice models that feel genuinely expressive. Curious what you think! Try the latest out in the MAI playground: playground.microsoft.ai
Agreed. We have to be very careful about this. I published an article in @Nature recently making similar arguments. mustafa-suleyman.ai/we-mustnt-let-… x.com/harari_yuval/s…
Had a great chat with Nilay Patel / Decoder about 7 new models we launched last week. Covered lots of important topics, from growing social angst, to whether AI is delivering enough value. Check it out: pod.link/decoder
.@ArtificialAnalysis’ graph shows that MAI-Transcribe-1.5 is in a league of its own
.@ArtificialAnalysis’ graph shows that MAI-Transcribe-1 is in a league of its own
Build was super fun! Here's a video recap of my presentation. Check it out: youtube.com/watch?v=OvLIae…
There are no shortcuts to the frontier. Disciplined, patient, meticulous attention to detail is critical. To give everyone a good sense of our progress we've published a very detailed technical report (109 pages!) outlining how we trained MAI-Thinking-1 and what we learned along
It’s time to move from renting intelligence to truly controlling your AI. Microsoft Frontier Tuning lets you take our models and make them uniquely your own, turning them from capable generalists to completely custom partners. It starts with reinforcement learning environments
So proud of the team today. Six months of super intense and outstanding work. I was honored to stand up and rep the work of the @MicrosoftAI lab at Build. Tons of technical detail we couldn't fit in the keynote, so we put it in a 109-page paper instead: microsoft.ai/wp-content/upl…
Today’s news all comes down to this: we’re putting our relentless hill-climbing machine at your service. From launching top tier models to helping you make them your own, our commitment as a platform company is to keep you at the absolute frontier. For all the details:
Proud that we’re collaborating with Mayo Clinic to build a frontier AI model for healthcare. Both our organizations exist to serve people at scale – and we believe this could be nothing short of transformative for global healthcare. news.microsoft.com/source/2026/06…
Awesome to have on you stage @steipete ! What an exciting time x.com/steipete/statu…
Super excited to announce seven new world-class MAI models today. They represent what we consider a new era in AI designed to keep you in control and on the frontier. First is our text foundation model, MAI-Thinking-1, exceptionally strong on reasoning and SWE tasks. - It’s a
Here we go! Tune in for the #MicrosoftBuild keynote starting now. I’m biased but you won’t want to miss it… youtube.com/watch?v=FFMm45…
One last run-through before Build tomorrow when I can finally share what we’ve been up to in the lab. Livestream starts at 9:30 am PT, and you can register or tune in here: build.microsoft.com/en-US/sessions…
Meet MAI-Image-2.5 - ranked third on the @arena text-to-image leaderboard. It's another great advance in quality. And with Build just a week away, there's much more to come from the @MicrosoftAI team. I can't wait.
Pleased to report that the model gets it right where it really matters: strong visual reasoning across objects, scene structure, lighting, scale, and spatial relationships, helping turn simple directions into polished images.
Little thought experiment to put AI chip improvements in perspective: Imagine that every person on Earth uses a calculator to perform one calculation per second. Everyone works 24 hours a day without rest. Every second, we all hit equals on the calculator for a long digit
Since I began work on AI in 2010, training compute for frontier models has grown by one trillion times. Now we're looking at something like another thousand-fold growth in effective compute by the end of 2028. 1000x the existing 1,000,000,000,000x. Extraordinary stuff.
Our paper landed in Nature Health today! Healthcare is one of the most high-stakes, high-potential applications of AI. So we set out to understand how people actually use it in our AI products today. nature.com/articles/s4436…
One insight that struck me: a lot of questions are actually about friends and loved ones. About 1 in 7 questions about symptoms and conditions are asked for someone else, like a child, aging parent, or partner.
Two models, two different parts of the creative process. MAI-Image-2-Efficient is a production workhorse. Volume, speed, tight cost control for iterative workflows. MAI-Image-2 is a precision tool. Highest fidelity, final deliverables, exact details, longer/more complex text.
Live on Microsoft Foundry + MAI Playground now: microsoft.ai/news/mai-image…
And you can try it now on MAI Playground too. Know some of you have hit regional/country restrictions - the team is working hard to bring Playground to more areas. Stay tuned! playground.microsoft.ai/chat
Meet MAI-Image-2-Efficient. Production-ready quality, 22% faster, and 4x more efficient than MAI-Image-2. Priced almost 41% lower too. Plus 40% average lower latency than other leading models. Live now in Microsoft Foundry + MAI Playground. microsoft.ai/news/mai-image…