MAI-Image-2.5: Microsoft's Quiet Contender for Image Editing
Microsoft's MAI-Image-2.5 debuted #2 on Arena for image editing from inside PowerPoint and OneDrive. What it edits, and how to run it on Versely for 5 credits.
Microsoft doesn't usually get filed under "AI image lab." It ships models the way it ships everything else: quietly, inside a product you already have open. That's exactly what happened with MAI-Image-2.5. It debuted at No. 2 on Arena's image-editing leaderboard and No. 3 for text-to-image — a genuinely strong showing against labs that do nothing but ship image models — and almost no one in the creator world noticed, because there's no standalone "MAI" app to download. It lives inside PowerPoint and OneDrive.
That distribution choice is the whole story. A frontier-adjacent editing model sitting behind slide-deck software isn't going to trend on its own. But the model underneath is worth understanding on its own terms, and you don't need a Microsoft 365 seat to use it — Versely runs both the generation and editing variants directly.
Built by the MAI Superintelligence Team
MAI-Image-2.5 comes out of Microsoft's MAI Superintelligence Team, the group Microsoft stood up to build first-party foundation models rather than exclusively relying on partner models. Image generation was an obvious early target: PowerPoint alone generates an enormous volume of "make me a picture for this slide" requests, and OneDrive is where a huge share of the world's product photos, screenshots, and marketing stills already sit waiting to be cleaned up.
According to Microsoft AI's announcement, the model is live on PowerPoint for high-quality image generation and rolling out to OneDrive for precise editing. Microsoft also shipped a Flash variant alongside the full model, aimed at faster, lower-cost generation and editing for scaled, high-volume workloads — the same pattern OpenAI, Google, and Anthropic have all followed with their own "fast" model tiers.
One model, not two
The detail that matters most for how you'll actually use it: MAI-Image-2.5 is a unified generation-and-editing model, not a generator paired with a bolted-on inpainting tool. Microsoft describes it as supporting "precise, localized edits, from replacing an object or updating text to removing motion blur" using the same underlying architecture that produces the image in the first place.
That matters because localized editing is where most image models fall apart. Ask a generalist model to "just change the sign in the background" and you frequently get a whole new image back — different lighting, a slightly different subject pose, a color grade that no longer matches. A model built for inpainting from the ground up is supposed to touch only the region you asked about and leave everything else — composition, lighting, grain, color — untouched. MAI-Image-2.5's Arena placement is specifically strong on the editing side, which is the harder half of that promise to keep.
What "localized edit" covers
Concretely, the categories MAI-Image-2.5 handles are the ones that come up constantly in ad and content production:
- Object swaps. Replace one product, prop, or background element with another without regenerating the whole scene.
- In-image text updates. Change a price, a headline, or a label on packaging that's already baked into the pixels — notoriously one of the hardest things for diffusion models to do cleanly.
- Background changes. Swap or clean up what's behind the subject: remove a distracting object, change a location, or neutralize a busy backdrop.
- Blur and artifact removal. Fix motion blur or compression noise on an otherwise usable shot instead of re-shooting or re-generating from scratch.
None of this requires masking tools, layer stacks, or a design background. You describe the change in plain language and the model applies it to the region it concerns.
Running it on Versely without a Microsoft 365 seat
Versely carries both halves of the MAI-Image-2.5 pair as separate, a la carte models — you don't need PowerPoint, OneDrive, or any Microsoft subscription to use either one.
Mai Image 2.5 Text to Image (model page) generates the base image from a prompt, at 5 credits per image, with sharp text rendering as one of its stated strengths — useful if your first pass needs a headline or label baked in correctly the first time.
Mai Image 2.5 Edit (model page) takes an existing image plus an edit instruction and returns the localized result, also at 5 credits.
A practical two-step workflow inside the AI photo editor:
- Generate the base shot with Mai Image 2.5 Text to Image: "Product photo of a matte-black water bottle on a marble countertop, soft studio lighting, minimal background, a small corner tag that reads 'SAMPLE'."
- Feed the result into Mai Image 2.5 Edit with a targeted instruction: "Change the corner tag text to read 'NEW' and swap the marble countertop for a light oak wood surface. Keep the bottle, lighting, and shadows unchanged."
Because it's the same underlying model handling both steps, the edit tends to preserve the lighting and shadow logic from the original generation rather than reinventing them — which is the entire point of a unified model over a generate-then-patch pipeline.
Where the same discipline carries into video
The instinct behind MAI-Image-2.5 — touch only what needs to change, leave the rest alone — is exactly what you want when the same kind of correction needs to happen on a moving frame instead of a still one. If you're cleaning up in-image text on a static ad, text-to-image plus Mai Image 2.5 Edit gets you there. If the text needs to live on top of a video instead — a caption, a price callout, a CTA card — that's a separate step handled with text overlays on video, since MAI-Image-2.5 itself only operates on still images.
FAQ
Do I need a Microsoft 365 or Copilot subscription to use MAI-Image-2.5? No. Microsoft's own rollout puts it inside PowerPoint and OneDrive, but Versely exposes the underlying Text to Image and Edit models directly, billed per call in credits.
What's the difference between the two Versely listings? Mai Image 2.5 Text to Image generates from a text prompt with no input image required. Mai Image 2.5 Edit requires an existing image and an instruction describing the localized change.
Is it good at in-image text specifically? Microsoft highlights sharp text rendering as a strength of the text-to-image variant, and in-image text updates as one of the editing model's named capabilities — text handling appears to be a deliberate focus area rather than an afterthought.
What is the Flash variant, and does Versely carry it? Flash is Microsoft's faster, lower-cost tier of the same model family, aimed at scaled workloads. Versely's current listings are the standard Text to Image and Edit models; the lineup expands as new variants roll out.
How many edits can I chain? There's no hard limit from Versely's side — each edit call takes whatever image you feed it, including the output of a previous edit — but compounding edits can drift from the original the more passes you run. Keep instructions specific and check the result after every one or two edits rather than chaining five in a row blind.
Closing takeaway
MAI-Image-2.5 is a legitimate top-three-on-Arena editing model that most creators will never encounter through Microsoft's own surfaces, because PowerPoint and OneDrive aren't where creative work happens for most people outside enterprise offices. That's an availability gap, not a quality gap. Versely closes it: generate with Mai Image 2.5 Text to Image, refine with Mai Image 2.5 Edit, and you get the same unified-model editing discipline Microsoft built for slide decks, pointed at whatever you're actually making.