Modelmultimodal2mo ago

MiniMax M3 review

MiniMax's third-generation flagship — M3 is an open-weight 428B-parameter MoE (~23B active per token) using MiniMax Sparse Attention (MSA), with a 1M-token context window and native multimodality (text, image, and video input). It's positioned as the first open-weight model to combine frontier coding, a 1M context, and native multimodality — and can even operate a desktop computer.

By Ravi Menon · Models & Benchmarks EditorVerified 2026-07-08
Maker
MiniMax
Launched
Jun 1, 2026
Pricing
paid
Visit official site
Firstlook

Our verdict

MiniMax M3 is the clearest statement yet from a team that has shipped quietly and well. It's an open-weight 428B MoE (~23B active) with MiniMax Sparse Attention, a 1M-token context, and native multimodality across text, image, and video — and MiniMax pitches it as the first open-weight model to combine frontier coding, a 1M context, and native multimodality in one release. A 59.0% SWE-Bench Pro score at $0.30/$1.20 per 1M tokens makes it economically interesting, and the fact that it can operate a desktop computer pushes it into agentic territory. This is a first look from the launch materials, so no hands-on score yet — but the combination of open weights, long context, and low price makes it worth testing directly.

First look — our read from the docs and sources below; not yet hands-on tested.

MiniMax has been easy to overlook. They don't stage viral launch spectacles or get the breathless coverage that OpenAI or Anthropic do. What they've consistently done instead is ship capable models that land on Hugging Face, benchmark well, and get adopted quietly by builders who find them while comparing options.

MiniMax M3 breaks that quiet. Released June 1, 2026, it's the company's third M-series model, and it's a genuine flagship: an open-weight 428B-parameter mixture-of-experts model with roughly 23B active parameters per token, using MiniMax Sparse Attention (MSA), a 1M-token context window, and native multimodality across text, image, and video. MiniMax frames it as the first open-weight model to bring frontier coding, a 1M context, and native multimodality together in one release — and it can even operate a desktop computer.

What "natively multimodal" means here

M3 takes text, images, and video as combined input and produces text output. You can show it a photo, a chart, a screenshot, a document page, or a video clip alongside a prompt, and it answers in text. The output is text — it doesn't generate images or video — but the input surface is wide.

That surface covers the practical work you'd normally reach a vision model for: document parsing, visual Q&A, chart reading, screenshot debugging, and now video understanding — extended by a 1M-token context so you can put a lot in front of it at once.

Coding, agentics, and price

Two things push M3 past "another multimodal model." First, coding: MiniMax reports 59.0% on SWE-Bench Pro, and the model is open-weight, so you can self-host or use the API. Second, agentics: M3 can operate a desktop computer, which puts it squarely in computer-use and tool-driven-agent territory.

The economics are notable too — API access runs $0.30 per 1M input tokens and $1.20 per 1M output tokens, low for a frontier-tier model. Between the open weights, the 1M context, and that price, M3 is worth testing directly rather than waiting for a news cycle.

MiniMax's positioning in 2026

The company started primarily as a text model provider and has steadily layered in capabilities. M3 is its clearest step into the frontier tier where GPT-4o and Gemini 2.5 Flash set the standard — with the added twist of open weights and a stated first-mover claim on combining coding, long context, and native multimodality.

What we haven't tested yet

This is a first look from the launch materials and the official announcement — no hands-on evaluation, so no score. The reported 59.0% SWE-Bench Pro figure, multimodal quality on real documents and video, and behavior on long-context and agentic desktop tasks are all things we'll want to run ourselves. But MiniMax M3 is the kind of real, open-weight release that tends to matter more than its launch buzz suggests.

Provider

Specs & key facts

What it isOpen-weight 428B-parameter MoE, ~23B active params per token[src]
AttentionMiniMax Sparse Attention (MSA)[src]
Context window1M tokens[src]
Input modalitiesText, image, and video (natively multimodal)[src]
Agentic capabilityCan operate a desktop computer[src]
Coding59.0% on SWE-Bench Pro[src]
LicenseOpen weights[src]
API pricing$0.30 / $1.20 per 1M tokens (input/output)[src]

Capabilities

Frontier codingYes (59.0% SWE-Bench Pro)
Native multimodalityYes (text, image, video input)
1M-token contextYes
Operate a desktop computerYes
Open weightsYes
Image generationNo (multimodal input, text output)

How to use it

  1. 1Grab the open weights from Hugging Face (huggingface.co/MiniMaxAI/MiniMax-M3) for self-hosting, or use the hosted API.
  2. 2M3 takes text, image, and video as input — reach for it on multimodal reasoning, document/chart/screenshot parsing, and long-context tasks up to 1M tokens.
  3. 3Try the agentic side: M3 can operate a desktop computer, so it fits computer-use and tool-driven workflows.
  4. 4For coding, benchmark it against your current model — M3 reports 59.0% on SWE-Bench Pro at $0.30/$1.20 per 1M tokens.

Pricing

API

$0.30 / $1.20 per 1M tokens

$0.30 per 1M input tokens, $1.20 per 1M output tokens. Weights are also open, so self-hosting is an option.

API pricing is $0.30 input / $1.20 output per 1M tokens. Because M3 is open-weight, it can also be self-hosted. Verified 2026-07-08.

Pros & cons

Pros

  • Open weights — self-host it, or use the API; you're not locked in.
  • Rare combination in one model: frontier coding, 1M-token context, and native multimodality (text/image/video).
  • Low API price ($0.30 / $1.20 per 1M tokens) for a frontier-tier model, plus agentic desktop-computer operation.

Cons

  • Multimodal input, text output — it doesn't generate images or video.
  • 428B total parameters means self-hosting the open weights needs substantial GPU memory despite the sparse ~23B active path.
  • First look only — no independent hands-on evaluation of the reported 59.0% SWE-Bench Pro number yet.

Alternatives

FAQ

Sources

  1. 1.MiniMax M3 — open-weight 428B MoE, MSA sparse attention, 1M context, native multimodality, frontier coding, desktop-computer operation; pricing and benchmarkshttps://www.minimax.io/blog/minimax-m3Verified 2026-07-08
  2. 2.Open weights available on Hugging Facehttps://huggingface.co/MiniMaxAI/MiniMax-M3Verified 2026-07-08

More coverage

News & first-looks about this release. Coming soon.
Head-to-head comparisons. Coming soon.