Language
| Models โ โ โ โ โ | Description โ โ โ โ โ | Features โ โ โ โ โ โ โ โ |
|---|---|---|
| MiniMax-M3 | Frontier multimodal coding model with 1M context window | โข Multimodal โข 1M context window โข Frontier coding |
| MiniMax-M2.7 | Beginning the journey of recursive self-improvement | โข Top real-world engineering โข Professional office delivery โข Character-rich interaction |
| MiniMax-M2.7-highspeed | Same performance as M2.7 โข Significantly faster inference | โข Polyglot code mastery โข Precision code refactoring โข Low latency |
Legacy Models
Legacy Models
| Models โ โ โ โ โ | Description โ โ โ โ โ | Features โ โ โ โ โ โ โ โ |
|---|---|---|
| MiniMax-M2.5 | โข Optimized for code generation and refactoring | โข Peak Performance. Ultimate Value. Master the Complex. |
| MiniMax-M2.5-highspeed | โข Same performance as M2.5 โข Significantly faster inference | โข Polyglot code mastery โข Precision code refactoring โข Low latency |
| MiniMax-M2.1 | โข 230B total parameters with 10B activated per inference โข Optimized for code generation and refactoring | โข Polyglot code mastery โข Precision code refactoring โข Enhanced reasoning |
| MiniMax-M2.1-highspeed | โข Same performance as M2.1 โข Significantly faster inference | โข Polyglot code mastery โข Precision code refactoring โข Low latency |
| MiniMax-M2 | โข Context Length: 200k tokens โข Maximum Output: 128k tokens (including CoT) | โข Agentic capabilities โข Function calling โข Advanced reasoning โข Real-time streaming |
Video
| Models โ โ โ โ โ โ | Description โ โ โ โ โ โ โ โ โ | Res.& Dur. โ โ โ โ | FPS โ โ โ โ |
|---|---|---|---|
| MiniMax H3 | Next-gen open general-purpose multimodal video model โข Text-to-Video / Image-to-Video / FirstโLast Frame / Multimodal reference | โข 768P / 2K โข 4โ15s | 24 fps |
Legacy Models
Legacy Models
| Models โ โ โ โ โ โ | Description โ โ โ โ โ โ โ โ โ | Res.& Dur. โ โ โ โ | FPS โ โ โ โ |
|---|---|---|---|
| MiniMax Hailuo 2.3 | โข Text to Video & Image to Video โข SOTA instruction following โข Extreme physics mastery | โข 1080p 6s โข 768p 6s, 10s | 24 fps |
| MiniMax Hailuo 2.3Fast | โข Image to Video โข Extreme physics mastery โข Value and Efficiency | โข 1080p 6s โข 768p 6s, 10s | 24 fps |
| MiniMax Hailuo 02 | โข Text to Video & Image to Video โข SOTA instruction following โข Extreme physics mastery | โข 1080p 6s โข 768p 6s, 10s โข 512p 6s, 10s | 24 fps |
Audio
| Models โ โ โ โ โ โ | Description โ โ โ โ โ | Features โ โ โ โ โ โ โ โ โ โ โ โ โ |
|---|---|---|
| speech-2.8-hd | โข Ultra-realistic quality featuring sound tags | โข 40 languages supported โข 7 emotions supported โข specified languages and dialects supported |
| speech-2.8-turbo | โข Seamless speed meets natural flow | โข 40 languages supported โข 7 emotions supported โข specified languages and dialects supported |
Legacy Models
Legacy Models
| Models โ โ โ โ โ โ | Description โ โ โ โ โ | Features โ โ โ โ โ โ โ โ โ โ โ โ โ |
|---|---|---|
| speech-2.6-hd | โข Ultimate Similarity โข Ultra-High Quality | โข 40 languages supported โข 7 emotions supported โข specified languages and dialects supported |
| speech-2.6-turbo | โข Ultimate Value โข Low latency | โข 40 languages supported โข 7 emotions supported โข specified languages and dialects supported |
| speech-02-hd | โข Stronger replication similarity โข High quality voice generation | โข 24 languages supported โข 7 emotions supported โข specified languages and dialects supported |
| speech-02-turbo | โข Superior rhythm and stability โข Low latency | โข 24 languages supported โข 7 emotions supported โข specified languages and dialects supported |
Music
Starting August 20, 2026, the paid APIs (Music Generation and Lyrics Generation) will no longer be available to new users; existing paying users can continue to use the current API services. The free music generation APIs (Music-3.0-free, Music-2.6-free, music-cover-free) will be discontinued.To experience or use music generation capabilities, please visit MiniMax Audio, or use the open-source MiniMax Music 3 model on Hugging Face.
| Models โ โ โ โ โ โ | Description โ โ โ โ โ | Features โ โ โ โ โ โ โ โ โ โ โ โ โ |
|---|---|---|
| music-3.0 | โข New Music Generation Capabilities | โข Intent Understood โข Sound Elevated โข Vocals Humanized |
| music-2.6 | โข Cover Reborn. Bass Redefined. | โข Cover Reborn. Bass Redefined. |
| music-cover | โข Generate cover versions from reference audio | โข One-step cover generation โข Two-step cover with lyrics modification โข Style transfer โข Auto lyrics extraction |
Legacy Models
Legacy Models
| Models โ โ โ โ โ โ | Description โ โ โ โ โ | Features โ โ โ โ โ โ โ โ โ โ โ โ โ |
|---|---|---|
| music-2.0 | โข Text to Music โข Enhanced musicality โข Natural vocals and smooth melodies | โข Human-like performance โข Riche emotional expression โข Enhanced tone control |