MiniMax-M3.1-Flash-Preview is available only through M Plan and MiniMax Code for now. Get your Subscription Key.
Model Overview
MiniMax offers multiple LLMs to meet different scenario requirements. MiniMax-M3.1-Flash-Preview is the latest M-series language model for agentic reasoning, tool use, coding, and long-context tasks, and it supports tuning thinking depth witheffort. MiniMax-M3, MiniMax-M2.7, and MiniMax-M2.7-highspeed also remain available. Earlier models are listed under Legacy Models below.
Supported Models
Legacy Models
Legacy Models
For details on how tps (Tokens Per Second) is calculated, please refer to FAQ > About APIs.
MiniMax M3.1-Flash-Preview Key Highlights
1M-token context
1M-token context
MiniMax-M3.1-Flash-Preview supports up to a 1,000,000-token context window for long documents, codebases, and multi-step agent sessions.
Agent and coding workflows
Agent and coding workflows
MiniMax-M3.1-Flash-Preview is designed for agentic reasoning, tool use, coding, and structured task execution.
Tunable thinking depth (effort)
Tunable thinking depth (effort)
Tune thinking depth with
effort, which accepts low, medium, high, xhigh, and max; higher levels think more thoroughly. When omitted, the default is max. See Thinking.Multimodal chat input
Multimodal chat input
Supports multimodal inputs, including text, images, and video, for a wide range of content understanding and analysis scenarios.
URL Configuration
Before calling MiniMax models, prepare the following:Calling Example
MiniMax accepts both Anthropic-style and OpenAI-style request formats. The two examples below are equivalent non-streaming calls; flipstream to true to switch to streaming responses.
Anthropic-Compatible (Recommended)
Supports thinking blocks, interleaved thinking, and other advanced features — this is the default path.OpenAI-Compatible
Already wired up to the OpenAI SDK? Swapbase_url and model for the values below and you can keep using your existing client without migrating to a new SDK.
Thinking
MiniMax-M3.1-Flash-Preview reasons before it answers, breaking a complex problem into steps before producing a response. This measurably improves accuracy on agentic reasoning, tool use, coding, and maths. Thinking is on by default and needs no configuration. Thinking content comes back separately from the answer: on the OpenAI-compatible protocol, thinking is always returned onreasoning_content while content holds only the final answer, ready to display without parsing it out of <think> tags.
Thinking depth (effort)
effort accepts low, medium, high, xhigh, and max. Higher levels think more thoroughly and produce more output tokens at higher latency. For MiniMax-M3.1-Flash-Preview, omitting effort defaults to max. The field name differs per protocol:
For the Anthropic-compatible API, explicitly set
output_config.effort to max:
Limitations
Thinking cannot be turned off. Sendingthinking: {"type": "disabled"} or effort: "none" returns 400:
effort level rather than trying to disable thinking.
API Reference
Anthropic API Compatible (Recommended)
Call MiniMax models via Anthropic SDK, supporting streaming output and Interleaved Thinking
OpenAI API Compatible
Call MiniMax models via OpenAI SDK
Using MiniMax M-series Models in AI Coding Tools
Use MiniMax M-series models in Claude Code, Cursor, and other tools
Chat Model
M2-her chat model, designed for role-playing and multi-turn dialogue scenarios
Contact Us
If you encounter any issues while using MiniMax models:- Contact our technical support team through official channels such as email Model@minimax.io
- Submit an Issue on our Github repository