Skip to main content
MiniMax-M3.1-Flash-Preview is available only through M Plan and MiniMax Code for now. Get your Subscription Key.

Model Overview

MiniMax offers multiple LLMs to meet different scenario requirements. MiniMax-M3.1-Flash-Preview is the latest M-series language model for agentic reasoning, tool use, coding, and long-context tasks, and it supports tuning thinking depth with effort. MiniMax-M3, MiniMax-M2.7, and MiniMax-M2.7-highspeed also remain available. Earlier models are listed under Legacy Models below.

Supported Models

For details on how tps (Tokens Per Second) is calculated, please refer to FAQ > About APIs.

MiniMax M3.1-Flash-Preview Key Highlights

MiniMax-M3.1-Flash-Preview supports up to a 1,000,000-token context window for long documents, codebases, and multi-step agent sessions.
MiniMax-M3.1-Flash-Preview is designed for agentic reasoning, tool use, coding, and structured task execution.
Tune thinking depth with effort, which accepts low, medium, high, xhigh, and max; higher levels think more thoroughly. When omitted, the default is max. See Thinking.
Supports multimodal inputs, including text, images, and video, for a wide range of content understanding and analysis scenarios.

URL Configuration

Before calling MiniMax models, prepare the following:

Calling Example

MiniMax accepts both Anthropic-style and OpenAI-style request formats. The two examples below are equivalent non-streaming calls; flip stream to true to switch to streaming responses. Supports thinking blocks, interleaved thinking, and other advanced features — this is the default path.

OpenAI-Compatible

Already wired up to the OpenAI SDK? Swap base_url and model for the values below and you can keep using your existing client without migrating to a new SDK.

Thinking

MiniMax-M3.1-Flash-Preview reasons before it answers, breaking a complex problem into steps before producing a response. This measurably improves accuracy on agentic reasoning, tool use, coding, and maths. Thinking is on by default and needs no configuration. Thinking content comes back separately from the answer: on the OpenAI-compatible protocol, thinking is always returned on reasoning_content while content holds only the final answer, ready to display without parsing it out of <think> tags.

Thinking depth (effort)

effort accepts low, medium, high, xhigh, and max. Higher levels think more thoroughly and produce more output tokens at higher latency. For MiniMax-M3.1-Flash-Preview, omitting effort defaults to max. The field name differs per protocol: For the Anthropic-compatible API, explicitly set output_config.effort to max:

Limitations

Thinking cannot be turned off. Sending thinking: {"type": "disabled"} or effort: "none" returns 400:
To reduce the latency and token usage of thinking, lower the effort level rather than trying to disable thinking.

API Reference

Anthropic API Compatible (Recommended)

Call MiniMax models via Anthropic SDK, supporting streaming output and Interleaved Thinking

OpenAI API Compatible

Call MiniMax models via OpenAI SDK

Using MiniMax M-series Models in AI Coding Tools

Use MiniMax M-series models in Claude Code, Cursor, and other tools

Chat Model

M2-her chat model, designed for role-playing and multi-turn dialogue scenarios

Contact Us

If you encounter any issues while using MiniMax models:
  • Contact our technical support team through official channels such as email Model@minimax.io
  • Submit an Issue on our Github repository