Notice
Wan3.0 Officially Launched: Wan3.0 is a native large-scale model for generating 30-second videos, seamlessly orchestrating all audiovisual elements and supporting multi-modal references, delivering viOriginal Manufacturer Authorized · Genuine and TraceableEnterprise‑grade large‑model intelligent gateway platform—provides enterprises with compliant, controllable, stable, and efficient AI model API aggregation and governance services; one‑stop access to⚡️This site has been integrated with the official versions of GLM 5.3 and DeepSeek V4❗️

Vision / MiniMax

MiniMax-M3

API name: minimax-m3

MultimodalCodeDeep ThinkingTool Use

Overview

Leveraging industry-leading coding and agentic capabilities, a 1M‑token ultra‑long context window, and native multimodal support, MiniMax M3 excels at enterprise‑level tasks such as long‑document understanding, high‑quality content generation, code writing, bug fixing, and native app development. Its robust agentic capabilities enable end‑to‑end workflow integration, while its native multimodality delivers a seamless, natural, and fluid mixed‑media interaction experience.

  • Model type: Vision
  • Input: Text / Image / Video
  • Output: Text
  • Endpoints: chat / chat_completion
Tiered billing details
TierInputOutputCache hit
Fallback tier$0.0006/1K$0.0024/1K$0.0001/1K

Billed bylen (input context token count). Coefficient is the $ / 1M tokens price. Default group price (other groups converted by discount).

Capabilities

  • Intelligent Reasoning: Supports multi-step thinking and complex logical inference (with a switchable “thinking” mode).
  • Code Generation: Cutting-edge coding capabilities, covering mainstream languages such as Python, JavaScript, and Go, with code ready for direct deployment.
  • Tool Invocation: Native function calls and tool usage, featuring robust agent‑based functionality and advanced multi‑step tool orchestration.
  • Extended Context: Up to 1 million tokens, with the MSA architecture significantly enhancing long‑context processing efficiency.
  • Native Multimodality: Supports image and video inputs, enabling deep semantic alignment.

Use cases

  • Agent Orchestration: Long-Chain Multi-Round Workflows and Autonomous Task Execution
  • Long-Range Coding: Complex Software Engineering, Performance Optimization, and Long-Term Development Tasks
  • Code Review: Batch PR Review and Code Issue Localization
  • RAG Q&A: Combining Ultra-Long Context with Enterprise Knowledge Bases
  • Multimodal Understanding: Image/Video Analysis and Desktop Operation Assistance

Code example

Call with minimax-m3 :

curl https://www.starunion.net/v1/chat/completions \
  -H "Authorization: Bearer YOUR_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "model": "minimax-m3",
    "messages": [{
      "role": "user",
      "content": [
        { "type": "text", "text": "Describe this image" },
        { "type": "image_url", "image_url": { "url": "https://example.com/photo.jpg" } }
      ]
    }]
  }'