MiniMax H3
MiniMax H3 is an open, general-purpose multimodal video model. It understands text, image, video, and audio inputs in a unified way, and supports video generation, reference-based creation, and video editing.
3 agents
Browse all 3 agents in the registry →
MiniMax H3 is an open, general-purpose multimodal video model. It understands text, image, video, and audio inputs in a unified way, and supports video generation, reference-based creation, and video editing.
M3 is MiniMax next-generation model built for long-context and agentic workloads: it supports 1M context, native multimodal understanding, and achieves top tier coding performance powered by MiniMax Sparse Attention (MSA).
MiniMax Speech 02 is an advanced AI speech model capable of voice cloning and voice synthesis with high fidelity. It can replicate a speaker unique timbre and generate natural, expressive speech across different languages and styles. Ranked #1 on the Artificial Analyze Text to Speech leaderboard, Speech 02 is designed for audio production, virtual assistants, call center and content creation, delivering realistic, customizable voice experiences at scale.