Auto
Auto
自动路由:根据请求内容自动选择最合适的底层模型。
Search and compare text, image, video and audio models from leading providers. Review capabilities and pricing, then integrate with one API key.
自动路由:根据请求内容自动选择最合适的底层模型。
Explore the GPT Image 2.5 API.

claude-fable-5-1 Next generation intelligence for long-running agents

Explore the GLM 5.3 API.

Access to Claude Fable 5 has been restored. It brings 5th-generation intelligence to your most ambitious coding and professional work.
Excels at agentic reasoning, knowledge work, and tool use.
Super powerful video generation model, with sound effects, supports chat format.
OpenAI's most capable model, engineered for the hardest tasks and long-running agentic workflows. GPT-5.5 Pro excels at complex coding, computer use, deep research, data analysis, and scientific reasoning — delivering frontier-level intelligence at GPT-5.4 latency with greater token efficiency. Ideal for enterprise and professional use cases demanding the highest standard of accuracy and autonomous task execution.
DeepSeek V4 Flash is an efficiency-optimized Mixture-of-Experts model from DeepSeek with 284B total parameters and 13B activated parameters, supporting a 1M-token context window. It is designed for fast inference and high-throughput workloads, while maintaining strong reasoning and coding performance.
OpenAI's most capable image generation model, featuring near-perfect text rendering across multiple languages, up to 4K resolution, and reasoning-powered Thinking Mode. Built for production workflows that demand accuracy, speed, and on-brand visual output.
xAI's latest image-to-video model features native synchronized audio generation — video and sound are produced in a single inference pass. Supports 480p/720p output, clips up to 15 seconds, and ranked #1 on the Image-to-Video Arena leaderboard.
OpenAI's smartest and most intuitive flagship model, designed for complex coding, agentic workflows, computer use, data analysis, and in-depth research. Delivers frontier-level intelligence with the same low latency as GPT-5.4, while completing tasks with greater token efficiency. The go-to model for demanding professional and enterprise workloads.
GPT Image 2 is openai state-of-the-art image generation model for fast, high-quality image generation and editing. It supports flexible image sizes and high-fidelity image inputs.

gpt-5.2-chat-latest is the Chat-optimized snapshot of OpenAI’s GPT-5.2 family (branded in ChatGPT as GPT-5.2 Instant). It is the model for interactive/chat use cases that need a blend of speed, long-context handling, multimodal inputs and reliable conversational behaviour.
MiniMax-M2.7 offers the same top-tier intelligence as the standard version—including recursive self-evolution and expert-level office productivity—but is designed for applications requiring sub-second latency and high-speed token generation. Leveraging an enhanced inference backbone architecture, its output speed is 66% faster than the standard model (reaching 100 tps). It is the preferred choice for interactive programming assistants, real-time agent loop execution, and high-throughput enterprise pipelines with stringent completion time requirements.
MiMo-V2.5 is Xiaomi's native full-modal model. It achieves professional-grade agent performance at about half the cost of inference, while outperforming MiMo-V2-Omni in multimodal perception in image and video understanding tasks.
MiMo-V2.5-Pro is Xiaomi's flagship model, excelling in general-purpose agent capabilities and complex software engineering.
Qwen 3.6-Plus is now available, featuring enhanced code development capabilities and improved efficiency in multimodal recognition and inference, making the Vibe Coding experience even better.
GLM-5.1 (released April 2026), purpose-built for long-horizon autonomous tasks. Unlike traditional models optimized for short interactions, GLM-5.1 excels at maintaining goal alignment, reducing strategy drift, and delivering production-grade results over extended periods — up to 8 hours of continuous autonomous work on a single complex task. It represents a major leap in agentic engineering, shifting evaluation from single-turn intelligence to real-world sustained execution.
Kimi K2.6 is Kimi's latest and most intelligent model, possessing stronger and more stable long-term code writing capabilities, significantly improved instruction compliance and self-correction abilities, and supporting text, image, and video input, thinking and non-thinking modes, and dialogue and agent tasks.
Claude Mythos Preview is our most capable frontier model to date, and shows a striking leap in scores on many evaluation benchmarks compared to our previous frontier model, Claude Opus 4.6.
MiMo-V2-Omni is a frontier omni-modal model that natively processes image, video, and audio inputs within a unified architecture. It combines strong multimodal perception with agentic capability - visual grounding, multi-step planning, tool use, and code execution - making it well-suited for complex real-world tasks that span modalities. 256K context window.
MiMo-V2-Pro is Xiaomi's flagship foundation model, featuring over 1T total parameters and a 1M context length, deeply optimized for agentic scenarios. It is highly adaptable to general agent frameworks like OpenClaw. It ranks among the global top tier in the standard PinchBench and ClawBench benchmarks, with perceived performance approaching that of Opus 4.6. MiMo-V2-Pro is designed to serve as the brain of agent systems, orchestrating complex workflows, driving production engineering tasks, and delivering results reliably.
GLM-5 Turbo is a new model from Z.ai designed for fast inference and strong performance in agent-driven environments such as OpenClaw scenarios.
GPT-5.3 Instant model used in ChatGPT
Sora 2 Pro is our most advanced and powerful media generation model, capable of generating videos with synchronized Audio. It can create detailed, dynamic video clips from natural language or images.
Midjourney video generation
Explore the mj_turbo_imagine API.
Midjourney drawing
Generate videos from text prompts, animate still images, or edit existing videos with natural language. The API supports configurable duration, aspect ratio, and resolution for generated videos — with the SDK handling the asynchronous polling automatically.
The best voice model for audio in, audio out with Chat Completions.
The best voice model for audio in, audio out.