279 models
Open multilingual model optimized for generation across 23 languages
Compact open multilingual model optimized for generation across 23 languages
Open multilingual vision model for OCR, visual reasoning, and image question answering
Compact open multilingual vision model for OCR and visual question answering
Claude model for creative writing, analysis, and controlled agent workflows
Legacy model retained for compatibility with older integrations
Fast Claude model for responsive assistance, classification, and lightweight agents
Fast Claude model for responsive assistance, classification, and lightweight agents
Fast Claude lane for lightweight agents, office tasks, and responsive chat
Flagship Claude model for deep reasoning, coding, and long-horizon agents
Flagship Claude model for deep reasoning, coding, and long-horizon agents
Flagship Claude model for deep reasoning, coding, and long-horizon agents
Flagship Claude model for deep reasoning, coding, and long-horizon agents
Flagship Claude model for deep reasoning, coding, and long-horizon agents
Flagship Claude model for deep reasoning, coding, and long-horizon agents
High-end Claude for difficult coding, planning, and slower expert reasoning
Stronger Opus tier for advanced software work and high-stakes reasoning
Top Claude Opus tier for the hardest reasoning, coding, and long-horizon agents
Strongest Claude Opus model for coding, agents, and professional work
Balanced Claude model for coding, analysis, agent workflows, and cost control
Balanced Claude model for coding, analysis, agent workflows, and cost control
Balanced Claude model for coding, analysis, agent workflows, and cost control
Balanced Claude model for coding, analysis, agent workflows, and cost control
Balanced Claude model for coding, analysis, agent workflows, and cost control
Balanced Claude model for coding, analysis, agent workflows, and cost control
Claude workhorse for coding agents, careful analysis, and production cost control
Everyday Claude agent model for coding, planning, browsing, and general work
Mistral code model for completions, refactors, and developer IDE workflows
Cohere command model for multilingual enterprise agents, tools, and chat
Cohere's stronger command model for multilingual agents and enterprise workflows
Cohere reasoning model for multilingual enterprise agents, tools, and complex workflows
Translation model for multilingual conversion, localization, and cross-language workflows
Cohere vision model for multilingual document analysis, OCR, and image understanding
Cohere retrieval model for long-context chat and enterprise RAG workflows
Cohere's RAG workhorse for long-context enterprise search and tool use
Cohere retrieval model for long-context chat and enterprise RAG workflows
Open Command R model optimized for Arabic enterprise chat, RAG, and cultural knowledge
Maximum-comprehensiveness agentic researcher for multi-step investigation, synthesis, and cited reports
DeepSeek chat model for instruction following, coding, and analysis
DeepSeek reasoning model for multi-step analysis, math, coding, and tools
Fast DeepSeek V4 lane for economical reasoning, coding, and long-context work
Open MoE flagship with million-token context for coding and long agent runs
Classic open reasoning model for transparent math, coding, and deliberate problem solving
Mistral's coding-agent model for repository work, terminal tasks, and software fixes
Mistral coding agent model for repository tasks and software engineering workflows
Mistral coding agent model for repository tasks and software engineering workflows
Mistral coding agent model for repository tasks and software engineering workflows
Multi-agent model for routing expert agents across complex analytical tasks
Quality-first multi-agent model for hard research, analysis, and competitions
Earlier Gemini Flash workhorse for responsive multimodal apps and tool use
Low-latency Gemini model for high-volume multimodal and agent workloads
Specialized Gemini 2.5 model for browser-control agents that automate UI tasks
Fast Gemini workhorse for multimodal apps where latency and price matter
Speech generation model for controllable voice, narration, and audio delivery
Lean Gemini 2.5 lane for cheap multimodal traffic and quick agents
Google's proven reasoning model for coding, math, and multimodal analysis
Speech generation model for controllable voice, narration, and audio delivery
New Gemini flash lane bringing frontier-style multimodal reasoning to cheaper runs
Preview Gemini flagship for complex reasoning, coding, and rich multimodal prompts
Low-latency Gemini model for high-volume multimodal and agent workloads
Low-latency Gemini model for high-volume multimodal and agent workloads
High-quality, low-latency Live API model for real-time dialogue and voice-first AI applications
Low-latency speech generation with steerable prompts and expressive audio tags
Reasoning-first Gemini preview for agentic coding and complex problem solving
Advanced Gemini model for complex reasoning, coding, and multimodal analysis
Fast Gemini model balancing multimodal reasoning, tool use, and cost
Fast Gemini model balancing multimodal reasoning, tool use, and cost
Low-latency audio-to-audio model for real-time speech translation across 70+ languages
Fast Gemini model balancing multimodal reasoning, tool use, and cost
Agentic model for autonomous multi-step research, synthesis, and cited reports
Embedding model for semantic search, retrieval, clustering, and ranking pipelines
Multimodal embedding model mapping text, images, video, audio, and PDFs into a unified embedding space
Fast Gemini model balancing multimodal reasoning, tool use, and cost
Low-latency Gemini model for high-volume multimodal and agent workloads
Video generation and editing model for fast, conversational text- and image-to-video workflows
Vision-language model for embodied reasoning: spatial understanding, task planning, and physical-world agentic robotics
Open Gemma instruction model for efficient chat and self-hosted deployments
Largest Gemma 4 instruction model for open, self-hosted chat and reasoning
Open Gemma instruction model for efficient chat and self-hosted deployments
Open Gemma instruction model for efficient chat and self-hosted deployments
Hybrid-reasoning GLM release that made the 4.5 line broadly useful
Lighter GLM-4.5 variant for fast coding assistance and cheaper agents
Efficient GLM model for fast reasoning, coding, and agent workflows
GLM vision model for visual reasoning, documents, and multimodal agents
Late GLM-4 workhorse for coding agents, reasoning, and structured tasks
GLM vision model for visual reasoning, documents, and multimodal agents
Mature GLM model for dependable coding, reasoning, and structured agent tasks
Budget GLM lane for fast coding help, routing, and everyday automation
Efficient GLM model for fast reasoning, coding, and agent workflows
General GLM flagship for coding, analysis, and tool-heavy engineering workflows
Faster GLM-5 lane for coding agents that need lower latency
Strong GLM coding model for agentic engineering, terminals, and repository generation
Open flagship GLM for long-horizon coding agents and million-token context work
Fast GLM vision model for screenshots, documents, and multimodal agent tasks
Open GPT reasoning model for self-hosted agents and controllable deployments
Open GPT reasoning model for self-hosted agents and controllable deployments
Safety model for policy screening, moderation, and risk-aware routing workflows
Streaming speech-to-text model for low-latency transcript deltas from live audio
Compact GPT model for low-latency assistance and high-volume workloads
GPT model for general reasoning, writing, coding, and tool-assisted tasks
Compact GPT model for low-latency assistance and high-volume workloads
Long-lived GPT workhorse for coding, instruction following, and production apps
Affordable GPT-4.1 lane for fast coding help and structured extraction
Tiny GPT-4.1 option for classification, routing, and very high-volume tasks
Omni-era GPT for multimodal chat, practical coding, and general assistants
GPT model for general reasoning, writing, coding, and tool-assisted tasks
GPT model for general reasoning, writing, coding, and tool-assisted tasks
GPT model for general reasoning, writing, coding, and tool-assisted tasks
Small omni GPT for cheap multimodal assistance and production-scale traffic
Original GPT-5 workhorse for reasoning, coding, writing, and tool workflows
Chat-tuned GPT model for conversational assistance, writing, and tool workflows
Small GPT-5 for responsive agents, coding help, and everyday automation
Tiny GPT-5 lane for routing, extraction, classification, and bulk jobs
Higher-accuracy GPT-5 tier for tough analysis, coding reviews, and planning
Coding-optimized GPT model for repository edits, reviews, and agentic software work
Sharper GPT-5 generation for coding, product work, and tool-assisted tasks
Chat-tuned GPT-5.1 for polished assistants, writing, and product conversations
Codex GPT for repository edits, code review, and practical software agents
Coding-optimized GPT model for repository edits, reviews, and agentic software work
Coding-optimized GPT model for repository edits, reviews, and agentic software work
Reliable GPT generation for broad coding, writing, and tool-assisted product work
Chat-tuned GPT model for conversational assistance, writing, and tool workflows
Code-specialist GPT for repository edits, reviews, and long-running software agents
Higher-accuracy GPT-5.2 variant for tougher reasoning and review workflows
Chat-tuned GPT model for conversational assistance, writing, and tool workflows
Coding-optimized GPT model for repository edits, reviews, and agentic software work
Agent-ready GPT for coding and computer-use workflows at a lower cost
Strong small GPT for coding subagents, quick tool use, and high-volume work
Cheapest GPT-5.4 lane for simple routing, extraction, and bulk automation
More exact GPT-5.4 tier for demanding professional reasoning and agent tasks
Default frontier GPT for coding, computer use, research, and knowledge work
Compact GPT model for low-latency assistance and high-volume workloads
Highest-accuracy GPT-5.5 tier for slower, precision-heavy reasoning and coding
Cost-efficient GPT-5.6 model for fast, high-volume workloads
Frontier GPT-5.6 model for complex professional work, coding, and agentic workflows
Balanced GPT-5.6 model for capable, cost-efficient everyday work
OpenAI image model for production generation, edits, and brand-safe visual workflows
Image model for prompt-driven generation, editing, and visual design workflows
Image model for prompt-driven generation, editing, and visual design workflows
Realtime speech-to-speech model with configurable reasoning, tool use, and robust voice-agent behavior
Grok model for agentic tool use, reasoning, coding, and live assistance
Reasoning Grok for document-heavy analysis and long-horizon tool use
xAI's default Grok for chat, coding, agentic tools, and lower hallucination risk
xAI's latest Grok for chat, coding, agentic tools, and lower hallucination risk
Fast Grok coding model tuned for agentic engineering and iterative edits
Video model for image-to-video generation, editing, and extension workflows
Tencent Hy reasoning model for coding, instruction following, and agent tasks
Tencent Hy reasoning model for coding, instruction following, and agent tasks
Multimodal MoE reasoning model (975B total, 41B active) for text, image, and audio
Thinking Kimi model for slower research passes, planning, and hard technical questions
Kimi reasoning model for long-horizon research, planning, and tool use
Earlier Kimi frontier model for long-context agents, coding, and multimodal work
Multimodal Kimi workhorse for agent loops, coding tasks, and visual context
Coding-focused Kimi model, stronger on long-horizon repo work with less overthinking
Lower-latency Kimi Code variant for interactive edits and coding-agent loops
Multimodal Kimi model with 1M context and toggleable max-effort thinking for long-horizon agent work
Poolside's open-weight model for agentic coding and long-horizon work
Agentic coding model from Poolside in the XS size class for local deployment
Agentic coding model from Poolside in the XS size class for local deployment
Agentic coding model from Poolside in the XS size class for local deployment
Nemotron model for efficient reasoning, coding, and specialized AI agents
Safety model for policy screening, moderation, and risk-aware routing workflows
Flagship Nemotron model for high-throughput reasoning and complex agents
Nemotron model for efficient reasoning, coding, and specialized AI agents
Nemotron model for efficient reasoning, coding, and specialized AI agents
Open multimodal Llama for strong reasoning with efficient everyday serving
Open Llama with long-context vision for efficient multimodal agents
Embedding model for semantic search, retrieval, clustering, and ranking pipelines
Reranking model for improving retrieval quality in search and recommendation systems
Popular open Llama workhorse for multilingual chat, coding, and self-hosting
Meituan LongCat-2.0, a reasoning model with tool calling and a 1M-token context window
Music generation model for short 30-second clips, loops, and previews from text or image prompts
Music generation model for full-length songs from text or images with vocals and structure
Mistral reasoning model for transparent analysis, math, and complex decisions
Microsoft coding model built for fast, efficient assistance in everyday developer workflows
MiMo flash model for fast multimodal assistance and agent workflows
MiMo omni model for text, image, video, audio, and agents
Earlier MiMo Pro model for multimodal agents, reasoning, and code tasks
Open MiMo model for multimodal coding agents and long-context automation
Stronger MiMo Pro tier for multimodal reasoning and coding-agent execution
MiMo pro model for strong multimodal reasoning and agent execution
Efficient open MiniMax model built for coding agents and tool-heavy workflows
Earlier MiniMax agent model for practical coding and productivity tasks
Prior MiniMax coding model for agent workflows, office edits, and automation
High-speed MiniMax model for low-latency coding and agent workflows
Open MiniMax flagship for coding agents, office automation, and complex environments
Low-latency M2.7 variant for interactive coding plans and agent loops
MiniMax multimodal model for long-context coding, perception, and agent planning
Flagship Mistral model for advanced reasoning, coding, and multilingual work
Flagship Mistral model for advanced reasoning, coding, and multilingual work
Mistral's largest general model for enterprise agents, coding, and multilingual reasoning
Balanced Mistral model for enterprise assistants, multilingual work, and tools
Mistral model for multilingual chat, reasoning, and tool-assisted workflows
Balanced Mistral model for enterprise assistants, multilingual work, and tools
Efficient Mistral-NVIDIA open model for multilingual chat and local deployment
Mistral model for multilingual chat, reasoning, and tool-assisted workflows
Efficient Mistral model for fast chat, extraction, and production assistants
Efficient Mistral model for fast chat, extraction, and production assistants
Fast Mistral production model for chat, extraction, and cost-sensitive agents
Muse Spark is a natively multimodal reasoning model with support for tool-use, visual chain of thought, and multi-agent orchestration.
Nano Banana image model for fast generation, edits, and character-consistent assets
Image model for prompt-driven generation, editing, and visual design workflows
Image model for prompt-driven generation, editing, and visual design workflows
Fastest, most cost-efficient Gemini image model for high-volume 1K generation and editing
Nano Banana Pro for higher-fidelity image generation and design-heavy edits
Nano Banana Pro for higher-fidelity image generation and design-heavy edits
Safety model for policy screening, moderation, and risk-aware routing workflows
Small Nemotron 3 MoE for efficient coding, math, and long-context agents
Open Nemotron omni model combining reasoning with text, vision, and audio
Nemotron middle tier for collaborative agents and high-volume reasoning workloads
Largest Nemotron 3 model for maximum open-weight reasoning and agent accuracy
Safety model for policy screening, moderation, and risk-aware routing workflows
Nemotron model for efficient reasoning, coding, and specialized AI agents
Safety model for policy screening, moderation, and risk-aware routing workflows
Compact Nemotron model for efficient reasoning and deployable AI agents
Nemotron multimodal model for visual reasoning and agentic AI workflows
Compact Nemotron model for efficient reasoning and deployable AI agents
Nemotron multimodal model for visual reasoning and agentic AI workflows
Cohere coding model for practical software engineering and agentic edits
O-series reasoning model for hard analysis, math, coding, and planning
O-series reasoning model for hard analysis, math, coding, and planning
Deliberate o-series reasoner for hard math, coding, and multi-step analysis
Research model for long-horizon investigation, synthesis, and analytical reports
Smaller o-series reasoner for economical coding, math, and planning tasks
High-effort o3 tier for difficult technical reasoning and careful answers
Fast o-series model for compact reasoning, coding, and tool use
Research model for long-horizon investigation, synthesis, and analytical reports
Open coding-reasoning model for repository tasks and self-improving agents
Large coding-reasoning model for agentic software tasks and RL search
Large coding-reasoning model for agentic software tasks and RL search
Open coding-reasoning model for repository tasks and self-improving agents
Mistral vision-language model for image understanding and multimodal chat
Mistral's larger vision model for document-heavy image understanding and chat
Efficient Qwen model for fast chat, extraction, and high-volume workloads
Flagship Qwen model for complex reasoning, coding, and agentic workflows
Qwen instruction model for multilingual chat, reasoning, and tool use
Efficient Qwen model for fast chat, extraction, and high-volume workloads
Qwen omni model for text, vision, audio, and multimodal agent tasks
Qwen vision-language model for visual reasoning, documents, and agent tasks
Qwen vision-language model for visual reasoning, documents, and agent tasks
Qwen vision-language model for visual reasoning, documents, and agent tasks
Large open Qwen MoE for multilingual reasoning, coding, and tool use
Dense open Qwen model for self-hosted chat, reasoning, and coding
Qwen coding model for software agents, repository edits, and code reasoning
Hosted Qwen coder for software agents, repo edits, and long-context code
Flagship Qwen3 model for coding agents, complex reasoning, and tool use
Smaller Qwen coder for efficient local agents and repo-level fixes
Open Qwen coding heavyweight for repository reasoning and agentic engineering
Efficient Qwen thinking model for local reasoning, math, and coding agents
Qwen instruction model for multilingual chat, reasoning, and tool use
Qwen vision-language model for visual reasoning, documents, and agent tasks
Qwen vision-language model for visual reasoning, documents, and agent tasks
Qwen vision-language model for visual reasoning, documents, and agent tasks
Qwen vision-language model for visual reasoning, documents, and agent tasks
Large open Qwen multimodal MoE for visual agents and long technical tasks
Qwen instruction model for multilingual chat, reasoning, and tool use
Qwen vision-language model for visual reasoning, documents, and agent tasks
Qwen vision-language model for visual reasoning, documents, and agent tasks
Open multimodal Qwen MoE for local agents that need vision, audio, and code
Qwen vision-language model for visual reasoning, documents, and agent tasks
Flagship Qwen model for complex reasoning, coding, and agentic workflows
Earlier Qwen multimodal workhorse for million-token agent and document tasks
Qwen frontier model tuned for agent frameworks, coding assistants, and long tasks
Multimodal Qwen workhorse for long-context agents, visual inputs, and coding
Preview Qwen flagship for million-token multimodal reasoning and long-horizon agentic workflows
Qwen reasoning model for deliberate problem solving, math, and coding
Flagship Indian-language reasoning model for enterprise multilingual applications
Efficient Indian-language reasoning model for chat, coding, and multilingual work
Fast web-grounded Sonar for current answers, citations, and lightweight retrieval
Deeper Sonar search model with broader retrieval and stronger synthesis
Web-grounded Sonar for multi-step research questions that need cited reasoning
StepFun flash lane for quick multimodal reasoning and coding assistance
StepFun flash model for efficient multimodal reasoning, coding, and tool use
Newer StepFun flash model for faster agents, coding, and multimodal prompts
Video model for prompt-guided generation, editing, and motion workflows
Video model for prompt-guided generation, editing, and motion workflows
Video model for prompt-guided generation, editing, and motion workflows
Open Whisper checkpoint for robust multilingual transcription and captioning
Speech transcription model for accurate audio-to-text and captioning workflows