Viberank
Trending models
Popular Labs
Hottest providers
Top Agents
Login
Sign up
← All providers
Vercel AI Gateway
Provider ID: vercel
Details
SDK package
@ai-sdk/gateway
Models
310
Links
Documentation
Auth: AI_GATEWAY_API_KEY
Models (310)
Qwen3-14B
Qwen instruction model for multilingual chat, reasoning, and tool use
40,960 ctx
reasoning
Qwen3 235B A22B Instruct 2507
Qwen instruction model for multilingual chat, reasoning, and tool use
262,144 ctx
reasoning
Qwen3-30B-A3B
Qwen instruction model for multilingual chat, reasoning, and tool use
40,960 ctx
reasoning
Qwen 3.32B
Qwen instruction model for multilingual chat, reasoning, and tool use
128,000 ctx
reasoning
Qwen 3.6 Max Preview
Flagship Qwen model for complex reasoning, coding, and agentic workflows
240,000 ctx
reasoning
Qwen3 235B A22B Thinking 2507
Qwen vision-language model for visual reasoning, documents, and agent tasks
131,072 ctx
reasoning
Qwen3 Coder 480B A35B Instruct
Qwen coding model for software agents, repository edits, and code reasoning
262,144 ctx
reasoning
Qwen 3 Coder 30B A3B Instruct
Qwen coding model for software agents, repository edits, and code reasoning
262,144 ctx
reasoning
Qwen3 Coder Next
Qwen coding model for software agents, repository edits, and code reasoning
256,000 ctx
reasoning
Qwen3 Coder Plus
Hosted Qwen coder for software agents, repo edits, and long-context code
1,000,000 ctx
Qwen3 Embedding 0.6B
Embedding model for semantic search, retrieval, clustering, and ranking pipelines
32,768 ctx
Qwen3 Embedding 4B
Embedding model for semantic search, retrieval, clustering, and ranking pipelines
32,768 ctx
Qwen3 Embedding 8B
Embedding model for semantic search, retrieval, clustering, and ranking pipelines
32,768 ctx
Qwen3 Max
Flagship Qwen3 model for coding agents, complex reasoning, and tool use
262,144 ctx
Qwen3 Max Preview
Flagship Qwen model for complex reasoning, coding, and agentic workflows
262,144 ctx
Qwen 3 Max Thinking
Qwen reasoning model for deliberate problem solving, math, and coding
256,000 ctx
reasoning
Qwen3 Next 80B A3B Instruct
Qwen instruction model for multilingual chat, reasoning, and tool use
131,072 ctx
Qwen3 Next 80B A3B Thinking
Efficient Qwen thinking model for local reasoning, math, and coding agents
131,072 ctx
reasoning
Qwen3 VL 235B A22B Instruct
Qwen vision-language model for visual reasoning, documents, and agent tasks
131,072 ctx
Qwen3 VL Instruct
Qwen vision-language model for visual reasoning, documents, and agent tasks
131,072 ctx
Qwen3 VL Thinking
Qwen vision-language model for visual reasoning, documents, and agent tasks
131,072 ctx
reasoning
Qwen 3.5 Flash
Qwen vision-language model for visual reasoning, documents, and agent tasks
1,000,000 ctx
reasoning
Qwen 3.5 Plus
Qwen vision-language model for visual reasoning, documents, and agent tasks
1,000,000 ctx
reasoning
Qwen 3.6 27B
Qwen vision-language model for visual reasoning, documents, and agent tasks
256,000 ctx
reasoning
Qwen 3.6 Plus
Earlier Qwen multimodal workhorse for million-token agent and document tasks
1,000,000 ctx
reasoning
Qwen 3.7 Max
Qwen frontier model tuned for agent frameworks, coding assistants, and long tasks
991,000 ctx
reasoning
Qwen 3.7 Plus
Multimodal Qwen workhorse for long-context agents, visual inputs, and coding
1,000,000 ctx
reasoning
Wan v2.5 Text-to-Video Preview
Video model for prompt-guided generation, editing, and motion workflows
0
Wan v2.6 Image-to-Video
Image model for prompt-driven generation, editing, and visual design workflows
0
Wan v2.6 Image-to-Video Flash
Image model for prompt-driven generation, editing, and visual design workflows
0
Wan v2.6 Reference-to-Video
Video model for prompt-guided generation, editing, and motion workflows
0
Wan v2.6 Reference-to-Video Flash
Video model for prompt-guided generation, editing, and motion workflows
0
Wan v2.6 Text-to-Video
Video model for prompt-guided generation, editing, and motion workflows
0
Wan v2.7 Reference-to-Video
Video model for prompt-guided generation, editing, and motion workflows
0
Wan v2.7 Text-to-Video
Video model for prompt-guided generation, editing, and motion workflows
0
Nova 2 Lite
Multimodal reasoning model for visual analysis, planning, and tool use
1,000,000 ctx
reasoning
Nova Lite
Efficient model for low-latency assistance, extraction, and routine automation
300,000 ctx
Nova Micro
Efficient model for low-latency assistance, extraction, and routine automation
128,000 ctx
Nova Pro
Flagship model for demanding analysis, coding, and production agent workflows
300,000 ctx
Titan Text Embeddings V2
Embedding model for semantic search, retrieval, clustering, and ranking pipelines
8,192 ctx
Claude Haiku 3
Legacy model retained for compatibility with older integrations
200,000 ctx
Claude Fable 5
Claude model for creative writing, analysis, and controlled agent workflows
1,000,000 ctx
reasoning
Claude Haiku 4.5
Fast Claude lane for lightweight agents, office tasks, and responsive chat
200,000 ctx
reasoning
Claude Opus 4
Flagship Claude model for deep reasoning, coding, and long-horizon agents
200,000 ctx
reasoning
Claude Opus 4.1
Flagship Claude model for deep reasoning, coding, and long-horizon agents
200,000 ctx
reasoning
Claude Opus 4.5
Flagship Claude model for deep reasoning, coding, and long-horizon agents
200,000 ctx
reasoning
Claude Opus 4.6
High-end Claude for difficult coding, planning, and slower expert reasoning
1,000,000 ctx
reasoning
Claude Opus 4.7
Stronger Opus tier for advanced software work and high-stakes reasoning
1,000,000 ctx
reasoning
Claude Opus 4.7 (Fast)
Stronger Opus tier for advanced software work and high-stakes reasoning
1,000,000 ctx
reasoning
Claude Opus 4.8
Top Claude Opus tier for the hardest reasoning, coding, and long-horizon agents
1,000,000 ctx
reasoning
Claude Opus 4.8 (Fast)
Top Claude Opus tier for the hardest reasoning, coding, and long-horizon agents
1,000,000 ctx
reasoning
Claude Opus 5
Flagship Claude model for deep reasoning, coding, and long-horizon agents
1,000,000 ctx
reasoning
Claude Opus 5 (Fast)
Strongest Claude Opus model for coding, agents, and professional work
1,000,000 ctx
reasoning
Claude Sonnet 4
Balanced Claude model for coding, analysis, agent workflows, and cost control
1,000,000 ctx
reasoning
Claude Sonnet 4.5
Balanced Claude model for coding, analysis, agent workflows, and cost control
1,000,000 ctx
reasoning
Claude Sonnet 4.6
Claude workhorse for coding agents, careful analysis, and production cost control
1,000,000 ctx
reasoning
Claude Sonnet 5
Everyday Claude agent model for coding, planning, browsing, and general work
1,000,000 ctx
reasoning
Trinity Large Thinking
Flagship model for demanding analysis, coding, and production agent workflows
262,100 ctx
reasoning
Trinity Mini
Efficient model for low-latency assistance, extraction, and routine automation
131,072 ctx
FLUX.2 [flex]
Image model for prompt-driven generation, editing, and visual design workflows
0
FLUX.2 [klein] 4B
Image model for prompt-driven generation, editing, and visual design workflows
0
FLUX.2 [klein] 9B
Image model for prompt-driven generation, editing, and visual design workflows
0
FLUX.2 [max]
Image model for prompt-driven generation, editing, and visual design workflows
67,300 ctx
FLUX.2 [pro]
Image model for prompt-driven generation, editing, and visual design workflows
67,300 ctx
FLUX.1 Kontext Max
Image model for prompt-driven generation, editing, and visual design workflows
512 ctx
FLUX.1 Kontext Pro
Image model for prompt-driven generation, editing, and visual design workflows
512 ctx
FLUX.1 Fill [pro]
Image model for prompt-driven generation, editing, and visual design workflows
512 ctx
FLUX1.1 [pro]
Image model for prompt-driven generation, editing, and visual design workflows
512 ctx
FLUX1.1 [pro] Ultra
Image model for prompt-driven generation, editing, and visual design workflows
512 ctx
Seed 1.6
Multimodal reasoning model for visual analysis, planning, and tool use
256,000 ctx
reasoning
Seed 1.8
Multimodal reasoning model for visual analysis, planning, and tool use
256,000 ctx
reasoning
Seedance 2.0
Video model for prompt-guided generation, editing, and motion workflows
0
Seedance 2.0 Fast
Video model for prompt-guided generation, editing, and motion workflows
0
Seedance v1.0 Pro
Video model for prompt-guided generation, editing, and motion workflows
0
Seedance v1.0 Pro Fast
Video model for prompt-guided generation, editing, and motion workflows
0
Seedance v1.5 Pro
Video model for prompt-guided generation, editing, and motion workflows
0
Seedream 4.0
Image model for prompt-driven generation, editing, and visual design workflows
0
Seedream 4.5
Image model for prompt-driven generation, editing, and visual design workflows
0
Seedream 5.0 Lite
Image model for prompt-driven generation, editing, and visual design workflows
0
Seedream 5.0 Pro
Image model for prompt-driven generation, editing, and visual design workflows
0
Command A
Cohere command model for multilingual enterprise agents, tools, and chat
256,000 ctx
Embed v4.0
Embedding model for semantic search, retrieval, clustering, and ranking pipelines
128,000 ctx
Cohere Rerank 3.5
Reranking model for improving retrieval quality in search and recommendation systems
4,096 ctx
Cohere Rerank 4 Fast
Reranking model for improving retrieval quality in search and recommendation systems
32,000 ctx
Cohere Rerank 4 Pro
Reranking model for improving retrieval quality in search and recommendation systems
32,000 ctx
DeepSeek-R1
DeepSeek reasoning model for multi-step analysis, math, coding, and tools
128,000 ctx
reasoning
DeepSeek V3 0324
DeepSeek chat model for instruction following, coding, and analysis
163,840 ctx
DeepSeek-V3.1
DeepSeek chat model for instruction following, coding, and analysis
163,840 ctx
reasoning
DeepSeek V3.1 Terminus
DeepSeek chat model for instruction following, coding, and analysis
131,072 ctx
reasoning
DeepSeek V3.2
DeepSeek chat model for instruction following, coding, and analysis
128,000 ctx
DeepSeek V3.2 Thinking
DeepSeek chat model for instruction following, coding, and analysis
128,000 ctx
reasoning
DeepSeek V4 Flash
Fast DeepSeek V4 lane for economical reasoning, coding, and long-context work
1,000,000 ctx
reasoning
DeepSeek V4 Pro
Open MoE flagship with million-token context for coding and long agent runs
1,000,000 ctx
reasoning
Gemini 2.5 Flash
Fast Gemini workhorse for multimodal apps where latency and price matter
1,048,576 ctx
reasoning
Nano Banana (Gemini 2.5 Flash Image)
Nano Banana image model for fast generation, edits, and character-consistent assets
32,768 ctx
Gemini 2.5 Flash Lite
Lean Gemini 2.5 lane for cheap multimodal traffic and quick agents
1,048,576 ctx
reasoning
Gemini 2.5 Pro
Google's proven reasoning model for coding, math, and multimodal analysis
1,048,576 ctx
reasoning
Gemini 3 Flash
Fast Gemini model balancing multimodal reasoning, tool use, and cost
1,000,000 ctx
reasoning
Nano Banana Pro
Nano Banana Pro for higher-fidelity image generation and design-heavy edits
65,536 ctx
Gemini 3 Pro Preview
Preview Gemini flagship for complex reasoning, coding, and rich multimodal prompts
1,000,000 ctx
reasoning
Nano Banana 2
Image model for prompt-driven generation, editing, and visual design workflows
131,072 ctx
reasoning
Gemini 3.1 Flash Image Preview (Nano Banana 2)
Image model for prompt-driven generation, editing, and visual design workflows
131,072 ctx
reasoning
Gemini 3.1 Flash Lite
Low-latency Gemini model for high-volume multimodal and agent workloads
1,000,000 ctx
reasoning
Gemini 3.1 Flash Lite Image (Nano Banana 2 Lite)
Image model for prompt-driven generation, editing, and visual design workflows
65,536 ctx
reasoning
Gemini 3.1 Flash Lite Preview
Low-latency Gemini model for high-volume multimodal and agent workloads
1,000,000 ctx
reasoning
Gemini 3.1 Pro Preview
Reasoning-first Gemini preview for agentic coding and complex problem solving
1,000,000 ctx
reasoning
Gemini 3.5 Flash
Fast Gemini model balancing multimodal reasoning, tool use, and cost
1,000,000 ctx
reasoning
Gemini 3.5 Flash Lite
Low-latency Gemini model for high-volume multimodal and agent workloads
1,000,000 ctx
reasoning
Gemini 3.6 Flash
Fast Gemini model balancing multimodal reasoning, tool use, and cost
1,000,000 ctx
reasoning
Gemini Embedding 001
Embedding model for semantic search, retrieval, clustering, and ranking pipelines
8,192 ctx
Gemini Embedding 2
Embedding model for semantic search, retrieval, clustering, and ranking pipelines
0
Gemini Omni Flash Preview
Omni-modal model for text, vision, audio, and multimodal agent tasks
1,000,000 ctx
reasoning
Gemma 4 26B A4B IT
Open Gemma instruction model for efficient chat and self-hosted deployments
262,144 ctx
reasoning
Gemma 4 31B IT
Largest Gemma 4 instruction model for open, self-hosted chat and reasoning
262,144 ctx
Imagen 4 Fast
Image model for prompt-driven generation, editing, and visual design workflows
480 ctx
Imagen 4
Image model for prompt-driven generation, editing, and visual design workflows
480 ctx
Imagen 4 Ultra
Image model for prompt-driven generation, editing, and visual design workflows
480 ctx
Text Embedding 005
Embedding model for semantic search, retrieval, clustering, and ranking pipelines
8,192 ctx
Text Multilingual Embedding 002
Embedding model for semantic search, retrieval, clustering, and ranking pipelines
8,192 ctx
Veo 3.0 Fast Generate
Video model for prompt-guided generation, editing, and motion workflows
0
Veo 3.0
Video model for prompt-guided generation, editing, and motion workflows
0
Veo 3.1 Fast Generate
Video model for prompt-guided generation, editing, and motion workflows
0
Veo 3.1
Video model for prompt-guided generation, editing, and motion workflows
0
Mercury 2
Reasoning model for deliberate analysis, multi-step problem solving, and tool use
128,000 ctx
reasoning
Mercury Coder Small Beta
Coding model for repository understanding, refactors, and agentic engineering tasks
32,000 ctx
Ling 3.0 Flash
Efficient model for low-latency assistance, extraction, and routine automation
256,000 ctx
reasoning
Interfaze Beta
Multimodal reasoning model for visual analysis, planning, and tool use
1,000,000 ctx
reasoning
Kling v2.5 Turbo Image-to-Video
Image model for prompt-driven generation, editing, and visual design workflows
0
Kling v2.5 Turbo Text-to-Video
Video model for prompt-guided generation, editing, and motion workflows
0
Kling v2.6 Image-to-Video
Image model for prompt-driven generation, editing, and visual design workflows
0
Kling v2.6 Motion Control
Video model for prompt-guided generation, editing, and motion workflows
0
Kling v2.6 Text-to-Video
Video model for prompt-guided generation, editing, and motion workflows
0
Kling v3.0 Image-to-Video
Image model for prompt-driven generation, editing, and visual design workflows
0
Kling v3.0 Motion Control
Video model for prompt-guided generation, editing, and motion workflows
0
Kling v3.0 Text-to-Video
Video model for prompt-guided generation, editing, and motion workflows
0
Kat Coder Air V2.5
Coding model for repository understanding, refactors, and agentic engineering tasks
256,000 ctx
reasoning
KAT-Coder-Pro V1
Coding model for repository understanding, refactors, and agentic engineering tasks
256,000 ctx
reasoning
Kat Coder Pro V2
Coding model for repository understanding, refactors, and agentic engineering tasks
256,000 ctx
reasoning
Kat Coder Pro V2.5
Coding model for repository understanding, refactors, and agentic engineering tasks
256,000 ctx
reasoning
Llama 3.1 70B Instruct
Open Llama instruction model for multilingual chat, reasoning, and coding
128,000 ctx
Llama 3.1 8B Instruct
Open Llama instruction model for multilingual chat, reasoning, and coding
128,000 ctx
Llama-3.3-70B-Instruct
Open Llama instruction model for multilingual chat, reasoning, and coding
128,000 ctx
Llama-4-Maverick-17B-128E-Instruct-FP8
Open multimodal Llama model for strong reasoning and fast responses
128,000 ctx
Llama-4-Scout-17B-16E-Instruct-FP8
Open multimodal Llama model for long-context analysis and efficient agents
128,000 ctx
Muse Spark 1.1
Open Llama instruction model for multilingual chat, reasoning, and coding
1,048,576 ctx
reasoning
MiniMax M2
Efficient open MiniMax model built for coding agents and tool-heavy workflows
205,000 ctx
reasoning
MiniMax M2.1
Earlier MiniMax agent model for practical coding and productivity tasks
204,800 ctx
reasoning
MiniMax M2.1 Lightning
High-speed MiniMax model for low-latency coding and agent workflows
204,800 ctx
reasoning
MiniMax M2.5
Prior MiniMax coding model for agent workflows, office edits, and automation
204,800 ctx
reasoning
MiniMax M2.5 High Speed
High-speed MiniMax model for low-latency coding and agent workflows
204,800 ctx
reasoning
Minimax M2.7
Open MiniMax flagship for coding agents, office automation, and complex environments
204,800 ctx
reasoning
MiniMax M2.7 High Speed
Low-latency M2.7 variant for interactive coding plans and agent loops
204,800 ctx
reasoning
MiniMax M3
MiniMax multimodal model for long-context coding, perception, and agent planning
1,000,000 ctx
reasoning
Codestral (latest)
Mistral code model for completions, refactors, and developer IDE workflows
256,000 ctx
Codestral Embed
Embedding model for semantic search, retrieval, clustering, and ranking pipelines
8,192 ctx
Devstral 2
Mistral coding agent model for repository tasks and software engineering workflows
256,000 ctx
Devstral Small 2
Mistral coding agent model for repository tasks and software engineering workflows
256,000 ctx
Magistral Medium (latest)
Mistral reasoning model for transparent analysis, math, and complex decisions
128,000 ctx
reasoning
Magistral Small
Mistral reasoning model for transparent analysis, math, and complex decisions
128,000 ctx
reasoning
Ministral 14B
Compact Mistral model for edge, latency-sensitive, and cost-efficient workloads
256,000 ctx
Ministral 3B (latest)
Compact Mistral model for edge, latency-sensitive, and cost-efficient workloads
128,000 ctx
Ministral 8B (latest)
Compact Mistral model for edge, latency-sensitive, and cost-efficient workloads
128,000 ctx
Mistral Embed
Embedding model for semantic search, retrieval, clustering, and ranking pipelines
8,192 ctx
Mistral Large 3
Flagship Mistral model for advanced reasoning, coding, and multilingual work
256,000 ctx
Mistral Medium 3.1
Mistral model for multilingual chat, reasoning, and tool-assisted workflows
128,000 ctx
Mistral Medium Latest
Mistral model for multilingual chat, reasoning, and tool-assisted workflows
256,000 ctx
reasoning
Mistral Nemo
Mistral model for multilingual chat, reasoning, and tool-assisted workflows
128,000 ctx
Mistral Small (latest)
Efficient Mistral model for fast chat, extraction, and production assistants
32,000 ctx
Pixtral 12B
Mistral vision-language model for image understanding and multimodal chat
128,000 ctx
Kimi K2 Instruct
Kimi model for long-context chat, coding, and agentic reasoning
131,072 ctx
Kimi K2 Thinking
Thinking Kimi model for slower research passes, planning, and hard technical questions
216,144 ctx
reasoning
Kimi K2.5
Earlier Kimi frontier model for long-context agents, coding, and multimodal work
262,114 ctx
reasoning
Kimi K2.6
Multimodal Kimi workhorse for agent loops, coding tasks, and visual context
262,000 ctx
reasoning
Kimi K2.7 Code
Coding-focused Kimi model, stronger on long-horizon repo work with less overthinking
256,000 ctx
reasoning
Kimi K2.7 Code High Speed
Lower-latency Kimi Code variant for interactive edits and coding-agent loops
262,144 ctx
reasoning
Kimi K3
Kimi multimodal agent model for visual understanding, coding, and planning
1,000,000 ctx
reasoning
Morph v3 Fast
Efficient model for low-latency assistance, extraction, and routine automation
16,000 ctx
Morph v3 Large
Flagship model for demanding analysis, coding, and production agent workflows
32,000 ctx
Nemotron 3 Nano 30B A3B
Small Nemotron 3 MoE for efficient coding, math, and long-context agents
262,144 ctx
reasoning
NVIDIA Nemotron 3 Super 120B A12B
Nemotron middle tier for collaborative agents and high-volume reasoning workloads
256,000 ctx
reasoning
Nemotron 3 Ultra
Largest Nemotron 3 model for maximum open-weight reasoning and agent accuracy
1,000,000 ctx
reasoning
Nvidia Nemotron Nano 12B V2 VL
Nemotron multimodal model for visual reasoning and agentic AI workflows
131,072 ctx
reasoning
Nvidia Nemotron Nano 9B V2
Compact Nemotron model for efficient reasoning and deployable AI agents
131,072 ctx
reasoning
GPT-3.5 Turbo
Compact GPT model for low-latency assistance and high-volume workloads
16,385 ctx
GPT-4 Turbo
Compact GPT model for low-latency assistance and high-volume workloads
128,000 ctx
GPT-4.1
Long-lived GPT workhorse for coding, instruction following, and production apps
1,047,576 ctx
GPT-4.1 mini
Affordable GPT-4.1 lane for fast coding help and structured extraction
1,047,576 ctx
GPT-4.1 nano
Tiny GPT-4.1 option for classification, routing, and very high-volume tasks
1,047,576 ctx
GPT-4o
Omni-era GPT for multimodal chat, practical coding, and general assistants
128,000 ctx
GPT-4o mini
Small omni GPT for cheap multimodal assistance and production-scale traffic
128,000 ctx
GPT 4o Mini Search Preview
Compact GPT model for low-latency assistance and high-volume workloads
128,000 ctx
GPT-4o mini Transcribe
Speech transcription model for accurate audio-to-text and captioning workflows
0
GPT-4o Transcribe
Speech transcription model for accurate audio-to-text and captioning workflows
0
GPT-5
Original GPT-5 workhorse for reasoning, coding, writing, and tool workflows
400,000 ctx
reasoning
GPT-5 Chat
Image model for prompt-driven generation, editing, and visual design workflows
128,000 ctx
reasoning
GPT-5-Codex
Coding-optimized GPT model for repository edits, reviews, and agentic software work
400,000 ctx
reasoning
GPT-5 Mini
Small GPT-5 for responsive agents, coding help, and everyday automation
400,000 ctx
reasoning
GPT-5 Nano
Tiny GPT-5 lane for routing, extraction, classification, and bulk jobs
400,000 ctx
reasoning
GPT-5 pro
Higher-accuracy GPT-5 tier for tough analysis, coding reviews, and planning
400,000 ctx
reasoning
GPT-5.1-Codex
Codex GPT for repository edits, code review, and practical software agents
400,000 ctx
reasoning
GPT 5.1 Codex Max
Coding-optimized GPT model for repository edits, reviews, and agentic software work
400,000 ctx
reasoning
GPT-5.1 Codex mini
Coding-optimized GPT model for repository edits, reviews, and agentic software work
400,000 ctx
reasoning
GPT-5.1 Instant
Compact GPT model for low-latency assistance and high-volume workloads
128,000 ctx
GPT 5.1 Thinking
Image model for prompt-driven generation, editing, and visual design workflows
400,000 ctx
reasoning
GPT-5.2
Reliable GPT generation for broad coding, writing, and tool-assisted product work
400,000 ctx
reasoning
GPT-5.2 Chat
Chat-tuned GPT model for conversational assistance, writing, and tool workflows
128,000 ctx
GPT-5.2-Codex
Code-specialist GPT for repository edits, reviews, and long-running software agents
400,000 ctx
reasoning
GPT 5.2
Higher-accuracy GPT-5.2 variant for tougher reasoning and review workflows
400,000 ctx
reasoning
GPT-5.3 Chat
Chat-tuned GPT model for conversational assistance, writing, and tool workflows
128,000 ctx
GPT 5.3 Codex
Coding-optimized GPT model for repository edits, reviews, and agentic software work
400,000 ctx
reasoning
GPT 5.4
Agent-ready GPT for coding and computer-use workflows at a lower cost
1,050,000 ctx
reasoning
GPT 5.4 Mini
Strong small GPT for coding subagents, quick tool use, and high-volume work
400,000 ctx
reasoning
GPT 5.4 Nano
Cheapest GPT-5.4 lane for simple routing, extraction, and bulk automation
400,000 ctx
reasoning
GPT 5.4 Pro
More exact GPT-5.4 tier for demanding professional reasoning and agent tasks
1,050,000 ctx
reasoning
GPT 5.5
Default frontier GPT for coding, computer use, research, and knowledge work
1,000,000 ctx
reasoning
GPT 5.5 Pro
Highest-accuracy GPT-5.5 tier for slower, precision-heavy reasoning and coding
1,000,000 ctx
reasoning
GPT 5.6 Luna
GPT model for general reasoning, writing, coding, and tool-assisted tasks
1,050,000 ctx
reasoning
GPT 5.6 Sol
GPT model for general reasoning, writing, coding, and tool-assisted tasks
1,050,000 ctx
reasoning
GPT 5.6 Terra
GPT model for general reasoning, writing, coding, and tool-assisted tasks
1,050,000 ctx
reasoning
GPT Image 1
OpenAI image model for production generation, edits, and brand-safe visual workflows
0
GPT Image 1 Mini
Image model for prompt-driven generation, editing, and visual design workflows
0
GPT Image 1.5
Image model for prompt-driven generation, editing, and visual design workflows
0
GPT Image 2
Image model for prompt-driven generation, editing, and visual design workflows
0
GPT OSS 120B
Open GPT reasoning model for self-hosted agents and controllable deployments
131,072 ctx
reasoning
GPT OSS 20B
Open-weight GPT model for self-hosted reasoning and instruction-following workloads
131,072 ctx
reasoning
gpt-oss-safeguard-20b
Safety model for policy screening, moderation, and risk-aware routing workflows
131,072 ctx
reasoning
GPT-Realtime-1.5
Speech generation model for controllable voice, narration, and audio delivery
0
gpt-realtime-2
Speech generation model for controllable voice, narration, and audio delivery
0
gpt-realtime-2.1
Speech generation model for controllable voice, narration, and audio delivery
128,000 ctx
reasoning
GPT-Realtime mini
Speech generation model for controllable voice, narration, and audio delivery
0
gpt-realtime-whisper
Streaming speech-to-text model for low-latency transcript deltas from live audio
0
o1
O-series reasoning model for hard analysis, math, coding, and planning
200,000 ctx
reasoning
o3
Deliberate o-series reasoner for hard math, coding, and multi-step analysis
200,000 ctx
reasoning
o3-deep-research
Research model for long-horizon investigation, synthesis, and analytical reports
200,000 ctx
reasoning
o3-mini
Smaller o-series reasoner for economical coding, math, and planning tasks
200,000 ctx
reasoning
o3 Pro
High-effort o3 tier for difficult technical reasoning and careful answers
200,000 ctx
reasoning
o4-mini
Fast o-series model for compact reasoning, coding, and tool use
200,000 ctx
reasoning
text-embedding-3-large
Embedding model for semantic search, retrieval, clustering, and ranking pipelines
8,192 ctx
text-embedding-3-small
Embedding model for semantic search, retrieval, clustering, and ranking pipelines
8,192 ctx
text-embedding-ada-002
Embedding model for semantic search, retrieval, clustering, and ranking pipelines
8,192 ctx
TTS-1
Speech generation model for controllable voice, narration, and audio delivery
0
TTS-1 HD
Speech generation model for controllable voice, narration, and audio delivery
0
Whisper
Speech transcription model for accurate audio-to-text and captioning workflows
0
Sonar
Sonar search model for current answers, retrieval, and citation-backed chat
127,000 ctx
Sonar Pro
Advanced Sonar search model for deeper research and cited synthesis
200,000 ctx
Sonar Reasoning Pro
Web-grounded reasoning model for multi-step research and cited answers
127,000 ctx
reasoning
Laguna S 2.1
Agentic coding model from Poolside in the XS size class for local deployment
1,000,000 ctx
reasoning
Laguna S 2.1 Free
Free provider route for experiments, demos, and cost-sensitive chat workloads
256,000 ctx
reasoning
Flux Schnell
Image model for prompt-driven generation, editing, and visual design workflows
512 ctx
Arrow 1.1
Image model for prompt-driven generation, editing, and visual design workflows
131,072 ctx
Recraft V2
Image model for prompt-driven generation, editing, and visual design workflows
512 ctx
Recraft V3
Image model for prompt-driven generation, editing, and visual design workflows
512 ctx
Recraft V4
Image model for prompt-driven generation, editing, and visual design workflows
0
Recraft V4 Pro
Image model for prompt-driven generation, editing, and visual design workflows
0
Recraft V4.1
Image model for prompt-driven generation, editing, and visual design workflows
0
Recraft V4.1 Pro
Image model for prompt-driven generation, editing, and visual design workflows
0
Recraft V4.1 Utility
Image model for prompt-driven generation, editing, and visual design workflows
0
Recraft V4.1 Utility Pro
Image model for prompt-driven generation, editing, and visual design workflows
0
Fugu Ultra
Quality-first multi-agent model for hard research, analysis, and competitions
1,000,000 ctx
reasoning
StepFun 3.5 Flash
StepFun flash lane for quick multimodal reasoning and coding assistance
262,114 ctx
reasoning
Step 3.7 Flash
Newer StepFun flash model for faster agents, coding, and multimodal prompts
256,000 ctx
reasoning
Hy3
Tencent Hy reasoning model for coding, instruction following, and agent tasks
262,144 ctx
reasoning
Inkling
Multimodal MoE reasoning model (975B total, 41B active) for text, image, and audio
256,000 ctx
reasoning
Voyage Rerank 2.5
Reranking model for improving retrieval quality in search and recommendation systems
32,000 ctx
Voyage Rerank 2.5 Lite
Reranking model for improving retrieval quality in search and recommendation systems
32,000 ctx
voyage-3-large
Flagship model for demanding analysis, coding, and production agent workflows
8,192 ctx
voyage-3.5
General-purpose chat model for instruction following, writing, and analysis
8,192 ctx
voyage-3.5-lite
Efficient model for low-latency assistance, extraction, and routine automation
8,192 ctx
voyage-4
General-purpose chat model for instruction following, writing, and analysis
32,000 ctx
voyage-4-large
Flagship model for demanding analysis, coding, and production agent workflows
32,000 ctx
voyage-4-lite
Efficient model for low-latency assistance, extraction, and routine automation
32,000 ctx
voyage-code-2
Coding model for repository understanding, refactors, and agentic engineering tasks
8,192 ctx
voyage-code-3
Coding model for repository understanding, refactors, and agentic engineering tasks
8,192 ctx
voyage-finance-2
General-purpose chat model for instruction following, writing, and analysis
8,192 ctx
voyage-law-2
General-purpose chat model for instruction following, writing, and analysis
8,192 ctx
Grok 4.1 Fast Non-Reasoning
Fast Grok model for responsive chat, reasoning, and tool-assisted work
1,000,000 ctx
Grok 4.1 Fast Reasoning
Fast Grok model for responsive chat, reasoning, and tool-assisted work
1,000,000 ctx
reasoning
Grok 4.20 Multi-Agent
Grok model for agentic tool use, reasoning, coding, and live assistance
2,000,000 ctx
reasoning
Grok 4.20 Multi Agent Beta
Grok model for agentic tool use, reasoning, coding, and live assistance
2,000,000 ctx
reasoning
Grok 4.20 Non-Reasoning
Grok model for agentic tool use, reasoning, coding, and live assistance
2,000,000 ctx
Grok 4.20 Beta Non-Reasoning
Grok model for agentic tool use, reasoning, coding, and live assistance
2,000,000 ctx
Grok 4.20 Reasoning
Grok model for agentic tool use, reasoning, coding, and live assistance
2,000,000 ctx
reasoning
Grok 4.20 Beta Reasoning
Grok model for agentic tool use, reasoning, coding, and live assistance
2,000,000 ctx
reasoning
Grok 4.3
xAI's default Grok for chat, coding, agentic tools, and lower hallucination risk
1,000,000 ctx
reasoning
Grok 4.5
Grok model for agentic tool use, reasoning, coding, and live assistance
500,000 ctx
reasoning
Grok Build 0.1
Fast Grok coding model tuned for agentic engineering and iterative edits
256,000 ctx
reasoning
Grok Imagine Image
Image model for prompt-driven generation, editing, and visual design workflows
0
Grok Imagine
Image model for prompt-driven generation, editing, and visual design workflows
0
Grok Imagine Video 1.5
Image model for prompt-driven generation, editing, and visual design workflows
0
Grok Imagine Video 1.5 Preview
Image model for prompt-driven generation, editing, and visual design workflows
0
Grok STT
Speech transcription model for accurate audio-to-text and captioning workflows
0
Grok TTS
Speech generation model for controllable voice, narration, and audio delivery
0
Grok Voice Think Fast 1.0
Speech generation model for controllable voice, narration, and audio delivery
0
MiMo M2.5
Open MiMo model for multimodal coding agents and long-context automation
1,050,000 ctx
reasoning
MiMo V2.5 Pro
Stronger MiMo Pro tier for multimodal reasoning and coding-agent execution
1,050,000 ctx
reasoning
GLM 4.5
Hybrid-reasoning GLM release that made the 4.5 line broadly useful
128,000 ctx
reasoning
GLM 4.5 Air
Lighter GLM-4.5 variant for fast coding assistance and cheaper agents
128,000 ctx
reasoning
GLM 4.5V
GLM vision model for visual reasoning, documents, and multimodal agents
66,000 ctx
reasoning
GLM 4.6
Late GLM-4 workhorse for coding agents, reasoning, and structured tasks
200,000 ctx
reasoning
GLM-4.6V
GLM vision model for visual reasoning, documents, and multimodal agents
128,000 ctx
reasoning
GLM-4.6V-Flash
GLM vision model for visual reasoning, documents, and multimodal agents
128,000 ctx
reasoning
GLM 4.7
Mature GLM model for dependable coding, reasoning, and structured agent tasks
200,000 ctx
reasoning
GLM 4.7 Flash
Budget GLM lane for fast coding help, routing, and everyday automation
200,000 ctx
reasoning
GLM 4.7 FlashX
Efficient GLM model for fast reasoning, coding, and agent workflows
200,000 ctx
reasoning
GLM-5
General GLM flagship for coding, analysis, and tool-heavy engineering workflows
202,800 ctx
reasoning
GLM 5 Turbo
Faster GLM-5 lane for coding agents that need lower latency
202,800 ctx
reasoning
GLM 5.1
Strong GLM coding model for agentic engineering, terminals, and repository generation
202,000 ctx
reasoning
GLM 5.2
Open flagship GLM for long-horizon coding agents and million-token context work
1,040,000 ctx
reasoning
GLM 5.2 Fast
Efficient GLM model for fast reasoning, coding, and agent workflows
1,000,000 ctx
reasoning
GLM 5V Turbo
Fast GLM vision model for screenshots, documents, and multimodal agent tasks
200,000 ctx
reasoning
Comments (0)
Sign in
to leave a comment
No comments yet.
Community Rating
—
No ratings yet
Sign in
to rate