Viberank
Trending models
Popular Labs
Hottest providers
Top Agents
Login
Sign up
← All providers
LLM Gateway
Provider ID: llmgateway
Details
SDK package
@ai-sdk/openai-compatible
Models
183
API
https://api.llmgateway.io/v1
Links
Documentation
Auth: LLMGATEWAY_API_KEY
Models (183)
Auto Route
Automatic model router for matching prompts to suitable backends and budgets
128,000 ctx
Claude 3 Opus
Flagship Claude model for deep reasoning, coding, and long-horizon agents
200,000 ctx
Claude Fable 5
Claude model for creative writing, analysis, and controlled agent workflows
1,000,000 ctx
reasoning
Claude Haiku 4.5 (latest)
Fast Claude lane for lightweight agents, office tasks, and responsive chat
200,000 ctx
reasoning
Claude Haiku 4.5
Fast Claude model for responsive assistance, classification, and lightweight agents
200,000 ctx
reasoning
Claude Haiku 4.5 (latest)
Fast Claude lane for lightweight agents, office tasks, and responsive chat
200,000 ctx
Claude Opus 4.1
Flagship Claude model for deep reasoning, coding, and long-horizon agents
200,000 ctx
reasoning
Claude Opus 4.5
Flagship Claude model for deep reasoning, coding, and long-horizon agents
200,000 ctx
reasoning
Claude Opus 4.6
High-end Claude for difficult coding, planning, and slower expert reasoning
1,000,000 ctx
reasoning
Claude Opus 4.7
Stronger Opus tier for advanced software work and high-stakes reasoning
1,000,000 ctx
reasoning
Claude Opus 4.8
Top Claude Opus tier for the hardest reasoning, coding, and long-horizon agents
1,000,000 ctx
reasoning
Claude Opus 5
Strongest Claude Opus model for coding, agents, and professional work
1,000,000 ctx
reasoning
Claude Sonnet 4.5 (latest)
Balanced Claude model for coding, analysis, agent workflows, and cost control
200,000 ctx
reasoning
Claude Sonnet 4.5
Balanced Claude model for coding, analysis, agent workflows, and cost control
200,000 ctx
reasoning
Claude Sonnet 4.6
Claude workhorse for coding agents, careful analysis, and production cost control
1,000,000 ctx
reasoning
Claude Sonnet 5
Everyday Claude agent model for coding, planning, browsing, and general work
1,000,000 ctx
reasoning
Codestral
Mistral coding model for code completion, generation, and developer workflows
256,000 ctx
Cosmos 3 Super Reasoner
Multimodal reasoning model for visual analysis, planning, and tool use
262,144 ctx
reasoning
Custom Model
Automatic model router for matching prompts to suitable backends and budgets
128,000 ctx
DeepSeek V3.2
DeepSeek chat model for instruction following, coding, and analysis
163,840 ctx
reasoning
DeepSeek V4 Flash
Fast DeepSeek V4 lane for economical reasoning, coding, and long-context work
1,050,000 ctx
reasoning
DeepSeek V4 Pro
Open MoE flagship with million-token context for coding and long agent runs
1,050,000 ctx
reasoning
Devstral 2
Mistral's coding-agent model for repository work, terminal tasks, and software fixes
262,144 ctx
Fugu Ultra
Quality-first multi-agent model for hard research, analysis, and competitions
1,000,000 ctx
reasoning
Gemini 2.5 Flash
Fast Gemini workhorse for multimodal apps where latency and price matter
1,048,576 ctx
reasoning
Gemini 2.5 Flash-Lite
Lean Gemini 2.5 lane for cheap multimodal traffic and quick agents
1,048,576 ctx
reasoning
Gemini 2.5 Pro
Google's proven reasoning model for coding, math, and multimodal analysis
1,048,576 ctx
reasoning
Gemini 3 Flash Preview
New Gemini flash lane bringing frontier-style multimodal reasoning to cheaper runs
1,048,576 ctx
reasoning
Gemini 3.1 Flash Lite
Low-latency Gemini model for high-volume multimodal and agent workloads
1,048,576 ctx
reasoning
Gemini 3.1 Pro Preview
Reasoning-first Gemini preview for agentic coding and complex problem solving
1,048,576 ctx
reasoning
Gemini 3.5 Flash
Fast Gemini model balancing multimodal reasoning, tool use, and cost
1,048,576 ctx
reasoning
Gemini 3.5 Flash Lite
Fast Gemini model balancing multimodal reasoning, tool use, and cost
1,048,576 ctx
reasoning
Gemini 3.6 Flash
Fast Gemini model balancing multimodal reasoning, tool use, and cost
1,048,576 ctx
reasoning
Gemini Pro Latest
Advanced Gemini model for complex reasoning, coding, and multimodal analysis
1,048,576 ctx
reasoning
Gemma 3 27B
Open Gemma instruction model for efficient chat and self-hosted deployments
110,000 ctx
Gemma 4 26B A4B IT
Open Gemma instruction model for efficient chat and self-hosted deployments
262,144 ctx
reasoning
Gemma 4 31B IT
Largest Gemma 4 instruction model for open, self-hosted chat and reasoning
262,144 ctx
reasoning
GLM-4 32B (0414-128k)
Flagship GLM model for hybrid reasoning, coding, and agentic engineering
128,000 ctx
GLM-4.5
Hybrid-reasoning GLM release that made the 4.5 line broadly useful
131,000 ctx
reasoning
GLM-4.5-Air
Lighter GLM-4.5 variant for fast coding assistance and cheaper agents
131,000 ctx
reasoning
GLM-4.5 AirX
Flagship GLM model for hybrid reasoning, coding, and agentic engineering
128,000 ctx
GLM-4.5 X
Flagship GLM model for hybrid reasoning, coding, and agentic engineering
128,000 ctx
reasoning
GLM-4.5V
GLM vision model for visual reasoning, documents, and multimodal agents
128,000 ctx
reasoning
GLM-4.6
Late GLM-4 workhorse for coding agents, reasoning, and structured tasks
204,800 ctx
reasoning
GLM-4.6V
GLM vision model for visual reasoning, documents, and multimodal agents
131,072 ctx
reasoning
GLM-4.6V FlashX
GLM vision model for visual reasoning, documents, and multimodal agents
128,000 ctx
reasoning
GLM-4.7
Mature GLM model for dependable coding, reasoning, and structured agent tasks
204,800 ctx
reasoning
GLM-4.7-Flash
Budget GLM lane for fast coding help, routing, and everyday automation
200,000 ctx
reasoning
GLM-4.7-FlashX
Efficient GLM model for fast reasoning, coding, and agent workflows
200,000 ctx
reasoning
GLM-5
General GLM flagship for coding, analysis, and tool-heavy engineering workflows
203,000 ctx
reasoning
GLM-5.1
Strong GLM coding model for agentic engineering, terminals, and repository generation
204,800 ctx
reasoning
GLM-5.2
Open flagship GLM for long-horizon coding agents and million-token context work
1,048,576 ctx
reasoning
GPT-3.5-turbo
Compact GPT model for low-latency assistance and high-volume workloads
16,385 ctx
GPT-4
GPT model for general reasoning, writing, coding, and tool-assisted tasks
8,192 ctx
GPT-4 Turbo
Compact GPT model for low-latency assistance and high-volume workloads
128,000 ctx
GPT-4.1
Long-lived GPT workhorse for coding, instruction following, and production apps
1,000,000 ctx
GPT-4.1 mini
Affordable GPT-4.1 lane for fast coding help and structured extraction
1,000,000 ctx
GPT-4.1 nano
Tiny GPT-4.1 option for classification, routing, and very high-volume tasks
1,000,000 ctx
GPT-4o
Omni-era GPT for multimodal chat, practical coding, and general assistants
128,000 ctx
GPT-4o mini
Small omni GPT for cheap multimodal assistance and production-scale traffic
128,000 ctx
GPT-4o Mini Search Preview
Compact GPT model for low-latency assistance and high-volume workloads
128,000 ctx
GPT-4o Search Preview
GPT model for general reasoning, writing, coding, and tool-assisted tasks
128,000 ctx
GPT-5
Original GPT-5 workhorse for reasoning, coding, writing, and tool workflows
400,000 ctx
reasoning
GPT-5 Chat (latest)
Chat-tuned GPT model for conversational assistance, writing, and tool workflows
400,000 ctx
reasoning
GPT-5 Mini
Small GPT-5 for responsive agents, coding help, and everyday automation
400,000 ctx
reasoning
GPT-5 Nano
Tiny GPT-5 lane for routing, extraction, classification, and bulk jobs
400,000 ctx
reasoning
GPT-5 Pro
Higher-accuracy GPT-5 tier for tough analysis, coding reviews, and planning
400,000 ctx
reasoning
GPT-5.1
Sharper GPT-5 generation for coding, product work, and tool-assisted tasks
400,000 ctx
reasoning
GPT-5.1 Codex
Codex GPT for repository edits, code review, and practical software agents
400,000 ctx
reasoning
GPT-5.1 Codex mini
Coding-optimized GPT model for repository edits, reviews, and agentic software work
400,000 ctx
reasoning
GPT-5.2
Reliable GPT generation for broad coding, writing, and tool-assisted product work
400,000 ctx
reasoning
GPT-5.2 Chat
Chat-tuned GPT model for conversational assistance, writing, and tool workflows
128,000 ctx
reasoning
GPT-5.2 Codex
Code-specialist GPT for repository edits, reviews, and long-running software agents
400,000 ctx
reasoning
GPT-5.2 Pro
Higher-accuracy GPT-5.2 variant for tougher reasoning and review workflows
400,000 ctx
reasoning
GPT-5.3 Chat (latest)
Chat-tuned GPT model for conversational assistance, writing, and tool workflows
128,000 ctx
GPT-5.3 Codex
Coding-optimized GPT model for repository edits, reviews, and agentic software work
400,000 ctx
reasoning
GPT-5.4
Agent-ready GPT for coding and computer-use workflows at a lower cost
1,050,000 ctx
reasoning
GPT-5.4 mini
Strong small GPT for coding subagents, quick tool use, and high-volume work
400,000 ctx
reasoning
GPT-5.4 nano
Cheapest GPT-5.4 lane for simple routing, extraction, and bulk automation
400,000 ctx
reasoning
GPT-5.4 Pro
More exact GPT-5.4 tier for demanding professional reasoning and agent tasks
1,050,000 ctx
reasoning
GPT-5.5
Default frontier GPT for coding, computer use, research, and knowledge work
1,050,000 ctx
reasoning
GPT-5.5 Pro
Highest-accuracy GPT-5.5 tier for slower, precision-heavy reasoning and coding
1,050,000 ctx
reasoning
GPT-5.6 Luna
Cost-efficient GPT-5.6 model for fast, high-volume workloads
1,050,000 ctx
reasoning
GPT-5.6 Sol
Frontier GPT-5.6 model for complex professional work, coding, and agentic workflows
1,050,000 ctx
reasoning
GPT-5.6 Terra
Balanced GPT-5.6 model for capable, cost-efficient everyday work
1,050,000 ctx
reasoning
GPT OSS 120B
Open-weight GPT model for self-hosted reasoning and instruction-following workloads
131,072 ctx
reasoning
GPT OSS 20B
Open-weight GPT model for self-hosted reasoning and instruction-following workloads
131,072 ctx
reasoning
Grok 4
Grok model for agentic tool use, reasoning, coding, and live assistance
256,000 ctx
Grok 4.1 Fast Non-Reasoning
Fast Grok model for responsive chat, reasoning, and tool-assisted work
2,000,000 ctx
Grok 4.1 Fast Reasoning
Fast Grok model for responsive chat, reasoning, and tool-assisted work
2,000,000 ctx
reasoning
Grok 4.20 (Non-Reasoning)
Grok model for agentic tool use, reasoning, coding, and live assistance
2,000,000 ctx
Grok 4.20 (Reasoning)
Reasoning Grok for document-heavy analysis and long-horizon tool use
2,000,000 ctx
reasoning
Grok 4.20 (Non-Reasoning)
Grok model for agentic tool use, reasoning, coding, and live assistance
2,000,000 ctx
Grok 4.20 (Reasoning)
Reasoning Grok for document-heavy analysis and long-horizon tool use
2,000,000 ctx
reasoning
Grok 4.3
xAI's default Grok for chat, coding, agentic tools, and lower hallucination risk
1,000,000 ctx
reasoning
Grok 4.5
Grok model for agentic tool use, reasoning, coding, and live assistance
500,000 ctx
reasoning
Grok Build 0.1
Fast Grok coding model tuned for agentic engineering and iterative edits
256,000 ctx
reasoning
Hermes 4 405B
Reasoning model for deliberate analysis, multi-step problem solving, and tool use
131,072 ctx
reasoning
Hermes 4 70B
Reasoning model for deliberate analysis, multi-step problem solving, and tool use
131,072 ctx
reasoning
Kimi K2
Kimi model for long-context chat, coding, and agentic reasoning
256,000 ctx
Kimi K2 Thinking
Thinking Kimi model for slower research passes, planning, and hard technical questions
262,144 ctx
reasoning
Kimi K2.5
Earlier Kimi frontier model for long-context agents, coding, and multimodal work
262,144 ctx
reasoning
Kimi K2.6
Multimodal Kimi workhorse for agent loops, coding tasks, and visual context
262,144 ctx
reasoning
Kimi K2.7 Code
Coding-focused Kimi model, stronger on long-horizon repo work with less overthinking
262,144 ctx
reasoning
Kimi K2.7 Code Highspeed
Lower-latency Kimi Code variant for interactive edits and coding-agent loops
262,144 ctx
reasoning
Kimi K3
Multimodal Kimi model with 1M context and toggleable max-effort thinking for long-horizon agent work
1,048,576 ctx
reasoning
Llama 3 70B Instruct
Open Llama instruction model for multilingual chat, reasoning, and coding
8,192 ctx
Llama 3.1 70B Instruct
Open Llama instruction model for multilingual chat, reasoning, and coding
128,000 ctx
Llama 3.1 Nemotron Ultra 253B
Flagship Nemotron model for high-throughput reasoning and complex agents
128,000 ctx
Llama 3.2 11B Instruct
Open Llama instruction model for multilingual chat, reasoning, and coding
128,000 ctx
Llama 3.2 3B Instruct
Open Llama instruction model for multilingual chat, reasoning, and coding
32,768 ctx
Llama-3.3-70B-Instruct
Popular open Llama workhorse for multilingual chat, coding, and self-hosting
131,072 ctx
Llama 4 Maverick 17B Instruct
Open multimodal Llama model for strong reasoning and fast responses
1,048,576 ctx
Llama 4 Scout 17B Instruct
Open multimodal Llama model for long-context analysis and efficient agents
131,072 ctx
MiMo-V2.5
Open MiMo model for multimodal coding agents and long-context automation
1,000,000 ctx
reasoning
MiMo-V2.5-Pro
Stronger MiMo Pro tier for multimodal reasoning and coding-agent execution
1,000,000 ctx
reasoning
MiniCPM-V 4.5
Multimodal model for analyzing text, images, documents, and rich media
32,000 ctx
MiniMax-M2
Efficient open MiniMax model built for coding agents and tool-heavy workflows
196,608 ctx
reasoning
MiniMax-M2.1
Earlier MiniMax agent model for practical coding and productivity tasks
204,800 ctx
reasoning
MiniMax M2.1 Lightning
High-speed MiniMax model for low-latency coding and agent workflows
196,608 ctx
reasoning
MiniMax-M2.5
Prior MiniMax coding model for agent workflows, office edits, and automation
228,700 ctx
reasoning
MiniMax-M2.5-highspeed
High-speed MiniMax model for low-latency coding and agent workflows
204,800 ctx
reasoning
MiniMax-M2.7
Open MiniMax flagship for coding agents, office automation, and complex environments
204,800 ctx
reasoning
MiniMax-M2.7-highspeed
Low-latency M2.7 variant for interactive coding plans and agent loops
204,800 ctx
reasoning
MiniMax-M3
MiniMax multimodal model for long-context coding, perception, and agent planning
1,048,576 ctx
reasoning
MiniMax Text 01
MiniMax model for chat, coding, office work, and agentic tasks
1,000,000 ctx
reasoning
Ministral 14B
Compact Mistral model for edge, latency-sensitive, and cost-efficient workloads
262,144 ctx
Ministral 3B
Compact Mistral model for edge, latency-sensitive, and cost-efficient workloads
131,072 ctx
Ministral 8B
Compact Mistral model for edge, latency-sensitive, and cost-efficient workloads
262,144 ctx
Mistral Large 3
Mistral's largest general model for enterprise agents, coding, and multilingual reasoning
262,144 ctx
Mistral Large (latest)
Flagship Mistral model for advanced reasoning, coding, and multilingual work
128,000 ctx
Mistral Small 3.2
Efficient Mistral model for fast chat, extraction, and production assistants
128,000 ctx
Muse Spark 1.1
Muse Spark is a natively multimodal reasoning model with support for tool-use, visual chain of thought, and multi-agent orchestration.
1,048,576 ctx
reasoning
Nemotron 3 Nano 30B
Compact Nemotron model for efficient reasoning and deployable AI agents
262,144 ctx
reasoning
Nemotron 3 Nano Omni
Omni-modal model for text, vision, audio, and multimodal agent tasks
262,144 ctx
reasoning
Nemotron 3 Super 120B
Nemotron model for efficient reasoning, coding, and specialized AI agents
262,144 ctx
reasoning
Nemotron 3 Ultra 550B A55B
Largest Nemotron 3 model for maximum open-weight reasoning and agent accuracy
1,048,576 ctx
reasoning
o1
O-series reasoning model for hard analysis, math, coding, and planning
200,000 ctx
reasoning
o3
Deliberate o-series reasoner for hard math, coding, and multi-step analysis
200,000 ctx
reasoning
o3-mini
Smaller o-series reasoner for economical coding, math, and planning tasks
200,000 ctx
reasoning
o4-mini
Fast o-series model for compact reasoning, coding, and tool use
200,000 ctx
reasoning
Qwen Coder Plus
Qwen coding model for software agents, repository edits, and code reasoning
131,072 ctx
Qwen Flash
Efficient Qwen model for fast chat, extraction, and high-volume workloads
1,000,000 ctx
reasoning
Qwen Max
Flagship Qwen model for complex reasoning, coding, and agentic workflows
32,768 ctx
Qwen Max Latest
Qwen vision-language model for visual reasoning, documents, and agent tasks
32,768 ctx
Qwen-Omni Turbo
Qwen omni model for text, vision, audio, and multimodal agent tasks
32,768 ctx
Qwen Plus
Qwen instruction model for multilingual chat, reasoning, and tool use
131,072 ctx
reasoning
Qwen Plus Latest
Qwen vision-language model for visual reasoning, documents, and agent tasks
1,000,000 ctx
Qwen2.5 VL 32B Instruct
Qwen vision-language model for visual reasoning, documents, and agent tasks
131,072 ctx
Qwen2.5-VL 72B Instruct
Qwen vision-language model for visual reasoning, documents, and agent tasks
32,000 ctx
Qwen3 235B A22B FP8
Qwen instruction model for multilingual chat, reasoning, and tool use
40,960 ctx
reasoning
Qwen3 235B A22B Instruct (2507)
Qwen instruction model for multilingual chat, reasoning, and tool use
262,144 ctx
Qwen3 235B A22B Thinking (2507)
Qwen reasoning model for deliberate problem solving, math, and coding
262,000 ctx
reasoning
Qwen3 30B A3B Instruct (2507)
Qwen instruction model for multilingual chat, reasoning, and tool use
262,000 ctx
Qwen3 32B
Dense open Qwen model for self-hosted chat, reasoning, and coding
40,960 ctx
reasoning
Qwen3-Coder 30B-A3B Instruct
Smaller Qwen coder for efficient local agents and repo-level fixes
262,000 ctx
Qwen3-Coder 480B-A35B Instruct
Open Qwen coding heavyweight for repository reasoning and agentic engineering
262,144 ctx
Qwen3 Coder Flash
Qwen coding model for software agents, repository edits, and code reasoning
1,000,000 ctx
Qwen3 Coder Next
Qwen coding model for software agents, repository edits, and code reasoning
262,144 ctx
reasoning
Qwen3 Coder Plus
Hosted Qwen coder for software agents, repo edits, and long-context code
1,000,000 ctx
Qwen3 Max
Flagship Qwen3 model for coding agents, complex reasoning, and tool use
262,144 ctx
Qwen3-Next 80B-A3B Instruct
Qwen instruction model for multilingual chat, reasoning, and tool use
131,072 ctx
Qwen3-Next 80B-A3B (Thinking)
Efficient Qwen thinking model for local reasoning, math, and coding agents
131,072 ctx
reasoning
Qwen3 VL 235B A22B Instruct
Qwen vision-language model for visual reasoning, documents, and agent tasks
131,072 ctx
Qwen3 VL 235B A22B Thinking
Qwen vision-language model for visual reasoning, documents, and agent tasks
131,072 ctx
reasoning
Qwen3 VL 30B A3B Instruct
Qwen vision-language model for visual reasoning, documents, and agent tasks
131,072 ctx
Qwen3 VL Flash
Qwen vision-language model for visual reasoning, documents, and agent tasks
262,144 ctx
Qwen3-VL Plus
Qwen vision-language model for visual reasoning, documents, and agent tasks
262,144 ctx
reasoning
Qwen3.5 9B
Qwen instruction model for multilingual chat, reasoning, and tool use
262,144 ctx
reasoning
Qwen3.6 35B-A3B
Open multimodal Qwen MoE for local agents that need vision, audio, and code
262,144 ctx
reasoning
Qwen3.6 Flash
Qwen vision-language model for visual reasoning, documents, and agent tasks
262,144 ctx
reasoning
Qwen3.6 Max Preview
Flagship Qwen model for complex reasoning, coding, and agentic workflows
262,144 ctx
reasoning
Qwen3.6 Plus
Earlier Qwen multimodal workhorse for million-token agent and document tasks
262,144 ctx
reasoning
Qwen3.7 Max
Qwen frontier model tuned for agent frameworks, coding assistants, and long tasks
1,000,000 ctx
reasoning
Qwen3.7 Plus
Multimodal Qwen workhorse for long-context agents, visual inputs, and coding
1,000,000 ctx
reasoning
Qwen3.5 397B-A17B
Large open Qwen multimodal MoE for visual agents and long technical tasks
262,144 ctx
reasoning
Seed 1.6 (250615)
Multimodal reasoning model for visual analysis, planning, and tool use
256,000 ctx
reasoning
Seed 1.6 (250915)
Multimodal reasoning model for visual analysis, planning, and tool use
256,000 ctx
reasoning
Seed 1.6 Flash (250715)
Multimodal reasoning model for visual analysis, planning, and tool use
256,000 ctx
reasoning
Seed 1.8 (251228)
Multimodal reasoning model for visual analysis, planning, and tool use
256,000 ctx
reasoning
Sonar
Fast web-grounded Sonar for current answers, citations, and lightweight retrieval
130,000 ctx
Sonar Pro
Deeper Sonar search model with broader retrieval and stronger synthesis
200,000 ctx
Sonar Reasoning Pro
Web-grounded Sonar for multi-step research questions that need cited reasoning
128,000 ctx
reasoning
Comments (0)
Sign in
to leave a comment
No comments yet.
Community Rating
—
No ratings yet
Sign in
to rate