Viberank
Trending models
Popular Labs
Hottest providers
Top Agents
Login
Sign up
← All providers
Kilo Gateway
Provider ID: kilo
Details
SDK package
@ai-sdk/openai-compatible
Models
346
API
https://api.kilo.ai/api/gateway
Links
Documentation
Auth: KILO_API_KEY
Models (346)
AI21: Jamba Large 1.7
Flagship model for demanding analysis, coding, and production agent workflows
256,000 ctx
AionLabs: Aion-1.0
Reasoning model for deliberate analysis, multi-step problem solving, and tool use
131,072 ctx
reasoning
AionLabs: Aion-1.0-Mini
Efficient model for low-latency assistance, extraction, and routine automation
131,072 ctx
reasoning
AionLabs: Aion-2.0
Reasoning model for deliberate analysis, multi-step problem solving, and tool use
131,072 ctx
reasoning
AionLabs: Aion-RP 1.0 (8B)
Open Llama instruction model for multilingual chat, reasoning, and coding
32,768 ctx
AlfredPros: CodeLLaMa 7B Instruct Solidity
Open-weight instruction model for adaptable chat and self-hosted production workloads
4,096 ctx
AllenAI: Olmo 3 32B Think
Reasoning model for deliberate analysis, multi-step problem solving, and tool use
65,536 ctx
reasoning
Amazon: Nova 2 Lite
Multimodal reasoning model for visual analysis, planning, and tool use
1,000,000 ctx
reasoning
Amazon: Nova Lite 1.0
Efficient model for low-latency assistance, extraction, and routine automation
300,000 ctx
Amazon: Nova Micro 1.0
Efficient model for low-latency assistance, extraction, and routine automation
128,000 ctx
Amazon: Nova Premier 1.0
Flagship model for demanding analysis, coding, and production agent workflows
1,000,000 ctx
Amazon: Nova Pro 1.0
Flagship model for demanding analysis, coding, and production agent workflows
300,000 ctx
Magnum v4 72B
Open-weight instruction model for adaptable chat and self-hosted production workloads
16,384 ctx
Anthropic: Claude 3 Haiku
Fast Claude model for responsive assistance, classification, and lightweight agents
200,000 ctx
Anthropic: Claude 3.5 Haiku
Fast Claude model for responsive assistance, classification, and lightweight agents
200,000 ctx
Anthropic: Claude Haiku 4.5
Fast Claude model for responsive assistance, classification, and lightweight agents
200,000 ctx
reasoning
Anthropic: Claude Opus 4
Flagship Claude model for deep reasoning, coding, and long-horizon agents
200,000 ctx
reasoning
Anthropic: Claude Opus 4.1
Flagship Claude model for deep reasoning, coding, and long-horizon agents
200,000 ctx
reasoning
Anthropic: Claude Opus 4.5
Flagship Claude model for deep reasoning, coding, and long-horizon agents
200,000 ctx
reasoning
Anthropic: Claude Opus 4.6
Flagship Claude model for deep reasoning, coding, and long-horizon agents
1,000,000 ctx
reasoning
Anthropic: Claude Opus 4.6 (Fast)
Flagship Claude model for deep reasoning, coding, and long-horizon agents
1,000,000 ctx
reasoning
Anthropic: Claude Opus 4.7
Flagship Claude model for deep reasoning, coding, and long-horizon agents
1,000,000 ctx
reasoning
Anthropic: Claude Opus 4.7 (Fast)
Flagship Claude model for deep reasoning, coding, and long-horizon agents
1,000,000 ctx
reasoning
Anthropic: Claude Sonnet 4
Balanced Claude model for coding, analysis, agent workflows, and cost control
200,000 ctx
reasoning
Anthropic: Claude Sonnet 4.5
Balanced Claude model for coding, analysis, agent workflows, and cost control
1,000,000 ctx
reasoning
Anthropic: Claude Sonnet 4.6
Balanced Claude model for coding, analysis, agent workflows, and cost control
1,000,000 ctx
reasoning
Arcee AI: Coder Large
Coding model for repository understanding, refactors, and agentic engineering tasks
32,768 ctx
Arcee AI: Maestro Reasoning
Open-weight instruction model for adaptable chat and self-hosted production workloads
131,072 ctx
Arcee AI: Spotlight
Multimodal model for analyzing text, images, documents, and rich media
131,072 ctx
Arcee AI: Trinity Large Thinking
Flagship model for demanding analysis, coding, and production agent workflows
262,144 ctx
reasoning
Arcee AI: Trinity Mini
Efficient model for low-latency assistance, extraction, and routine automation
131,072 ctx
reasoning
Arcee AI: Virtuoso Large
Flagship model for demanding analysis, coding, and production agent workflows
131,072 ctx
Baidu: CoBuddy (free)
Free provider route for experiments, demos, and cost-sensitive chat workloads
131,072 ctx
reasoning
Baidu: ERNIE 4.5 21B A3B
Open-weight instruction model for adaptable chat and self-hosted production workloads
120,000 ctx
Baidu: ERNIE 4.5 21B A3B Thinking
Reasoning model for deliberate analysis, multi-step problem solving, and tool use
131,072 ctx
reasoning
Baidu: ERNIE 4.5 300B A47B
Open-weight instruction model for adaptable chat and self-hosted production workloads
123,000 ctx
Baidu: ERNIE 4.5 VL 28B A3B
Multimodal reasoning model for visual analysis, planning, and tool use
30,000 ctx
reasoning
Baidu: ERNIE 4.5 VL 424B A47B
Multimodal reasoning model for visual analysis, planning, and tool use
123,000 ctx
reasoning
Baidu: Qianfan-OCR-Fast
OCR model for extracting structured text from documents and screenshots
65,536 ctx
ByteDance Seed: Seed 1.6
Multimodal reasoning model for visual analysis, planning, and tool use
262,144 ctx
reasoning
ByteDance Seed: Seed 1.6 Flash
Multimodal reasoning model for visual analysis, planning, and tool use
262,144 ctx
reasoning
ByteDance Seed: Seed-2.0-Lite
Multimodal reasoning model for visual analysis, planning, and tool use
262,144 ctx
reasoning
ByteDance Seed: Seed-2.0-Mini
Multimodal reasoning model for visual analysis, planning, and tool use
262,144 ctx
reasoning
ByteDance: UI-TARS 7B
Multimodal model for analyzing text, images, documents, and rich media
128,000 ctx
Cohere: Command A
Cohere command model for multilingual enterprise agents, tools, and chat
256,000 ctx
Cohere: Command R (08-2024)
Cohere retrieval model for long-context chat and enterprise RAG workflows
128,000 ctx
Cohere: Command R+ (08-2024)
Cohere retrieval model for long-context chat and enterprise RAG workflows
128,000 ctx
Cohere: Command R7B (12-2024)
Cohere command model for multilingual enterprise agents, tools, and chat
128,000 ctx
Deep Cogito: Cogito v2.1 671B
Reasoning model for deliberate analysis, multi-step problem solving, and tool use
128,000 ctx
reasoning
DeepSeek: DeepSeek V3
DeepSeek chat model for instruction following, coding, and analysis
163,840 ctx
DeepSeek: DeepSeek V3 0324
DeepSeek chat model for instruction following, coding, and analysis
163,840 ctx
DeepSeek: DeepSeek V3.1
DeepSeek chat model for instruction following, coding, and analysis
32,768 ctx
reasoning
DeepSeek: R1
DeepSeek reasoning model for multi-step analysis, math, coding, and tools
64,000 ctx
reasoning
DeepSeek: R1 0528
DeepSeek reasoning model for multi-step analysis, math, coding, and tools
163,840 ctx
reasoning
DeepSeek: R1 Distill Llama 70B
DeepSeek reasoning model for multi-step analysis, math, coding, and tools
131,072 ctx
reasoning
DeepSeek: R1 Distill Qwen 32B
Qwen instruction model for multilingual chat, reasoning, and tool use
32,768 ctx
reasoning
DeepSeek: DeepSeek V3.1 Terminus
DeepSeek chat model for instruction following, coding, and analysis
163,840 ctx
reasoning
DeepSeek: DeepSeek V3.2
DeepSeek chat model for instruction following, coding, and analysis
163,840 ctx
reasoning
DeepSeek: DeepSeek V3.2 Exp
DeepSeek chat model for instruction following, coding, and analysis
163,840 ctx
reasoning
DeepSeek: DeepSeek V3.2 Speciale
DeepSeek chat model for instruction following, coding, and analysis
163,840 ctx
reasoning
DeepSeek: DeepSeek V4 Flash
Fast DeepSeek model for efficient chat, coding help, and agent loops
1,048,576 ctx
reasoning
DeepSeek: DeepSeek V4 Pro
Flagship DeepSeek model for coding, reasoning, and agentic work
1,048,576 ctx
reasoning
EssentialAI: Rnj 1 Instruct
Open-weight instruction model for adaptable chat and self-hosted production workloads
32,768 ctx
Google: Gemini 2.0 Flash
Fast Gemini model balancing multimodal reasoning, tool use, and cost
1,048,576 ctx
Google: Gemini 2.0 Flash Lite
Low-latency Gemini model for high-volume multimodal and agent workloads
1,048,576 ctx
Google: Gemini 2.5 Flash
Fast Gemini model balancing multimodal reasoning, tool use, and cost
1,048,576 ctx
reasoning
Google: Nano Banana (Gemini 2.5 Flash Image)
Image model for prompt-driven generation, editing, and visual design workflows
32,768 ctx
Google: Gemini 2.5 Flash Lite
Low-latency Gemini model for high-volume multimodal and agent workloads
1,048,576 ctx
reasoning
Google: Gemini 2.5 Flash Lite Preview 09-2025
Low-latency Gemini model for high-volume multimodal and agent workloads
1,048,576 ctx
reasoning
Google: Gemini 2.5 Pro
Advanced Gemini model for complex reasoning, coding, and multimodal analysis
1,048,576 ctx
reasoning
Google: Gemini 2.5 Pro Preview 06-05
Advanced Gemini model for complex reasoning, coding, and multimodal analysis
1,048,576 ctx
reasoning
Google: Gemini 2.5 Pro Preview 05-06
Advanced Gemini model for complex reasoning, coding, and multimodal analysis
1,048,576 ctx
reasoning
Google: Gemini 3 Flash Preview
Fast Gemini model balancing multimodal reasoning, tool use, and cost
1,048,576 ctx
reasoning
Google: Nano Banana Pro (Gemini 3 Pro Image Preview)
Image model for prompt-driven generation, editing, and visual design workflows
65,536 ctx
reasoning
Google: Nano Banana 2 (Gemini 3.1 Flash Image Preview)
Image model for prompt-driven generation, editing, and visual design workflows
65,536 ctx
reasoning
Google: Gemini 3.1 Flash Lite
Low-latency Gemini model for high-volume multimodal and agent workloads
1,048,576 ctx
reasoning
Google: Gemini 3.1 Flash Lite Preview
Low-latency Gemini model for high-volume multimodal and agent workloads
1,048,576 ctx
reasoning
Google: Gemini 3.1 Pro Preview
Advanced Gemini model for complex reasoning, coding, and multimodal analysis
1,048,576 ctx
reasoning
Google: Gemini 3.1 Pro Preview Custom Tools
Advanced Gemini model for complex reasoning, coding, and multimodal analysis
1,048,576 ctx
reasoning
Google: Gemini 3.5 Flash
Fast Gemini model balancing multimodal reasoning, tool use, and cost
1,048,576 ctx
reasoning
Google: Gemma 2 27B
Open Gemma instruction model for efficient chat and self-hosted deployments
8,192 ctx
Google: Gemma 3 12B
Open Gemma instruction model for efficient chat and self-hosted deployments
131,072 ctx
Google: Gemma 3 27B
Open Gemma instruction model for efficient chat and self-hosted deployments
128,000 ctx
Google: Gemma 3 4B
Open Gemma instruction model for efficient chat and self-hosted deployments
131,072 ctx
Google: Gemma 3n 4B
Open Gemma instruction model for efficient chat and self-hosted deployments
32,768 ctx
Google: Gemma 4 26B A4B
Open Gemma instruction model for efficient chat and self-hosted deployments
262,144 ctx
reasoning
Google: Gemma 4 31B
Open Gemma instruction model for efficient chat and self-hosted deployments
262,144 ctx
reasoning
Google: Lyria 3 Clip Preview
Speech generation model for controllable voice, narration, and audio delivery
1,048,576 ctx
Google: Lyria 3 Pro Preview
Speech generation model for controllable voice, narration, and audio delivery
1,048,576 ctx
MythoMax 13B
Open-weight instruction model for adaptable chat and self-hosted production workloads
4,096 ctx
IBM: Granite 4.0 Micro
Efficient model for low-latency assistance, extraction, and routine automation
131,000 ctx
IBM: Granite 4.1 8B
Tool-capable chat model for instruction following and agentic application workflows
131,072 ctx
Inception: Mercury 2
Reasoning model for deliberate analysis, multi-step problem solving, and tool use
128,000 ctx
reasoning
inclusionAI: Ling-2.6-1T
Tool-capable chat model for instruction following and agentic application workflows
262,144 ctx
inclusionAI: Ling-2.6 Flash
Efficient model for low-latency assistance, extraction, and routine automation
262,144 ctx
inclusionAI: Ring-2.6-1T
Reasoning model for deliberate analysis, multi-step problem solving, and tool use
262,144 ctx
reasoning
Inflection: Inflection 3 Pi
General-purpose chat model for instruction following, writing, and analysis
8,000 ctx
Inflection: Inflection 3 Productivity
General-purpose chat model for instruction following, writing, and analysis
8,000 ctx
Kilo Auto Balanced
Automatic model router for matching prompts to suitable backends and budgets
204,800 ctx
reasoning
Kilo Auto Free
Automatic model router for matching prompts to suitable backends and budgets
204,800 ctx
reasoning
Kilo Auto Frontier
Automatic model router for matching prompts to suitable backends and budgets
1,000,000 ctx
reasoning
Kilo Auto Small
Automatic model router for matching prompts to suitable backends and budgets
400,000 ctx
reasoning
Kwaipilot: KAT-Coder-Pro V2
Coding model for repository understanding, refactors, and agentic engineering tasks
256,000 ctx
LiquidAI: LFM2-24B-A2B
Open-weight instruction model for adaptable chat and self-hosted production workloads
32,768 ctx
Mancer: Weaver (alpha)
General-purpose chat model for instruction following, writing, and analysis
8,000 ctx
Meta: Llama 3 70B Instruct
Open Llama instruction model for multilingual chat, reasoning, and coding
8,192 ctx
Meta: Llama 3 8B Instruct
Open Llama instruction model for multilingual chat, reasoning, and coding
8,192 ctx
Meta: Llama 3.1 70B Instruct
Open Llama instruction model for multilingual chat, reasoning, and coding
131,072 ctx
Meta: Llama 3.1 8B Instruct
Open Llama instruction model for multilingual chat, reasoning, and coding
16,384 ctx
Meta: Llama 3.2 11B Vision Instruct
Open Llama multimodal model for image understanding and text reasoning
131,072 ctx
Meta: Llama 3.2 1B Instruct
Open Llama instruction model for multilingual chat, reasoning, and coding
60,000 ctx
Meta: Llama 3.2 3B Instruct
Open Llama instruction model for multilingual chat, reasoning, and coding
80,000 ctx
Meta: Llama 3.3 70B Instruct
Open Llama instruction model for multilingual chat, reasoning, and coding
131,072 ctx
Meta: Llama 4 Maverick
Open multimodal Llama model for strong reasoning and fast responses
1,048,576 ctx
Meta: Llama 4 Scout
Open multimodal Llama model for long-context analysis and efficient agents
327,680 ctx
Llama Guard 3 8B
Safety model for policy screening, moderation, and risk-aware routing workflows
131,072 ctx
Meta: Llama Guard 4 12B
Safety model for policy screening, moderation, and risk-aware routing workflows
163,840 ctx
Microsoft: Phi 4
Open-weight instruction model for adaptable chat and self-hosted production workloads
16,384 ctx
Microsoft: Phi 4 Mini Instruct
Efficient model for low-latency assistance, extraction, and routine automation
128,000 ctx
WizardLM-2 8x22B
Open-weight instruction model for adaptable chat and self-hosted production workloads
65,535 ctx
MiniMax: MiniMax-01
MiniMax multimodal coding model for long-context reasoning and agent tasks
1,000,192 ctx
MiniMax: MiniMax M1
MiniMax model for chat, coding, office work, and agentic tasks
1,000,000 ctx
reasoning
MiniMax: MiniMax M2
MiniMax model for chat, coding, office work, and agentic tasks
196,608 ctx
reasoning
MiniMax: MiniMax M2-her
MiniMax model for chat, coding, office work, and agentic tasks
65,536 ctx
MiniMax: MiniMax M2.1
MiniMax model for chat, coding, office work, and agentic tasks
196,608 ctx
reasoning
MiniMax: MiniMax M2.5
MiniMax model for chat, coding, office work, and agentic tasks
196,608 ctx
reasoning
MiniMax: MiniMax M2.7
MiniMax model for chat, coding, office work, and agentic tasks
204,800 ctx
reasoning
MiniMax: MiniMax M3
MiniMax multimodal model for long-context coding, perception, and agent planning
1,048,576 ctx
reasoning
Mistral: Codestral 2508
Mistral coding model for code completion, generation, and developer workflows
256,000 ctx
Mistral: Devstral 2 2512
Mistral coding agent model for repository tasks and software engineering workflows
262,144 ctx
Mistral: Devstral Medium
Mistral coding agent model for repository tasks and software engineering workflows
131,072 ctx
Mistral: Devstral Small 1.1
Mistral coding agent model for repository tasks and software engineering workflows
131,072 ctx
Mistral: Ministral 3 14B 2512
Compact Mistral model for edge, latency-sensitive, and cost-efficient workloads
262,144 ctx
Mistral: Ministral 3 3B 2512
Compact Mistral model for edge, latency-sensitive, and cost-efficient workloads
131,072 ctx
Mistral: Ministral 3 8B 2512
Compact Mistral model for edge, latency-sensitive, and cost-efficient workloads
262,144 ctx
Mistral: Mistral 7B Instruct v0.1
Mistral model for multilingual chat, reasoning, and tool-assisted workflows
2,824 ctx
Mistral Large
Flagship Mistral model for advanced reasoning, coding, and multilingual work
128,000 ctx
Mistral Large 2407
Flagship Mistral model for advanced reasoning, coding, and multilingual work
131,072 ctx
Mistral Large 2411
Flagship Mistral model for advanced reasoning, coding, and multilingual work
131,072 ctx
Mistral: Mistral Large 3 2512
Flagship Mistral model for advanced reasoning, coding, and multilingual work
262,144 ctx
Mistral: Mistral Medium 3
Mistral model for multilingual chat, reasoning, and tool-assisted workflows
131,072 ctx
Mistral: Mistral Medium 3.5
Mistral model for multilingual chat, reasoning, and tool-assisted workflows
262,144 ctx
reasoning
Mistral: Mistral Medium 3.1
Mistral model for multilingual chat, reasoning, and tool-assisted workflows
131,072 ctx
Mistral: Mistral Nemo
Mistral model for multilingual chat, reasoning, and tool-assisted workflows
131,072 ctx
Mistral: Saba
Mistral model for multilingual chat, reasoning, and tool-assisted workflows
32,768 ctx
Mistral: Mistral Small 3
Efficient Mistral model for fast chat, extraction, and production assistants
32,768 ctx
Mistral: Mistral Small 4
Efficient Mistral model for fast chat, extraction, and production assistants
262,144 ctx
reasoning
Mistral: Mistral Small 3.1 24B
Efficient Mistral model for fast chat, extraction, and production assistants
128,000 ctx
Mistral: Mistral Small 3.2 24B
Efficient Mistral model for fast chat, extraction, and production assistants
131,072 ctx
Mistral: Mixtral 8x22B Instruct
Mistral model for multilingual chat, reasoning, and tool-assisted workflows
65,536 ctx
Mistral: Pixtral Large 2411
Mistral vision-language model for image understanding and multimodal chat
131,072 ctx
Mistral: Voxtral Small 24B 2507
Efficient Mistral model for fast chat, extraction, and production assistants
32,000 ctx
MoonshotAI: Kimi K2 0711
Kimi model for long-context chat, coding, and agentic reasoning
131,000 ctx
MoonshotAI: Kimi K2 0905
Kimi model for long-context chat, coding, and agentic reasoning
131,072 ctx
MoonshotAI: Kimi K2 Thinking
Kimi reasoning model for long-horizon research, planning, and tool use
131,072 ctx
reasoning
MoonshotAI: Kimi K2.5
Kimi multimodal agent model for visual understanding, coding, and planning
262,144 ctx
reasoning
MoonshotAI: Kimi K2.6
Kimi multimodal agent model for visual understanding, coding, and planning
262,144 ctx
reasoning
Morph: Morph V3 Fast
Efficient model for low-latency assistance, extraction, and routine automation
81,920 ctx
Morph: Morph V3 Large
Flagship model for demanding analysis, coding, and production agent workflows
262,144 ctx
Nex AGI: DeepSeek V3.1 Nex N1
DeepSeek chat model for instruction following, coding, and analysis
131,072 ctx
NousResearch: Hermes 2 Pro - Llama-3 8B
Open Llama instruction model for multilingual chat, reasoning, and coding
8,192 ctx
Nous: Hermes 3 405B Instruct
Open Llama instruction model for multilingual chat, reasoning, and coding
131,072 ctx
Nous: Hermes 3 70B Instruct
Open Llama instruction model for multilingual chat, reasoning, and coding
131,072 ctx
Nous: Hermes 4 405B
Reasoning model for deliberate analysis, multi-step problem solving, and tool use
131,072 ctx
reasoning
Nous: Hermes 4 70B
Reasoning model for deliberate analysis, multi-step problem solving, and tool use
131,072 ctx
reasoning
NVIDIA: Llama 3.3 Nemotron Super 49B V1.5
Nemotron model for efficient reasoning, coding, and specialized AI agents
131,072 ctx
reasoning
NVIDIA: Nemotron 3 Nano 30B A3B
Small Nemotron 3 MoE for efficient coding, math, and long-context agents
262,144 ctx
reasoning
NVIDIA: Nemotron 3 Nano Omni (free)
Open Nemotron omni model combining reasoning with text, vision, and audio
256,000 ctx
reasoning
NVIDIA: Nemotron 3 Super
Nemotron middle tier for collaborative agents and high-volume reasoning workloads
262,144 ctx
reasoning
NVIDIA: Nemotron 3 Super (free)
Nemotron middle tier for collaborative agents and high-volume reasoning workloads
262,144 ctx
reasoning
NVIDIA: Nemotron Nano 9B V2
Compact Nemotron model for efficient reasoning and deployable AI agents
131,072 ctx
reasoning
OpenAI: GPT-3.5 Turbo
Compact GPT model for low-latency assistance and high-volume workloads
16,385 ctx
OpenAI: GPT-3.5 Turbo (older v0613)
Compact GPT model for low-latency assistance and high-volume workloads
4,095 ctx
OpenAI: GPT-3.5 Turbo 16k
Compact GPT model for low-latency assistance and high-volume workloads
16,385 ctx
OpenAI: GPT-3.5 Turbo Instruct
Compact GPT model for low-latency assistance and high-volume workloads
4,095 ctx
OpenAI: GPT-4
GPT model for general reasoning, writing, coding, and tool-assisted tasks
8,191 ctx
OpenAI: GPT-4 (older v0314)
GPT model for general reasoning, writing, coding, and tool-assisted tasks
8,191 ctx
OpenAI: GPT-4 Turbo (older v1106)
Compact GPT model for low-latency assistance and high-volume workloads
128,000 ctx
OpenAI: GPT-4 Turbo
Compact GPT model for low-latency assistance and high-volume workloads
128,000 ctx
OpenAI: GPT-4 Turbo Preview
Compact GPT model for low-latency assistance and high-volume workloads
128,000 ctx
OpenAI: GPT-4.1
GPT model for general reasoning, writing, coding, and tool-assisted tasks
1,047,576 ctx
OpenAI: GPT-4.1 Mini
Compact GPT model for low-latency assistance and high-volume workloads
1,047,576 ctx
OpenAI: GPT-4.1 Nano
Compact GPT model for low-latency assistance and high-volume workloads
1,047,576 ctx
OpenAI: GPT-4o
GPT model for general reasoning, writing, coding, and tool-assisted tasks
128,000 ctx
OpenAI: GPT-4o (2024-05-13)
GPT model for general reasoning, writing, coding, and tool-assisted tasks
128,000 ctx
OpenAI: GPT-4o (2024-08-06)
GPT model for general reasoning, writing, coding, and tool-assisted tasks
128,000 ctx
OpenAI: GPT-4o (2024-11-20)
GPT model for general reasoning, writing, coding, and tool-assisted tasks
128,000 ctx
OpenAI: GPT-4o Audio
Speech generation model for controllable voice, narration, and audio delivery
128,000 ctx
OpenAI: GPT-4o-mini
Compact GPT model for low-latency assistance and high-volume workloads
128,000 ctx
OpenAI: GPT-4o-mini (2024-07-18)
Compact GPT model for low-latency assistance and high-volume workloads
128,000 ctx
OpenAI: GPT-4o-mini Search Preview
Compact GPT model for low-latency assistance and high-volume workloads
128,000 ctx
OpenAI: GPT-4o Search Preview
GPT model for general reasoning, writing, coding, and tool-assisted tasks
128,000 ctx
OpenAI: GPT-5
GPT model for general reasoning, writing, coding, and tool-assisted tasks
400,000 ctx
reasoning
OpenAI: GPT-5 Chat
Chat-tuned GPT model for conversational assistance, writing, and tool workflows
128,000 ctx
OpenAI: GPT-5 Codex
Coding-optimized GPT model for repository edits, reviews, and agentic software work
400,000 ctx
reasoning
OpenAI: GPT-5 Image
Image model for prompt-driven generation, editing, and visual design workflows
400,000 ctx
reasoning
OpenAI: GPT-5 Image Mini
Image model for prompt-driven generation, editing, and visual design workflows
400,000 ctx
reasoning
OpenAI: GPT-5 Mini
Compact GPT model for low-latency assistance and high-volume workloads
400,000 ctx
reasoning
OpenAI: GPT-5 Nano
Compact GPT model for low-latency assistance and high-volume workloads
400,000 ctx
reasoning
OpenAI: GPT-5 Pro
Frontier GPT model for professional reasoning, coding, and multimodal work
400,000 ctx
reasoning
OpenAI: GPT-5.1
GPT model for general reasoning, writing, coding, and tool-assisted tasks
400,000 ctx
reasoning
OpenAI: GPT-5.1 Chat
Chat-tuned GPT model for conversational assistance, writing, and tool workflows
128,000 ctx
OpenAI: GPT-5.1-Codex
Coding-optimized GPT model for repository edits, reviews, and agentic software work
400,000 ctx
reasoning
OpenAI: GPT-5.1-Codex-Max
Coding-optimized GPT model for repository edits, reviews, and agentic software work
400,000 ctx
reasoning
OpenAI: GPT-5.1-Codex-Mini
Coding-optimized GPT model for repository edits, reviews, and agentic software work
400,000 ctx
reasoning
OpenAI: GPT-5.2
GPT model for general reasoning, writing, coding, and tool-assisted tasks
400,000 ctx
reasoning
OpenAI: GPT-5.2 Chat
Chat-tuned GPT model for conversational assistance, writing, and tool workflows
128,000 ctx
OpenAI: GPT-5.2-Codex
Coding-optimized GPT model for repository edits, reviews, and agentic software work
400,000 ctx
reasoning
OpenAI: GPT-5.2 Pro
Frontier GPT model for professional reasoning, coding, and multimodal work
400,000 ctx
reasoning
OpenAI: GPT-5.3 Chat
Chat-tuned GPT model for conversational assistance, writing, and tool workflows
128,000 ctx
OpenAI: GPT-5.3-Codex
Coding-optimized GPT model for repository edits, reviews, and agentic software work
400,000 ctx
reasoning
OpenAI: GPT-5.4
Frontier GPT model for professional reasoning, coding, and multimodal work
1,050,000 ctx
reasoning
OpenAI: GPT-5.4 Image 2
Image model for prompt-driven generation, editing, and visual design workflows
272,000 ctx
reasoning
OpenAI: GPT-5.4 Mini
Compact GPT model for low-latency assistance and high-volume workloads
400,000 ctx
reasoning
OpenAI: GPT-5.4 Nano
Compact GPT model for low-latency assistance and high-volume workloads
400,000 ctx
reasoning
OpenAI: GPT-5.4 Pro
Frontier GPT model for professional reasoning, coding, and multimodal work
1,050,000 ctx
reasoning
OpenAI: GPT-5.5
Frontier GPT model for professional reasoning, coding, and multimodal work
1,050,000 ctx
reasoning
OpenAI: GPT-5.5 Pro
Frontier GPT model for professional reasoning, coding, and multimodal work
1,050,000 ctx
reasoning
OpenAI: GPT Audio
Speech generation model for controllable voice, narration, and audio delivery
128,000 ctx
OpenAI: GPT Audio Mini
Speech generation model for controllable voice, narration, and audio delivery
128,000 ctx
OpenAI: GPT Chat Latest
Chat-tuned GPT model for conversational assistance, writing, and tool workflows
400,000 ctx
reasoning
OpenAI: gpt-oss-120b
Open-weight GPT model for self-hosted reasoning and instruction-following workloads
131,072 ctx
reasoning
OpenAI: gpt-oss-20b
Open-weight GPT model for self-hosted reasoning and instruction-following workloads
131,072 ctx
reasoning
OpenAI: gpt-oss-safeguard-20b
Safety model for policy screening, moderation, and risk-aware routing workflows
131,072 ctx
reasoning
OpenAI: o1
O-series reasoning model for hard analysis, math, coding, and planning
200,000 ctx
OpenAI: o1-pro
O-series reasoning model for hard analysis, math, coding, and planning
200,000 ctx
reasoning
OpenAI: o3
O-series reasoning model for hard analysis, math, coding, and planning
200,000 ctx
reasoning
OpenAI: o3 Deep Research
Research model for long-horizon investigation, synthesis, and analytical reports
200,000 ctx
reasoning
OpenAI: o3 Mini
O-series reasoning model for hard analysis, math, coding, and planning
200,000 ctx
OpenAI: o3 Mini High
O-series reasoning model for hard analysis, math, coding, and planning
200,000 ctx
OpenAI: o3 Pro
O-series reasoning model for hard analysis, math, coding, and planning
200,000 ctx
reasoning
OpenAI: o4 Mini
O-series reasoning model for hard analysis, math, coding, and planning
200,000 ctx
reasoning
OpenAI: o4 Mini Deep Research
Research model for long-horizon investigation, synthesis, and analytical reports
200,000 ctx
reasoning
OpenAI: o4 Mini High
O-series reasoning model for hard analysis, math, coding, and planning
200,000 ctx
reasoning
Auto Router
Image model for prompt-driven generation, editing, and visual design workflows
2,000,000 ctx
reasoning
Body Builder (beta)
Preview model for early access evaluation, prototyping, and compatibility testing
128,000 ctx
Free Models Router
Multimodal reasoning model for visual analysis, planning, and tool use
200,000 ctx
reasoning
Owl Alpha
Reasoning model for deliberate analysis, multi-step problem solving, and tool use
1,048,756 ctx
reasoning
Pareto Code Router
Coding model for repository understanding, refactors, and agentic engineering tasks
200,000 ctx
Perceptron: Perceptron Mk1
Multimodal reasoning model for visual analysis, planning, and tool use
32,768 ctx
reasoning
Perplexity: Sonar
Sonar search model for current answers, retrieval, and citation-backed chat
127,072 ctx
Perplexity: Sonar Deep Research
Sonar search model for current answers, retrieval, and citation-backed chat
128,000 ctx
reasoning
Perplexity: Sonar Pro
Advanced Sonar search model for deeper research and cited synthesis
200,000 ctx
Perplexity: Sonar Pro Search
Advanced Sonar search model for deeper research and cited synthesis
200,000 ctx
reasoning
Perplexity: Sonar Reasoning Pro
Web-grounded reasoning model for multi-step research and cited answers
128,000 ctx
reasoning
Poolside: Laguna M.1 (free)
Free provider route for experiments, demos, and cost-sensitive chat workloads
262,144 ctx
reasoning
Poolside: Laguna XS.2 (free)
Free provider route for experiments, demos, and cost-sensitive chat workloads
262,144 ctx
reasoning
Prime Intellect: INTELLECT-3
Reasoning model for deliberate analysis, multi-step problem solving, and tool use
131,072 ctx
reasoning
Qwen2.5 72B Instruct
Qwen instruction model for multilingual chat, reasoning, and tool use
32,768 ctx
Qwen: Qwen2.5 7B Instruct
Qwen instruction model for multilingual chat, reasoning, and tool use
32,768 ctx
Qwen2.5 Coder 32B Instruct
Qwen coding model for software agents, repository edits, and code reasoning
32,768 ctx
Qwen: Qwen-Plus
Qwen instruction model for multilingual chat, reasoning, and tool use
1,000,000 ctx
Qwen: Qwen Plus 0728
Qwen instruction model for multilingual chat, reasoning, and tool use
1,000,000 ctx
Qwen: Qwen Plus 0728 (thinking)
Qwen reasoning model for deliberate problem solving, math, and coding
1,000,000 ctx
reasoning
Qwen: Qwen2.5 VL 72B Instruct
Qwen vision-language model for visual reasoning, documents, and agent tasks
32,768 ctx
Qwen: Qwen3 14B
Qwen instruction model for multilingual chat, reasoning, and tool use
40,960 ctx
reasoning
Qwen: Qwen3 235B A22B
Qwen instruction model for multilingual chat, reasoning, and tool use
131,072 ctx
reasoning
Qwen: Qwen3 235B A22B Instruct 2507
Qwen instruction model for multilingual chat, reasoning, and tool use
262,144 ctx
Qwen: Qwen3 235B A22B Thinking 2507
Qwen reasoning model for deliberate problem solving, math, and coding
262,144 ctx
reasoning
Qwen: Qwen3 30B A3B
Qwen instruction model for multilingual chat, reasoning, and tool use
40,960 ctx
reasoning
Qwen: Qwen3 30B A3B Instruct 2507
Qwen instruction model for multilingual chat, reasoning, and tool use
262,144 ctx
Qwen: Qwen3 30B A3B Thinking 2507
Qwen reasoning model for deliberate problem solving, math, and coding
32,768 ctx
reasoning
Qwen: Qwen3 32B
Qwen instruction model for multilingual chat, reasoning, and tool use
40,960 ctx
reasoning
Qwen: Qwen3 8B
Qwen instruction model for multilingual chat, reasoning, and tool use
40,960 ctx
reasoning
Qwen: Qwen3 Coder 480B A35B
Qwen coding model for software agents, repository edits, and code reasoning
262,144 ctx
Qwen: Qwen3 Coder 30B A3B Instruct
Qwen coding model for software agents, repository edits, and code reasoning
160,000 ctx
Qwen: Qwen3 Coder Flash
Qwen coding model for software agents, repository edits, and code reasoning
1,000,000 ctx
Qwen: Qwen3 Coder Next
Qwen coding model for software agents, repository edits, and code reasoning
262,144 ctx
Qwen: Qwen3 Coder Plus
Qwen coding model for software agents, repository edits, and code reasoning
1,000,000 ctx
Qwen: Qwen3 Max
Flagship Qwen model for complex reasoning, coding, and agentic workflows
262,144 ctx
Qwen: Qwen3 Max Thinking
Qwen reasoning model for deliberate problem solving, math, and coding
262,144 ctx
reasoning
Qwen: Qwen3 Next 80B A3B Instruct
Qwen instruction model for multilingual chat, reasoning, and tool use
131,072 ctx
Qwen: Qwen3 Next 80B A3B Thinking
Qwen reasoning model for deliberate problem solving, math, and coding
131,072 ctx
reasoning
Qwen: Qwen3 VL 235B A22B Instruct
Qwen vision-language model for visual reasoning, documents, and agent tasks
262,144 ctx
Qwen: Qwen3 VL 235B A22B Thinking
Qwen vision-language model for visual reasoning, documents, and agent tasks
131,072 ctx
reasoning
Qwen: Qwen3 VL 30B A3B Instruct
Qwen vision-language model for visual reasoning, documents, and agent tasks
131,072 ctx
Qwen: Qwen3 VL 30B A3B Thinking
Qwen vision-language model for visual reasoning, documents, and agent tasks
131,072 ctx
reasoning
Qwen: Qwen3 VL 32B Instruct
Qwen vision-language model for visual reasoning, documents, and agent tasks
131,072 ctx
Qwen: Qwen3 VL 8B Instruct
Qwen vision-language model for visual reasoning, documents, and agent tasks
131,072 ctx
Qwen: Qwen3 VL 8B Thinking
Qwen vision-language model for visual reasoning, documents, and agent tasks
131,072 ctx
reasoning
Qwen: Qwen3.5-122B-A10B
Qwen vision-language model for visual reasoning, documents, and agent tasks
262,144 ctx
reasoning
Qwen: Qwen3.5-27B
Qwen vision-language model for visual reasoning, documents, and agent tasks
262,144 ctx
reasoning
Qwen: Qwen3.5-35B-A3B
Qwen vision-language model for visual reasoning, documents, and agent tasks
262,144 ctx
reasoning
Qwen: Qwen3.5 397B A17B
Qwen vision-language model for visual reasoning, documents, and agent tasks
262,144 ctx
reasoning
Qwen: Qwen3.5-9B
Qwen vision-language model for visual reasoning, documents, and agent tasks
256,000 ctx
reasoning
Qwen: Qwen3.5-Flash
Qwen vision-language model for visual reasoning, documents, and agent tasks
1,000,000 ctx
reasoning
Qwen: Qwen3.5 Plus 2026-02-15
Qwen vision-language model for visual reasoning, documents, and agent tasks
1,000,000 ctx
reasoning
Qwen: Qwen3.5 Plus 2026-04-20
Qwen vision-language model for visual reasoning, documents, and agent tasks
1,000,000 ctx
reasoning
Qwen: Qwen3.6 27B
Qwen vision-language model for visual reasoning, documents, and agent tasks
256,000 ctx
reasoning
Qwen: Qwen3.6 35B A3B
Qwen vision-language model for visual reasoning, documents, and agent tasks
262,144 ctx
reasoning
Qwen: Qwen3.6 Flash
Qwen vision-language model for visual reasoning, documents, and agent tasks
1,000,000 ctx
reasoning
Qwen: Qwen3.6 Max Preview
Flagship Qwen model for complex reasoning, coding, and agentic workflows
262,144 ctx
reasoning
Qwen: Qwen3.6 Plus
Qwen vision-language model for visual reasoning, documents, and agent tasks
1,000,000 ctx
reasoning
Qwen: Qwen3.7 Max
Flagship Qwen model for complex reasoning, coding, and agentic workflows
1,000,000 ctx
reasoning
Reka Edge
Multimodal model for analyzing text, images, documents, and rich media
16,384 ctx
Reka Flash 3
Efficient model for low-latency assistance, extraction, and routine automation
65,536 ctx
reasoning
Relace: Relace Apply 3
General-purpose chat model for instruction following, writing, and analysis
256,000 ctx
Relace: Relace Search
Tool-capable chat model for instruction following and agentic application workflows
256,000 ctx
Sao10k: Llama 3 Euryale 70B v2.1
Open Llama instruction model for multilingual chat, reasoning, and coding
8,192 ctx
Sao10K: Llama 3 8B Lunaris
Open Llama instruction model for multilingual chat, reasoning, and coding
8,192 ctx
Sao10K: Llama 3.1 70B Hanami x1
Open Llama instruction model for multilingual chat, reasoning, and coding
16,000 ctx
Sao10K: Llama 3.1 Euryale 70B v2.2
Open Llama instruction model for multilingual chat, reasoning, and coding
131,072 ctx
Sao10K: Llama 3.3 Euryale 70B
Open Llama instruction model for multilingual chat, reasoning, and coding
131,072 ctx
Stealth: Claude Opus 4.6 (20% off)
Flagship Claude model for deep reasoning, coding, and long-horizon agents
1,000,000 ctx
reasoning
Stealth: Claude Opus 4.7 (20% off)
Flagship Claude model for deep reasoning, coding, and long-horizon agents
1,000,000 ctx
reasoning
Stealth: Claude Sonnet 4.6 (20% off)
Balanced Claude model for coding, analysis, agent workflows, and cost control
1,000,000 ctx
reasoning
StepFun: Step 3.5 Flash
StepFun flash model for efficient multimodal reasoning, coding, and tool use
256,000 ctx
reasoning
Switchpoint Router
Automatic model router for matching prompts to suitable backends and budgets
131,072 ctx
reasoning
Tencent: Hunyuan A13B Instruct
Tencent Hy reasoning model for coding, instruction following, and agent tasks
131,072 ctx
reasoning
Tencent: Hy3 Preview
Tencent Hy reasoning model for coding, instruction following, and agent tasks
262,144 ctx
reasoning
TheDrummer: Cydonia 24B V4.1
Open-weight instruction model for adaptable chat and self-hosted production workloads
131,072 ctx
TheDrummer: Rocinante 12B
Open-weight instruction model for adaptable chat and self-hosted production workloads
32,768 ctx
TheDrummer: Skyfall 36B V2
Open-weight instruction model for adaptable chat and self-hosted production workloads
32,768 ctx
TheDrummer: UnslopNemo 12B
Open-weight instruction model for adaptable chat and self-hosted production workloads
32,768 ctx
ReMM SLERP 13B
Open-weight instruction model for adaptable chat and self-hosted production workloads
6,144 ctx
Upstage: Solar Pro 3
Flagship model for demanding analysis, coding, and production agent workflows
128,000 ctx
reasoning
Writer: Palmyra X5
General-purpose chat model for instruction following, writing, and analysis
1,040,000 ctx
xAI: Grok 4.20
Grok model for agentic tool use, reasoning, coding, and live assistance
2,000,000 ctx
reasoning
xAI: Grok 4.20 Multi-Agent
Grok model for agentic tool use, reasoning, coding, and live assistance
2,000,000 ctx
reasoning
xAI: Grok 4.3
Grok model for agentic tool use, reasoning, coding, and live assistance
1,000,000 ctx
reasoning
xAI: Grok Build 0.1
Grok coding model for agentic engineering, edits, and codebase workflows
256,000 ctx
reasoning
Xiaomi: MiMo-V2-Flash
MiMo flash model for fast multimodal assistance and agent workflows
262,144 ctx
reasoning
Xiaomi: MiMo-V2-Omni
MiMo omni model for text, image, video, audio, and agents
262,144 ctx
reasoning
Xiaomi: MiMo-V2-Pro
Earlier MiMo Pro model for multimodal agents, reasoning, and code tasks
1,048,576 ctx
reasoning
Xiaomi: MiMo-V2.5
Open MiMo model for multimodal coding agents and long-context automation
1,048,576 ctx
reasoning
Xiaomi: MiMo V2.5 Pro
Stronger MiMo Pro tier for multimodal reasoning and coding-agent execution
1,048,576 ctx
reasoning
Z.ai: GLM 4 32B
Flagship GLM model for hybrid reasoning, coding, and agentic engineering
128,000 ctx
Z.ai: GLM 4.5
Flagship GLM model for hybrid reasoning, coding, and agentic engineering
131,072 ctx
reasoning
Z.ai: GLM 4.5 Air
Efficient GLM model for fast reasoning, coding, and agent workflows
131,072 ctx
reasoning
Z.ai: GLM 4.5V
GLM vision model for visual reasoning, documents, and multimodal agents
65,536 ctx
reasoning
Z.ai: GLM 4.6
Flagship GLM model for hybrid reasoning, coding, and agentic engineering
204,800 ctx
reasoning
Z.ai: GLM 4.6V
GLM vision model for visual reasoning, documents, and multimodal agents
131,072 ctx
reasoning
Z.ai: GLM 4.7
Flagship GLM model for hybrid reasoning, coding, and agentic engineering
202,752 ctx
reasoning
Z.ai: GLM 4.7 Flash
Efficient GLM model for fast reasoning, coding, and agent workflows
202,752 ctx
reasoning
Z.ai: GLM 5
Flagship GLM model for hybrid reasoning, coding, and agentic engineering
202,752 ctx
reasoning
Z.ai: GLM 5 Turbo
Efficient GLM model for fast reasoning, coding, and agent workflows
202,752 ctx
reasoning
Z.ai: GLM 5.1
Flagship GLM model for hybrid reasoning, coding, and agentic engineering
202,752 ctx
reasoning
Z.ai: GLM 5V Turbo
GLM vision model for visual reasoning, documents, and multimodal agents
202,752 ctx
reasoning
Anthropic: Claude Haiku Latest
Fast Claude model for responsive assistance, classification, and lightweight agents
200,000 ctx
reasoning
Anthropic: Claude Opus Latest
Flagship Claude model for deep reasoning, coding, and long-horizon agents
1,000,000 ctx
reasoning
Anthropic: Claude Sonnet Latest
Balanced Claude model for coding, analysis, agent workflows, and cost control
1,000,000 ctx
reasoning
Google: Gemini Flash Latest
Fast Gemini model balancing multimodal reasoning, tool use, and cost
1,048,576 ctx
reasoning
Google: Gemini Pro Latest
Advanced Gemini model for complex reasoning, coding, and multimodal analysis
1,048,576 ctx
reasoning
MoonshotAI: Kimi Latest
Kimi multimodal agent model for visual understanding, coding, and planning
262,142 ctx
reasoning
OpenAI: GPT Latest
GPT model for general reasoning, writing, coding, and tool-assisted tasks
1,050,000 ctx
reasoning
OpenAI: GPT Mini Latest
Compact GPT model for low-latency assistance and high-volume workloads
400,000 ctx
reasoning
Comments (0)
Sign in
to leave a comment
No comments yet.
Community Rating
—
No ratings yet
Sign in
to rate