Viberank
Trending models
Popular Labs
Hottest providers
Top Agents
Login
Sign up
← All providers
Venice AI
Provider ID: venice
Details
SDK package
venice-ai-sdk-provider
Models
90
Links
Documentation
Auth: VENICE_API_KEY
Models (90)
Aion 3.0
Reasoning model for deliberate analysis, multi-step problem solving, and tool use
128,000 ctx
reasoning
Aion 3.0 Mini
Efficient model for low-latency assistance, extraction, and routine automation
128,000 ctx
reasoning
Claude Fable 5
Claude model for creative writing, analysis, and controlled agent workflows
1,000,000 ctx
reasoning
Claude Opus 4.5
Flagship Claude model for deep reasoning, coding, and long-horizon agents
198,000 ctx
reasoning
Claude Opus 4.6
High-end Claude for difficult coding, planning, and slower expert reasoning
1,000,000 ctx
reasoning
Claude Opus 4.7
Stronger Opus tier for advanced software work and high-stakes reasoning
1,000,000 ctx
reasoning
Claude Opus 4.8
Top Claude Opus tier for the hardest reasoning, coding, and long-horizon agents
1,000,000 ctx
reasoning
Claude Opus 4.8 Fast
Top Claude Opus tier for the hardest reasoning, coding, and long-horizon agents
1,000,000 ctx
reasoning
Claude Opus 5
Flagship Claude model for deep reasoning, coding, and long-horizon agents
1,000,000 ctx
reasoning
Claude Opus 5 Fast
Flagship Claude model for deep reasoning, coding, and long-horizon agents
1,000,000 ctx
reasoning
Claude Sonnet 4.5
Balanced Claude model for coding, analysis, agent workflows, and cost control
198,000 ctx
reasoning
Claude Sonnet 4.6
Claude workhorse for coding agents, careful analysis, and production cost control
1,000,000 ctx
reasoning
Claude Sonnet 5
Balanced Claude model for coding, analysis, agent workflows, and cost control
1,000,000 ctx
reasoning
DeepSeek V3.2
DeepSeek chat model for instruction following, coding, and analysis
160,000 ctx
reasoning
DeepSeek V4 Flash
Fast DeepSeek V4 lane for economical reasoning, coding, and long-context work
1,000,000 ctx
reasoning
DeepSeek V4 Pro
Open MoE flagship with million-token context for coding and long agent runs
1,000,000 ctx
reasoning
Gemini 3.1 Pro Preview
Reasoning-first Gemini preview for agentic coding and complex problem solving
1,000,000 ctx
reasoning
Gemini 3.5 Flash
Fast Gemini model balancing multimodal reasoning, tool use, and cost
1,000,000 ctx
reasoning
Gemini 3.5 Flash-Lite
Low-latency Gemini model for high-volume multimodal and agent workloads
1,000,000 ctx
reasoning
Gemini 3.6 Flash
Fast Gemini model balancing multimodal reasoning, tool use, and cost
1,000,000 ctx
reasoning
Gemini 3 Flash Preview
New Gemini flash lane bringing frontier-style multimodal reasoning to cheaper runs
256,000 ctx
reasoning
Gemma 4 Uncensored
Open Gemma instruction model for efficient chat and self-hosted deployments
256,000 ctx
Google Gemma 3 27B Instruct
Open Gemma instruction model for efficient chat and self-hosted deployments
198,000 ctx
Google Gemma 4 26B A4B Instruct
Open Gemma instruction model for efficient chat and self-hosted deployments
256,000 ctx
reasoning
Google Gemma 4 31B Instruct
Largest Gemma 4 instruction model for open, self-hosted chat and reasoning
256,000 ctx
reasoning
Grok 4.20
Grok model for agentic tool use, reasoning, coding, and live assistance
2,000,000 ctx
reasoning
Grok 4.20 Multi-Agent
Grok model for agentic tool use, reasoning, coding, and live assistance
2,000,000 ctx
reasoning
Grok 4.3
xAI's default Grok for chat, coding, agentic tools, and lower hallucination risk
1,000,000 ctx
reasoning
Grok 4.5
Grok model for agentic tool use, reasoning, coding, and live assistance
500,000 ctx
reasoning
Grok Build 0.1
Fast Grok coding model tuned for agentic engineering and iterative edits
256,000 ctx
reasoning
Hermes 3 Llama 3.1 405b
Open Llama instruction model for multilingual chat, reasoning, and coding
128,000 ctx
Inkling
Multimodal reasoning model for visual analysis, planning, and tool use
1,000,000 ctx
reasoning
Kimi K2.5
Earlier Kimi frontier model for long-context agents, coding, and multimodal work
256,000 ctx
reasoning
Kimi K2.6
Multimodal Kimi workhorse for agent loops, coding tasks, and visual context
256,000 ctx
reasoning
Kimi K2.7 Code
Coding-focused Kimi model, stronger on long-horizon repo work with less overthinking
256,000 ctx
reasoning
Kimi K3
Kimi multimodal agent model for visual understanding, coding, and planning
1,000,000 ctx
reasoning
Llama 3.2 3B
Open Llama instruction model for multilingual chat, reasoning, and coding
128,000 ctx
Llama 3.3 70B
Open Llama instruction model for multilingual chat, reasoning, and coding
128,000 ctx
Mercury 2
Reasoning model for deliberate analysis, multi-step problem solving, and tool use
128,000 ctx
reasoning
MiniMax M2.5
Prior MiniMax coding model for agent workflows, office edits, and automation
198,000 ctx
reasoning
MiniMax M2.7
Open MiniMax flagship for coding agents, office automation, and complex environments
198,000 ctx
reasoning
MiniMax M3 Preview
MiniMax multimodal coding model for long-context reasoning and agent tasks
524,288 ctx
reasoning
Mistral Small 4
Fast Mistral production model for chat, extraction, and cost-sensitive agents
256,000 ctx
reasoning
Mistral Small 3.2 24B Instruct
Efficient Mistral model for fast chat, extraction, and production assistants
256,000 ctx
NVIDIA Nemotron 3 Nano 30B
Small Nemotron 3 MoE for efficient coding, math, and long-context agents
128,000 ctx
NVIDIA Nemotron 3 Ultra
Largest Nemotron 3 model for maximum open-weight reasoning and agent accuracy
256,000 ctx
reasoning
Nemotron Cascade 2 30B A3B
Nemotron model for efficient reasoning, coding, and specialized AI agents
256,000 ctx
reasoning
GLM 4.7 Flash Heretic
Efficient GLM model for fast reasoning, coding, and agent workflows
200,000 ctx
reasoning
GPT-4o
GPT model for general reasoning, writing, coding, and tool-assisted tasks
128,000 ctx
GPT-4o Mini
Small omni GPT for cheap multimodal assistance and production-scale traffic
128,000 ctx
GPT-5.2
Reliable GPT generation for broad coding, writing, and tool-assisted product work
256,000 ctx
reasoning
GPT-5.2 Codex
Code-specialist GPT for repository edits, reviews, and long-running software agents
256,000 ctx
reasoning
GPT-5.3 Codex
Coding-optimized GPT model for repository edits, reviews, and agentic software work
400,000 ctx
reasoning
GPT-5.4
Agent-ready GPT for coding and computer-use workflows at a lower cost
1,000,000 ctx
reasoning
GPT-5.4 Mini
Strong small GPT for coding subagents, quick tool use, and high-volume work
400,000 ctx
reasoning
GPT-5.4 Pro
More exact GPT-5.4 tier for demanding professional reasoning and agent tasks
1,000,000 ctx
reasoning
GPT-5.5
Default frontier GPT for coding, computer use, research, and knowledge work
1,000,000 ctx
reasoning
GPT-5.5 Pro
Highest-accuracy GPT-5.5 tier for slower, precision-heavy reasoning and coding
1,000,000 ctx
reasoning
GPT-5.6 Luna
GPT model for general reasoning, writing, coding, and tool-assisted tasks
1,000,000 ctx
reasoning
GPT-5.6 Luna Pro
Frontier GPT model for professional reasoning, coding, and multimodal work
1,000,000 ctx
reasoning
GPT-5.6 Sol
GPT model for general reasoning, writing, coding, and tool-assisted tasks
1,000,000 ctx
reasoning
GPT-5.6 Sol Pro
Frontier GPT model for professional reasoning, coding, and multimodal work
1,000,000 ctx
reasoning
GPT-5.6 Terra
GPT model for general reasoning, writing, coding, and tool-assisted tasks
1,000,000 ctx
reasoning
GPT-5.6 Terra Pro
Frontier GPT model for professional reasoning, coding, and multimodal work
1,000,000 ctx
reasoning
OpenAI GPT OSS 120B
Open GPT reasoning model for self-hosted agents and controllable deployments
128,000 ctx
reasoning
Qwen 3.6 Plus Uncensored
Earlier Qwen multimodal workhorse for million-token agent and document tasks
1,000,000 ctx
reasoning
Qwen 3.7 Max
Qwen frontier model tuned for agent frameworks, coding assistants, and long tasks
1,000,000 ctx
reasoning
Qwen 3.7 Plus
Multimodal Qwen workhorse for long-context agents, visual inputs, and coding
1,000,000 ctx
reasoning
Qwen 3 235B A22B Instruct 2507
Qwen instruction model for multilingual chat, reasoning, and tool use
128,000 ctx
Qwen 3 235B A22B Thinking 2507
Qwen reasoning model for deliberate problem solving, math, and coding
128,000 ctx
reasoning
Qwen 3.5 35B A3B
Qwen vision-language model for visual reasoning, documents, and agent tasks
256,000 ctx
reasoning
Qwen 3.5 397B
Large open Qwen multimodal MoE for visual agents and long technical tasks
128,000 ctx
reasoning
Qwen 3.5 9B
Qwen instruction model for multilingual chat, reasoning, and tool use
256,000 ctx
reasoning
Qwen 3.6 27B
Qwen vision-language model for visual reasoning, documents, and agent tasks
256,000 ctx
reasoning
Qwen 3.6 35B A3B
Qwen instruction model for multilingual chat, reasoning, and tool use
256,000 ctx
reasoning
Qwen 3 Coder 480B Turbo
Qwen coding model for software agents, repository edits, and code reasoning
256,000 ctx
Qwen 3 Next 80b
Qwen instruction model for multilingual chat, reasoning, and tool use
256,000 ctx
Qwen3 VL 235B
Multimodal model for analyzing text, images, documents, and rich media
128,000 ctx
Seed 2.1 Turbo
Multimodal reasoning model for visual analysis, planning, and tool use
256,000 ctx
reasoning
Venice Uncensored 1.2
Multimodal model for analyzing text, images, documents, and rich media
128,000 ctx
Venice Role Play Uncensored
Multimodal model for analyzing text, images, documents, and rich media
128,000 ctx
MiMo-V2.5
Open MiMo model for multimodal coding agents and long-context automation
1,000,000 ctx
reasoning
GLM 5 Turbo
Faster GLM-5 lane for coding agents that need lower latency
200,000 ctx
reasoning
GLM 5V Turbo
Fast GLM vision model for screenshots, documents, and multimodal agent tasks
200,000 ctx
reasoning
GLM 4.6
Late GLM-4 workhorse for coding agents, reasoning, and structured tasks
198,000 ctx
reasoning
GLM 4.7
Mature GLM model for dependable coding, reasoning, and structured agent tasks
198,000 ctx
reasoning
GLM 4.7 Flash
Budget GLM lane for fast coding help, routing, and everyday automation
128,000 ctx
reasoning
GLM 5
General GLM flagship for coding, analysis, and tool-heavy engineering workflows
198,000 ctx
reasoning
GLM 5.1
Strong GLM coding model for agentic engineering, terminals, and repository generation
200,000 ctx
reasoning
GLM 5.2
Open flagship GLM for long-horizon coding agents and million-token context work
1,000,000 ctx
reasoning
Comments (0)
Sign in
to leave a comment
No comments yet.
Community Rating
—
No ratings yet
Sign in
to rate