Viberank
Trending models
Popular Labs
Hottest providers
Top Agents
Login
Sign up
← All providers
Nvidia
Provider ID: nvidia
Details
SDK package
@ai-sdk/openai-compatible
Models
84
API
https://integrate.api.nvidia.com/v1
Links
Documentation
Auth: NVIDIA_API_KEY
Models (84)
dracarys-llama-3.1-70b-instruct
Open Llama instruction model for multilingual chat, reasoning, and coding
128,000 ctx
BGE M3
Flagship model for demanding analysis, coding, and production agent workflows
8,192 ctx
FLUX.1-dev
Image model for prompt-driven generation, editing, and visual design workflows
4,096 ctx
FLUX.1-Kontext-dev
Image model for prompt-driven generation, editing, and visual design workflows
40,960 ctx
FLUX.1-schnell
Image model for prompt-driven generation, editing, and visual design workflows
77 ctx
FLUX.2 Klein 4B
Image model for prompt-driven generation, editing, and visual design workflows
40,960 ctx
ByteDance-Seed/Seed-OSS-36B-Instruct
Tool-capable chat model for instruction following and agentic application workflows
262,000 ctx
DeepSeek V4 Flash
Fast DeepSeek V4 lane for economical reasoning, coding, and long-context work
1,048,576 ctx
reasoning
DeepSeek V4 Pro
Open MoE flagship with million-token context for coding and long agent runs
1,048,576 ctx
reasoning
Gemma 2 2b It
Open Gemma instruction model for efficient chat and self-hosted deployments
128,000 ctx
Gemma 3n E2b It
Open Gemma instruction model for efficient chat and self-hosted deployments
128,000 ctx
Gemma 3n E4b It
Open Gemma instruction model for efficient chat and self-hosted deployments
128,000 ctx
Gemma-4-31B-IT
Open Gemma instruction model for efficient chat and self-hosted deployments
256,000 ctx
reasoning
paligemma
Gemini multimodal model for text, image, audio, video, and document tasks
128,000 ctx
esm2-650m
Open Llama instruction model for multilingual chat, reasoning, and coding
128,000 ctx
esmfold
Open Llama instruction model for multilingual chat, reasoning, and coding
128,000 ctx
Llama 3.1 70b Instruct
Open Llama instruction model for multilingual chat, reasoning, and coding
128,000 ctx
Llama 3.1 8B Instruct
Open Llama instruction model for multilingual chat, reasoning, and coding
16,000 ctx
Llama 3.2 11b Vision Instruct
Open Llama multimodal model for image understanding and text reasoning
128,000 ctx
Llama 3.2 1b Instruct
Open Llama instruction model for multilingual chat, reasoning, and coding
128,000 ctx
Llama 3.2 3B Instruct
Open Llama instruction model for multilingual chat, reasoning, and coding
32,768 ctx
Llama-3.2-90B-Vision-Instruct
Open Llama multimodal model for image understanding and text reasoning
128,000 ctx
Llama 3.3 70b Instruct
Open Llama instruction model for multilingual chat, reasoning, and coding
128,000 ctx
Llama 4 Maverick 17b 128e Instruct
Open multimodal Llama model for strong reasoning and fast responses
128,000 ctx
Llama Guard 4 12B
Safety model for policy screening, moderation, and risk-aware routing workflows
128,000 ctx
Phi-4-Mini
Efficient model for low-latency assistance, extraction, and routine automation
131,072 ctx
Phi 4 Multimodal
General-purpose chat model for instruction following, writing, and analysis
128,000 ctx
MiniMax-M2.7
MiniMax model for chat, coding, office work, and agentic tasks
204,800 ctx
reasoning
MiniMax-M3
MiniMax multimodal model for long-context coding, perception, and agent planning
1,000,000 ctx
reasoning
Magistral Small 2506
Mistral reasoning model for transparent analysis, math, and complex decisions
32,768 ctx
Mistral-7B-Instruct-v0.3
Mistral model for multilingual chat, reasoning, and tool-assisted workflows
65,536 ctx
Mistral Large 3 675B Instruct 2512
Flagship Mistral model for advanced reasoning, coding, and multilingual work
262,144 ctx
Mistral Medium 3
Mistral model for multilingual chat, reasoning, and tool-assisted workflows
131,072 ctx
mistral-nemotron
Mistral model for multilingual chat, reasoning, and tool-assisted workflows
128,000 ctx
mistral-small-4-119b-2603
Efficient Mistral model for fast chat, extraction, and production assistants
128,000 ctx
reasoning
Mistral: Mixtral 8x22B Instruct
Mistral model for multilingual chat, reasoning, and tool-assisted workflows
65,536 ctx
Mistral: Mixtral 8x7B Instruct
Mistral model for multilingual chat, reasoning, and tool-assisted workflows
32,768 ctx
Kimi K2 0905
Kimi model for long-context chat, coding, and agentic reasoning
262,144 ctx
Kimi K2.6
Kimi multimodal agent model for visual understanding, coding, and planning
262,144 ctx
reasoning
Active Speaker Detection
Nemotron multimodal model for visual reasoning and agentic AI workflows
0
bevformer
Nemotron multimodal model for visual reasoning and agentic AI workflows
128,000 ctx
cosmos-predict1-5b
Video model for prompt-guided generation, editing, and motion workflows
0
cosmos-transfer1-7b
Video model for prompt-guided generation, editing, and motion workflows
0
cosmos-transfer2.5-2b
Video model for prompt-guided generation, editing, and motion workflows
0
gliner-pii
Nemotron model for efficient reasoning, coding, and specialized AI agents
128,000 ctx
llama-3.1-nemotron-safety-guard-8b-v3
Safety model for policy screening, moderation, and risk-aware routing workflows
128,000 ctx
llama-3_2-nemoretriever-300m-embed-v1
Embedding model for semantic search, retrieval, clustering, and ranking pipelines
32,768 ctx
llama-nemotron-embed-vl-1b-v2
Embedding model for semantic search, retrieval, clustering, and ranking pipelines
32,768 ctx
llama-nemotron-rerank-vl-1b-v2
Reranking model for improving retrieval quality in search and recommendation systems
128,000 ctx
magpie-tts-zeroshot
Speech generation model for controllable voice, narration, and audio delivery
0
nemotron-3-content-safety
Safety model for policy screening, moderation, and risk-aware routing workflows
128,000 ctx
nemotron-3-nano-30b-a3b
Small Nemotron 3 MoE for efficient coding, math, and long-context agents
131,072 ctx
reasoning
Nemotron 3 Nano Omni
Open Nemotron omni model combining reasoning with text, vision, and audio
256,000 ctx
reasoning
Nemotron 3 Super
Nemotron middle tier for collaborative agents and high-volume reasoning workloads
262,144 ctx
reasoning
Nemotron 3 Ultra 550B A55B
Largest Nemotron 3 model for maximum open-weight reasoning and agent accuracy
1,000,000 ctx
reasoning
nemotron-content-safety-reasoning-4b
Safety model for policy screening, moderation, and risk-aware routing workflows
128,000 ctx
reasoning
nemotron-mini-4b-instruct
Compact Nemotron model for efficient reasoning and deployable AI agents
128,000 ctx
nemotron-voicechat
Nemotron multimodal model for visual reasoning and agentic AI workflows
128,000 ctx
nv-embed-v1
Embedding model for semantic search, retrieval, clustering, and ranking pipelines
32,768 ctx
nv-embedcode-7b-v1
Nemotron model for efficient reasoning, coding, and specialized AI agents
32,768 ctx
nvidia-nemotron-nano-9b-v2
Compact Nemotron model for efficient reasoning and deployable AI agents
131,072 ctx
reasoning
rerank-qa-mistral-4b
Reranking model for improving retrieval quality in search and recommendation systems
128,000 ctx
riva-translate-4b-instruct-v1_1
Translation model for multilingual conversion, localization, and cross-language workflows
128,000 ctx
sparsedrive
Nemotron multimodal model for visual reasoning and agentic AI workflows
128,000 ctx
streampetr
Nemotron multimodal model for visual reasoning and agentic AI workflows
128,000 ctx
studiovoice
Nemotron model for efficient reasoning, coding, and specialized AI agents
128,000 ctx
synthetic-video-detector
Video model for prompt-guided generation, editing, and motion workflows
0
usdcode
Nemotron model for efficient reasoning, coding, and specialized AI agents
128,000 ctx
usdvalidate
Nemotron model for efficient reasoning, coding, and specialized AI agents
0
GPT-OSS-120B
Open GPT reasoning model for self-hosted agents and controllable deployments
128,000 ctx
reasoning
GPT OSS 20B
Open-weight GPT model for self-hosted reasoning and instruction-following workloads
131,072 ctx
reasoning
Whisper Large v3
Speech transcription model for accurate audio-to-text and captioning workflows
0
Qwen Image
Image model for prompt-driven generation, editing, and visual design workflows
0
Qwen Image Edit
Image model for prompt-driven generation, editing, and visual design workflows
0
Qwen2.5 Coder 32b Instruct
Qwen coding model for software agents, repository edits, and code reasoning
128,000 ctx
Qwen3 Coder 480B A35B Instruct
Qwen coding model for software agents, repository edits, and code reasoning
262,144 ctx
Qwen3-Next-80B-A3B-Instruct
Qwen instruction model for multilingual chat, reasoning, and tool use
262,144 ctx
Qwen3.5 122B-A10B
Qwen vision-language model for visual reasoning, documents, and agent tasks
262,144 ctx
reasoning
Qwen3.5-397B-A17B
Qwen vision-language model for visual reasoning, documents, and agent tasks
262,144 ctx
reasoning
sarvam-m
Efficient Indian-language reasoning model for chat, coding, and multilingual work
128,000 ctx
Step 3.5 Flash
StepFun flash model for efficient multimodal reasoning, coding, and tool use
256,000 ctx
reasoning
Step 3.7 Flash
StepFun flash model for efficient multimodal reasoning, coding, and tool use
256,000 ctx
reasoning
solar-10.7b-instruct
Open-weight instruction model for adaptable chat and self-hosted production workloads
128,000 ctx
GLM-5.2
Open flagship GLM for long-horizon coding agents and million-token context work
1,000,000 ctx
reasoning
Comments (0)
Sign in
to leave a comment
No comments yet.
Community Rating
—
No ratings yet
Sign in
to rate