Skip to main content

Providers & Models

AICP routes to 23 providers across four categories: commercial, European (GDPR-native), inference platforms, and self-hosted/OpenAI-compatible. Every request goes through one endpoint; which provider actually serves it is decided by your routing configuration and by which providers are connected (platform-wide by your admin, or by you via BYOK).

BYOK: bring your own key

Connect your own provider credentials and they're used for real inference immediately. Your own key always takes priority over any provider your admin has enabled platform-wide, for that same provider. Available on every plan, including Community (self-hosted).

await client.providers.add({ provider: 'openai', apiKey: 'sk-...' });

Models are discovered automatically from the provider's own /models endpoint, so you don't need to list them yourself. See the full guide for managing, testing, and disabling providers: JavaScript · Python.

Some providers need the dashboard, not the SDK, today

client.providers.add() only takes a single apiKey right now. Four providers, AWS Bedrock, AWS SageMaker, Azure OpenAI, and Azure AI Foundry, authenticate with structured, multi-field credentials (SigV4 keys, resource/deployment names, API versions) instead of a single key, and self-hosted/custom endpoints need a base_url. Both are only exposed through the dashboard's Providers tab right now, not the SDKs. See the Providers tab below for exactly which fields each one needs.

Browse the catalog

The same search and filters as the product's own Model Catalog page. This is a static snapshot for the public docs (there's no org/BYOK context here), so connection status isn't shown. Self-hosted providers (vLLM, Ollama, TGI, NVIDIA NIM, SGLang, Custom Endpoint) have no fixed model list by design: the model is whatever you've deployed there, so they aren't part of the searchable catalog below; see the self-hosted section underneath it.

50 models across 11 providers with pre-registered catalog entries. Self-hosted providers add whatever you deploy

Region
Type
gpt-3.5-turbo
OpenAI
Q 53

US · EU

StreamingVisionTool CallingJSON ModeReasoning+3

$0.50 in · $1.50 out / 1M tokens · 16,385 ctx

gpt-4o-mini
OpenAI
Q 68

US · EU

StreamingVisionTool CallingJSON ModeReasoning+3

$0.15 in · $0.60 out / 1M tokens · 128,000 ctx

gpt-4o
OpenAI
Q 70

US · EU

StreamingVisionTool CallingJSON ModeReasoning+3

$2.50 in · $10.00 out / 1M tokens · 128,000 ctx

o4-mini
OpenAI
Q 83

US · EU

StreamingVisionTool CallingJSON ModeReasoning+3

$1.10 in · $4.40 out / 1M tokens · 200,000 ctx

o1
OpenAI
Q 90

US · EU

StreamingVisionTool CallingJSON ModeReasoning+3

$15.00 in · $60.00 out / 1M tokens · 200,000 ctx

gpt-4o-mini
Azure OpenAI
Q 68

US · EU

StreamingVisionTool CallingJSON Mode

$0.15 in · $0.60 out / 1M tokens · 128,000 ctx

gpt-4o
Azure OpenAI
Q 70

US · EU

StreamingVisionTool CallingJSON Mode

$2.50 in · $10.00 out / 1M tokens · 128,000 ctx

claude-haiku-4-5-20251001
Anthropic
Q 81

US · EU

StreamingVisionTool CallingJSON ModeLong Context+1

$1.00 in · $5.00 out / 1M tokens · 200,000 ctx

claude-sonnet-4-6
Anthropic
Q 84

US · EU

StreamingVisionTool CallingJSON ModeLong Context+1

$3.00 in · $15.00 out / 1M tokens · 1,000,000 ctx

claude-sonnet-5
Anthropic
Q 86

US · EU

StreamingVisionTool CallingJSON ModeLong Context+1

$3.00 in · $15.00 out / 1M tokens · 1,000,000 ctx

claude-opus-4-6
Anthropic
Q 87

US · EU

StreamingVisionTool CallingJSON ModeLong Context+1

$5.00 in · $25.00 out / 1M tokens · 1,000,000 ctx

claude-opus-4-7
Anthropic
Q 87

US · EU

StreamingVisionTool CallingJSON ModeLong Context+1

$5.00 in · $25.00 out / 1M tokens · 1,000,000 ctx

claude-opus-4-8
Anthropic
Q 87

US · EU

StreamingVisionTool CallingJSON ModeLong Context+1

$5.00 in · $25.00 out / 1M tokens · 1,000,000 ctx

claude-fable-5
Anthropic
Q 95

US · EU

StreamingVisionTool CallingJSON ModeLong Context+1

$10.00 in · $50.00 out / 1M tokens · 1,000,000 ctx

gemini-2.5-flash
Google Gemini
Q 75

US · EU

StreamingVisionTool CallingJSON ModeLong Context+2

$0.30 in · $2.50 out / 1M tokens · 1,000,000 ctx

gemini-2.5-pro
Google Gemini
Q 77

US · EU

StreamingVisionTool CallingJSON ModeLong Context+2

$1.25 in · $10.00 out / 1M tokens · 1,000,000 ctx

ministral-3b-latest
Mistral
Q 65

EU · US

StreamingTool CallingJSON ModeEmbeddings

$0.10 in · $0.10 out / 1M tokens · 128,000 ctx

ministral-8b-latest
Mistral
Q 77

EU · US

StreamingTool CallingJSON ModeEmbeddings

$0.15 in · $0.15 out / 1M tokens · 128,000 ctx

open-mistral-nemo
Mistral
Q 76

EU · US

StreamingTool CallingJSON ModeEmbeddings

$0.15 in · $0.15 out / 1M tokens · 128,000 ctx

mistral-small-latest
Mistral
Q 82

EU · US

StreamingTool CallingJSON ModeEmbeddings

$0.15 in · $0.60 out / 1M tokens · 131,072 ctx

codestral-latest
Mistral
Q 68

EU · US

StreamingTool CallingJSON ModeEmbeddings

$0.30 in · $0.90 out / 1M tokens · 32,000 ctx

magistral-small-latest
Mistral
Q 76

EU · US

StreamingTool CallingJSON ModeEmbeddings

$0.50 in · $1.50 out / 1M tokens · 128,000 ctx

mistral-medium-latest
Mistral
Q 79

EU · US

StreamingTool CallingJSON ModeEmbeddings

$1.50 in · $7.50 out / 1M tokens · 128,000 ctx

mistral-large-latest
Mistral
Q 83

EU · US

StreamingTool CallingJSON ModeEmbeddings

$0.50 in · $1.50 out / 1M tokens · 256,000 ctx

magistral-medium-latest
Mistral
Q 82

EU · US

StreamingTool CallingJSON ModeEmbeddings

$2.00 in · $5.00 out / 1M tokens · 128,000 ctx

command-light
Cohere
Q 51

US · EU

StreamingTool CallingEmbeddingsLong Context

$0.30 in · $0.60 out / 1M tokens · 4,096 ctx

command-r
Cohere
Q 70

US · EU

StreamingTool CallingEmbeddingsLong Context

$0.15 in · $0.60 out / 1M tokens · 128,000 ctx

command-r-plus
Cohere
Q 75

US · EU

StreamingTool CallingEmbeddingsLong Context

$2.50 in · $10.00 out / 1M tokens · 128,000 ctx

grok-3-mini
xAI (Grok)
Q 59

US

StreamingVisionTool CallingJSON ModeReasoning

$0.30 in · $0.50 out / 1M tokens · 131,072 ctx

grok-2
xAI (Grok)
Q 72

US

StreamingVisionTool CallingJSON ModeReasoning

$2.00 in · $10.00 out / 1M tokens · 131,072 ctx

grok-3
xAI (Grok)
Q 79

US

StreamingVisionTool CallingJSON ModeReasoning

$3.00 in · $15.00 out / 1M tokens · 131,072 ctx

amazon.titan-text-premier-v1:0
AWS Bedrock
Q 60

US · EU

StreamingTool CallingVision

$0.50 in · $1.50 out / 1M tokens · 42,000 ctx

meta.llama3-1-70b-instruct-v1:0
AWS Bedrock
Q 72

US · EU

StreamingTool CallingVision

$0.99 in · $0.99 out / 1M tokens · 128,000 ctx

anthropic.claude-3-5-sonnet-20241022-v2:0
AWS Bedrock
Q 80

US · EU

StreamingTool CallingVision

$3.00 in · $15.00 out / 1M tokens · 200,000 ctx

Phi-4
Azure AI Foundry
Q 62

US · EU

StreamingTool Calling

$0.07 in · $0.14 out / 1M tokens · 16,000 ctx

Llama-3.3-70B-Instruct
Azure AI Foundry
Q 73

US · EU

StreamingTool Calling

$0.71 in · $0.71 out / 1M tokens · 128,000 ctx

Mistral-Large-2411
Azure AI Foundry
Q 78

US · EU

StreamingTool Calling

$2.00 in · $6.00 out / 1M tokens · 128,000 ctx

llama3.1-8b
Cerebras
Q 50

US

StreamingTool Calling

$0.10 in · $0.10 out / 1M tokens · 128,000 ctx

llama3.1-70b
Cerebras
Q 71

US

StreamingTool Calling

$0.60 in · $0.60 out / 1M tokens · 128,000 ctx

google/gemma-3-27b-it
Nebius
Q 73

EU

StreamingTool CallingEmbeddings

$0.10 in · $0.20 out / 1M tokens · 131,072 ctx

nvidia/NVIDIA-Nemotron-3-Nano-30B-A3B
Nebius
Q 84

EU

StreamingTool CallingEmbeddings

$0.10 in · $0.30 out / 1M tokens · 131,072 ctx

Qwen/Qwen3-30B-A3B-Instruct-2507
Nebius
Q 82

EU

StreamingTool CallingEmbeddings

$0.10 in · $0.30 out / 1M tokens · 262,144 ctx

Qwen/Qwen3-32B
Nebius
Q 84

EU

StreamingTool CallingEmbeddings

$0.10 in · $0.30 out / 1M tokens · 131,072 ctx

meta-llama/Llama-3.3-70B-Instruct
Nebius
Q 74

EU

StreamingTool CallingEmbeddings

$0.13 in · $0.40 out / 1M tokens · 131,072 ctx

openai/gpt-oss-120b
Nebius
Q 82

EU

StreamingTool CallingEmbeddings

$0.15 in · $0.60 out / 1M tokens · 131,072 ctx

Qwen/Qwen3-235B-A22B-Instruct-2507
Nebius
Q 84

EU

StreamingTool CallingEmbeddings

$0.20 in · $0.60 out / 1M tokens · 40,960 ctx

MiniMaxAI/MiniMax-M3
Nebius
Q 84

EU

StreamingTool CallingEmbeddings

$0.30 in · $1.20 out / 1M tokens · 1,000,000 ctx

moonshotai/Kimi-K2.6
Nebius
Q 85

EU

StreamingTool CallingEmbeddings

$0.50 in · $2.00 out / 1M tokens · 262,144 ctx

nvidia/Nemotron-3-Ultra-550b-a55b
Nebius
Q 83

EU

StreamingTool CallingEmbeddings

$0.60 in · $1.80 out / 1M tokens · 131,072 ctx

NousResearch/Hermes-4-405B
Nebius
Q 68

EU

StreamingTool CallingEmbeddings

$0.50 in · $1.50 out / 1M tokens · 131,072 ctx

Self-hosted / OpenAI-compatible

Point AICP at your base_url (dashboard only, not yet exposed via SDK) and every model your deployment serves becomes available. Any endpoint that speaks the OpenAI chat completions API works, even if it's not one of the named runtimes below. That's what custom is for.

ProviderWhat it isOpen source
vllmHigh-throughput GPU inference server, PagedAttention for max efficiencyYes
ollamaRun LLMs locally on any Mac, Linux, or Windows machineYes
tgi (Hugging Face TGI)Text Generation Inference, optimized server for HF Transformer modelsYes
nvidia-nimOptimized GPU inference microservices for NVIDIA hardwareNo
sglangFast structured generation runtime, RadixAttention for prefix cachingYes
customAny OpenAI-compatible HTTP endpoint, bring your own deploymentNo