Providers & Models
AICP routes to 23 providers across four categories: commercial, European (GDPR-native), inference platforms, and self-hosted/OpenAI-compatible. Every request goes through one endpoint; which provider actually serves it is decided by your routing configuration and by which providers are connected (platform-wide by your admin, or by you via BYOK).
BYOK: bring your own key
Connect your own provider credentials and they're used for real inference immediately. Your own key always takes priority over any provider your admin has enabled platform-wide, for that same provider. Available on every plan, including Community (self-hosted).
await client.providers.add({ provider: 'openai', apiKey: 'sk-...' });
Models are discovered automatically from the provider's own /models endpoint, so you don't
need to list them yourself. See the full guide for managing, testing, and disabling
providers: JavaScript · Python.
client.providers.add() only takes a single apiKey right now. Four providers,
AWS Bedrock, AWS SageMaker, Azure OpenAI, and Azure AI Foundry, authenticate
with structured, multi-field credentials (SigV4 keys, resource/deployment names, API
versions) instead of a single key, and self-hosted/custom endpoints need a base_url. Both
are only exposed through the dashboard's Providers tab right now, not the SDKs. See the
Providers tab below for exactly which fields each one needs.
Browse the catalog
The same search and filters as the product's own Model Catalog page. This is a static snapshot for the public docs (there's no org/BYOK context here), so connection status isn't shown. Self-hosted providers (vLLM, Ollama, TGI, NVIDIA NIM, SGLang, Custom Endpoint) have no fixed model list by design: the model is whatever you've deployed there, so they aren't part of the searchable catalog below; see the self-hosted section underneath it.
50 models across 11 providers with pre-registered catalog entries. Self-hosted providers add whatever you deploy
$0.50 in · $1.50 out / 1M tokens · 16,385 ctx
$0.15 in · $0.60 out / 1M tokens · 128,000 ctx
$2.50 in · $10.00 out / 1M tokens · 128,000 ctx
$1.10 in · $4.40 out / 1M tokens · 200,000 ctx
$15.00 in · $60.00 out / 1M tokens · 200,000 ctx
$0.15 in · $0.60 out / 1M tokens · 128,000 ctx
$2.50 in · $10.00 out / 1M tokens · 128,000 ctx
$1.00 in · $5.00 out / 1M tokens · 200,000 ctx
$3.00 in · $15.00 out / 1M tokens · 1,000,000 ctx
$3.00 in · $15.00 out / 1M tokens · 1,000,000 ctx
$5.00 in · $25.00 out / 1M tokens · 1,000,000 ctx
$5.00 in · $25.00 out / 1M tokens · 1,000,000 ctx
$5.00 in · $25.00 out / 1M tokens · 1,000,000 ctx
$10.00 in · $50.00 out / 1M tokens · 1,000,000 ctx
$0.30 in · $2.50 out / 1M tokens · 1,000,000 ctx
$1.25 in · $10.00 out / 1M tokens · 1,000,000 ctx
$0.10 in · $0.10 out / 1M tokens · 128,000 ctx
$0.15 in · $0.15 out / 1M tokens · 128,000 ctx
$0.15 in · $0.15 out / 1M tokens · 128,000 ctx
$0.15 in · $0.60 out / 1M tokens · 131,072 ctx
$0.30 in · $0.90 out / 1M tokens · 32,000 ctx
$0.50 in · $1.50 out / 1M tokens · 128,000 ctx
$1.50 in · $7.50 out / 1M tokens · 128,000 ctx
$0.50 in · $1.50 out / 1M tokens · 256,000 ctx
$2.00 in · $5.00 out / 1M tokens · 128,000 ctx
$0.30 in · $0.60 out / 1M tokens · 4,096 ctx
$0.15 in · $0.60 out / 1M tokens · 128,000 ctx
$2.50 in · $10.00 out / 1M tokens · 128,000 ctx
$0.30 in · $0.50 out / 1M tokens · 131,072 ctx
$2.00 in · $10.00 out / 1M tokens · 131,072 ctx
$3.00 in · $15.00 out / 1M tokens · 131,072 ctx
$0.50 in · $1.50 out / 1M tokens · 42,000 ctx
$0.99 in · $0.99 out / 1M tokens · 128,000 ctx
$3.00 in · $15.00 out / 1M tokens · 200,000 ctx
$0.07 in · $0.14 out / 1M tokens · 16,000 ctx
$0.71 in · $0.71 out / 1M tokens · 128,000 ctx
$2.00 in · $6.00 out / 1M tokens · 128,000 ctx
$0.10 in · $0.10 out / 1M tokens · 128,000 ctx
$0.60 in · $0.60 out / 1M tokens · 128,000 ctx
$0.10 in · $0.20 out / 1M tokens · 131,072 ctx
$0.10 in · $0.30 out / 1M tokens · 131,072 ctx
$0.10 in · $0.30 out / 1M tokens · 262,144 ctx
$0.10 in · $0.30 out / 1M tokens · 131,072 ctx
$0.13 in · $0.40 out / 1M tokens · 131,072 ctx
$0.15 in · $0.60 out / 1M tokens · 131,072 ctx
$0.20 in · $0.60 out / 1M tokens · 40,960 ctx
$0.30 in · $1.20 out / 1M tokens · 1,000,000 ctx
$0.50 in · $2.00 out / 1M tokens · 262,144 ctx
$0.60 in · $1.80 out / 1M tokens · 131,072 ctx
$0.50 in · $1.50 out / 1M tokens · 131,072 ctx
Self-hosted / OpenAI-compatible
Point AICP at your base_url (dashboard only, not yet exposed via SDK) and every model your
deployment serves becomes available. Any endpoint that speaks the OpenAI chat completions
API works, even if it's not one of the named runtimes below. That's what custom is for.
| Provider | What it is | Open source |
|---|---|---|
vllm | High-throughput GPU inference server, PagedAttention for max efficiency | Yes |
ollama | Run LLMs locally on any Mac, Linux, or Windows machine | Yes |
tgi (Hugging Face TGI) | Text Generation Inference, optimized server for HF Transformer models | Yes |
nvidia-nim | Optimized GPU inference microservices for NVIDIA hardware | No |
sglang | Fast structured generation runtime, RadixAttention for prefix caching | Yes |
custom | Any OpenAI-compatible HTTP endpoint, bring your own deployment | No |