Self Host LLM Gateway
Deploy LLM Gateway on your own infrastructure with Docker, Docker Compose, or Kubernetes on AWS, GCP, and Azure.
LLM Gateway is a self-hostable platform that provides a unified API gateway for multiple LLM providers. Run it on your own infrastructure to keep full control over your data and avoid platform fees.
Pick the deployment path that matches where you're running it.
Single host
Which option should I choose?
- Trying it out or running a single low-traffic instance? Start with Docker or Docker Compose on one machine.
- Running in production? Deploy to Kubernetes with our Helm chart, and use a managed Postgres and Redis from your cloud.
- On a specific cloud? Follow the AWS, Google Cloud, or Azure guide for the exact managed services to provision.
- Unlocking Enterprise features? Install the signed token using the Enterprise licensing guide.
What you'll need
Every deployment is built from the same pieces:
- Stateless services — the gateway, API, UI, and a background worker. Scale these freely; they hold no data between requests.
- PostgreSQL — the source of truth for users, projects, keys, and usage records.
- Redis — response caching and the queue that feeds the worker.
- Provider API keys — the OpenAI, Anthropic, Google, and other credentials the gateway uses, injected as secrets.
In production, run PostgreSQL and Redis as managed services from your cloud provider so backups, failover, and patching are handled for you.
Product links
For self-hosted DevPass and Airside, set UI_URL and PLAYGROUND_URL to your Gateway dashboard and Lounge origins. When unset, these apps use https://llmgateway.io and https://lounge.llmgateway.io; with NODE_ENV=development, they use http://localhost:3002 and http://localhost:3003.
Admin dashboard access
Grant admin dashboard access with comma-separated email lists on the API service. Only verified email addresses count.
| Variable | Role | Access |
|---|---|---|
ADMIN_FULL_ACCESS_EMAILS | Admin | Full access |
ADMIN_SUPPORT_EMAILS | Support | Read-only, plus issuing DevPass refunds |
ADMIN_VIEWER_EMAILS | Viewer | Read-only |
Support and viewer accounts do not see platform-wide revenue or profit: the dashboard home, Global Stats, LLM SDK, DevPass and Lounge plan KPIs, and provider credential spend are hidden, and the gateway's own margin and fee fields — including Airside carrier margins and routing adjustments — are removed from every response. A subscriber's plan margin stays visible, since it is the plan price minus usage these roles already see. Request payloads in logs are shown unmodified.
Gateway load
The admin dashboard's Gateway Load → Requests view reports peak throughput from completed intervals: minutes for short model/provider views, and hours for longer ranges or tenant-filtered views. Daily charts retain hourly peaks, so switching from 7 to 30 days does not smooth away a spike. Peak timestamps identify the start of the interval; the current incomplete interval is excluded.
Switch to Tokens for platform, model, or provider throughput over 7 or 30 days. Choose total, input, or output tokens and a billing mode. Average TPM includes quiet minutes; average tokens/day normalizes the same window to 24 hours. Peak TPM is the busiest completed UTC minute, and peak tokens/day is the busiest complete UTC calendar day. Partial days appear faded in the daily chart and do not count toward daily peaks. The table ranks the top ten models or providers by token volume and supports CSV export.
Token metrics use recorded upstream usage, excluding gateway response-cache hits. Minute history is retained for 30 days. Measurements follow worker refreshes; use Refresh to update them. These fixed UTC intervals do not reproduce a provider's rolling rate-limit or quota accounting. Organization, project, and API-key filters apply only to the Requests view.
Support chat diagnostics
Support chat errors can arrive inside an HTTP 200 event stream. In the API
logs, search for the conversation's conversationId, then follow requestId
for each visitor attempt and gatewayRequestId into the gateway logs.
Chat support request finished records whether the attempt completed, failed,
returned an empty answer, was rejected, was aborted, or stayed silent because
the conversation was escalated. Failure events include the stage and available
provider error code or HTTP status. Gateway and step events record timing,
selected models, token counts, and tool-call counts without recording chat
contents or raw SDK error payloads.
Sign-up email restrictions
With HOSTED=true, administrators can manage the custom email-domain blocklist in
Admin → Settings → Blocked sign-up email domains. Paste domains one per line
or separated by commas, then save. Entries also block subdomains; updates apply
to subsequent email sign-ups without a restart. Existing users can still sign in.
The custom list starts empty and has no built-in entries. Clearing and saving it removes custom domain blocks. The separate disposable-email and plus-address checks remain active in hosted mode.
How is this guide?
Last updated on