Documentation
Build on one API. Reach every model.
Neural Router is an OpenAI-compatible gateway that routes each request to the best model across hundreds of options by cost, latency, or quality, with governance, observability, and failover built in.
What is Neural Router?
Neural Router sits in front of every major provider and hundreds of models behind a single, OpenAI-compatible endpoint. You send a chat completion request; the router scores eligible models against your objective and policies and dispatches to the best one, automatically failing over if a provider degrades.
Because the API mirrors the OpenAI Chat Completions schema, you can point an existing OpenAI SDK at Neural Router by changing only the base URL and API key, then layer routing, budgets, caching, and residency on top.
Start here
Quickstart
Make your first routed request in minutes.
Provider routing
Load balancing, fallbacks, and provider filters.
Model routing
Auto-select models by objective.
Models
Browse the catalog and pin or restrict models.
Structured outputs
Type-safe JSON responses.
API reference
Endpoints, parameters, and schemas.
Core concepts
- One key, one schema, every major model.
- Route by quality-per-dollar, lowest-cost, lowest-latency, or highest-quality.
- Ordered fallback chains and automatic provider failover.
- Budgets, guardrails, semantic caching, data residency, and an audit log per workspace.