Multi-provider orchestration
One API key reaches many model families. Noxery selects the provider and deployment per request, balances load across deployments, and fails over automatically when a provider throttles.
Noxery is an OpenAI-compatible inference gateway: it routes your prompts to large language models from providers such as Azure OpenAI, Alibaba Cloud and Groq, then returns the answers through one stable API.
One API key reaches many model families. Noxery selects the provider and deployment per request, balances load across deployments, and fails over automatically when a provider throttles.
Full Responses and Chat Completions compatibility with streaming, function calling and tool use, so coding agents such as Codex, OpenCode and Claude Code work out of the box — including long multi-turn sessions.
Token-aware scheduling, prompt caching, long-context management and per-plan quotas keep shared capacity responsive: heavy background jobs never starve interactive users.
Every request is metered per token for transparent billing, and upstream content-safety filtering applies on every completion.
Create a free API key and call your first model in minutes. Local pricing and support included.