Menu

How Noxery uses AI

Our AI approach

AI runs through every layer of Noxery

Noxery is an OpenAI-compatible inference gateway: it routes your prompts to large language models from providers such as Azure OpenAI, Alibaba Cloud and Groq, then returns the answers through one stable API.

Multi-provider orchestration

One API key reaches many model families. Noxery selects the provider and deployment per request, balances load across deployments, and fails over automatically when a provider throttles.

Built for AI agents

Full Responses and Chat Completions compatibility with streaming, function calling and tool use, so coding agents such as Codex, OpenCode and Claude Code work out of the box — including long multi-turn sessions.

Fair and reliable under load

Token-aware scheduling, prompt caching, long-context management and per-plan quotas keep shared capacity responsive: heavy background jobs never starve interactive users.

Metered and safe by design

Every request is metered per token for transparent billing, and upstream content-safety filtering applies on every completion.

Build with Noxery

Create a free API key and call your first model in minutes. Local pricing and support included.

Get started Docs