The request path.
Your application calls the gateway. The gateway sends a copy of the request to the Routing Engine, receives back a model choice and sends the original request to the model. The classifier is not the model producing the answer.
Solid lines carry requests and routing decisions. Dashed lines carry evidence and telemetry.
Your application calls the gateway. The gateway sends a copy of the request to the Routing Engine, receives back a model choice and sends the original request to the model. The classifier is not the model producing the answer.
Usage, costs and outcomes feed Assessment and Monitoring. They generate deployment recommendations and prepared classifier-training datasets. A classifier model update is tested separately before an approved release changes production routing.
| Connection | What it carries | Used by |
|---|---|---|
| Billing and workload data | Model and provider usage, token costs, gateway logs, agreed work samples and infrastructure information. | Assessment establishes a baseline and benchmarks model and deployment options. |
| Gateway routing interface | The request to classify, followed by the chosen approved destination returned to the gateway. | Routing makes the model decision. The gateway remains responsible for executing it. |
| Gateway and quality outcomes | Selected models, request costs, task-quality signals, routing failures and fallback reasons. | Monitoring attributes outcomes; selected records feed customer-specific classifier evaluation and training. |
| Infrastructure telemetry | GPU and host readings alongside grid carbon-intensity data, PUE and stated measurement coverage. | Monitoring connects workload activity with energy and estimated carbon; Assessment uses the evidence for deployment recommendations. |
| Approved changes | The reviewed model mix, routing policy, deployment plan or tested classifier version. | The agreed production rollout. A recommendation is not an unapproved server or model change. |
Route between models within the provider you already use. Other providers and private open-weight models can be added where the savings, workload and data rules justify them.
The Platform is fully compatible with Bifrost, LiteLLM, Portkey or your own gateway where its interfaces support the connection.
We can set up a gateway. When you lack suitable inference capacity, or reasons to run it yourself, managed hosting can run the platform, open-weight models, or both.