Works with your existing AI stack.

The Go Local AI platform is a cost, environmental and transparency optimisation layer around your existing AI deployment.
Go Local AI modular architecture Applications call a gateway. The gateway can be existing or set up by Go Local AI and sends requests to approved models. Assessment, Routing and Monitoring are independent modules that can be deployed separately, in combination or as the full platform. REQUEST PATH Your applications Agents · tools · workflows AI gateway Keep yours, or we set one up Executes the selected route APPROVED MODEL ESTATE Current provider · OpenAI · Anthropic · other Route between available approved API models Private / open-weight models Other approved APIs ADD THE ENGINE(S) YOU NEED ASSESSMENT Reads workload + cost evidence Model + capacity recommendations ROUTING Returns the live model decision Gateway executes the route Customer-specific classifier MONITORING Reads outcomes + telemetry Cost · quality · energy · carbon decision One engine · any combination · full platform The modules share evidence when deployed together, but none requires the other two. Solid = request / routing decision Dashed = evidence / telemetry Go Local AI modular architecture Applications call a gateway, which can be existing or set up by Go Local AI. The gateway sends requests to approved models. Assessment, Routing and Monitoring can each be deployed separately, in combination or together. REQUEST PATH Your applications Agents · tools · workflows AI gateway Keep yours, or we set one up Executes the selected route APPROVED MODEL ESTATE Current provider OpenAI · Anthropic · other API models Private / open-weight models Other approved APIs ADD THE ENGINE(S) YOU NEED ASSESSMENT Workload, cost + benchmark evidence ROUTING Returns the live model decision Gateway executes the route MONITORING Outcomes, cost + telemetry One engine · any combination · full platform Modules share evidence when used together. Solid = request path · engine links omitted here for clarity

Solid lines carry requests and routing decisions. Dashed lines carry evidence and telemetry.

Live requests and
background optimisation

The request path.

Your application calls the gateway. The gateway sends a copy of the request to the Routing Engine, receives back a model choice and sends the original request to the model. The classifier is not the model producing the answer.

The improvement path.

Usage, costs and outcomes feed Assessment and Monitoring. They generate deployment recommendations and prepared classifier-training datasets. A classifier model update is tested separately before an approved release changes production routing.

What connects where.

View connection matrix
Connection What it carries Used by
Billing and workload data Model and provider usage, token costs, gateway logs, agreed work samples and infrastructure information. Assessment establishes a baseline and benchmarks model and deployment options.
Gateway routing interface The request to classify, followed by the chosen approved destination returned to the gateway. Routing makes the model decision. The gateway remains responsible for executing it.
Gateway and quality outcomes Selected models, request costs, task-quality signals, routing failures and fallback reasons. Monitoring attributes outcomes; selected records feed customer-specific classifier evaluation and training.
Infrastructure telemetry GPU and host readings alongside grid carbon-intensity data, PUE and stated measurement coverage. Monitoring connects workload activity with energy and estimated carbon; Assessment uses the evidence for deployment recommendations.
Approved changes The reviewed model mix, routing policy, deployment plan or tested classifier version. The agreed production rollout. A recommendation is not an unapproved server or model change.

One engine, a combination,
or the full platform.

Assessment on its own

Start with billing, logs and workload samples. Get the comparison and recommendations without changing the live request path.

Routing on its own

Use a general or tuned classifier with your gateway and model deployment. Supply the workload examples and evaluation criteria needed to train and test it.

Monitoring on its own

Connect gateway, quality and infrastructure evidence for cost, routing and environmental visibility, without requiring a Go Local AI classifier.

Connect all three

Monitoring triggers a new assessment. Its results update the deployment recommendation or supply new training data for the Routing Engine.

Keep what already works.

Stay with your model provider.

Route between models within the provider you already use. Other providers and private open-weight models can be added where the savings, workload and data rules justify them.

Integrates into your gateway.

The Platform is fully compatible with Bifrost, LiteLLM, Portkey or your own gateway where its interfaces support the connection.

No gateway or server capacity?

We can set up a gateway. When you lack suitable inference capacity, or reasons to run it yourself, managed hosting can run the platform, open-weight models, or both.

Find out what your AI
should be costing you.