Three engines.
Use what you need.

Compare models, route requests and track results. Use the engines separately or together.

Find what your work
actually needs.

The Assessment Engine compares models on your work and recommends where to run them. It checks cost, quality, data requirements and the hardware each option needs.

Assessment Engine: from usage data to a deployment plan Ingest gateway, billing and workflow data. Anonymise, clean, classify and map privacy constraints. Benchmark current, cloud and open-weight models. Blind quality tests lead to a cost baseline, validated savings estimate and model and routing architecture. INGEST Gateway logs Model + provider usage Billing + tokens Spend + volume Prompts + workflows Tasks + infrastructure Anonymise. Clean. Classify. Privacy constraints + cost and usage baseline Current setup Provider + model Cloud alternatives Approved APIs Open-weight models Private inference BENCHMARK Blind quality testing Your quality threshold Model + routing architecture Savings estimate · impact projection · migration roadmap Assessment Engine: from usage data to a deployment plan Ingest gateway, billing and workflow data. Anonymise, clean, classify and map privacy constraints. Benchmark current, cloud and open-weight models. Blind quality tests lead to a cost baseline, validated savings estimate and model and routing architecture. INGEST Your usage data Gateway logs · billing · tokens Prompts · workflows · infrastructure STRUCTURE Anonymise. Clean. Classify. Map privacy constraints Build a cost + usage baseline ANALYSE Benchmark the alternatives Current setup · cloud models Open-weight models Local vs API economics Blind quality testing Your quality threshold RECOMMEND Your deployment plan Model + routing architecture Savings + impact estimates Migration roadmap

From source data
to a decision.

Collect evidence

The Engine reads API and gateway logs, provider usage, billing and token data, alongside prompts, workflows and infrastructure information. It establishes which models handle the work and what they cost.

Prepare the data

The Engine anonymises sensitive records, clean and deduplicate the data, and group similar workloads. It maps the privacy constraints and quality requirements before comparing the options.

Compare the options

The Engine benchmarks current, cloud and open-weight models on the same tasks. It blind-tests their responses and compare API costs with the infrastructure needed to run private AI inference.

Recommend a setup

The Engine produces model and routing choices, a costed deployment plan and an environmental projection. It recommends keeping the existing setup where a change does not improve the result.

Insights beyond price per token.

View comparison
Requirement What is compared What the recommendation answers
Cost and usage Provider bills, token use, workload volumes and infrastructure costs. Where to reduce spend within the current provider, route between models or run suitable work privately.
Output quality Current and candidate responses on the same work, tested blind against the task’s agreed threshold. Which alternatives meet the requirement, rather than simply costing less.
Infrastructure fit The workload mix and volume against the model and inference capacity needed to serve it. Whether existing capacity still fits, a different model mix is needed, or private inference is justified.
Data and impact Approved data routes, compute requirements, energy coverage and grid carbon intensity. Where the work may run and the estimated operational impact of each suitable option.

A plan you can
act on.

The output connects the evidence to a specific model, routing and deployment recommendation. Savings and environmental projections remain estimates until checked in production.

  • Savings, environmental projections and an optimisation or deployment plan.
  • Model mix, routing architecture and inference-capacity recommendations
  • Benchmark results, blind test comparisons, and all the data and evidence collected

A classifier trained
on your work.

The Routing Engine uses a small model trained on your work to choose the lowest cost suitable AI model for each request.

How the classifier improves
Routing Engine: a customer-tuned add-in for your gateway The gateway asks the customer-tuned classifier for a routing decision. Data boundaries, workload requirements and model suitability determine the route. The gateway sends the request to a private small model, private large model or approved cloud API. LIVE REQUEST ROUTING Your applications Your gateway Bifrost · LiteLLM · other TRAIN + EVALUATE GO LOCAL AI Your routing classifier Data boundaries Workload + quality Suitable model + cost CLASSIFY DECISION Private small Routine work Private large More demanding work Cloud / API Approved fallback LOWEST-COST SUITABLE APPROVED DESTINATION Routing Engine: a customer-tuned add-in for your gateway The gateway asks the customer-tuned classifier for a routing decision. Data boundaries, workload requirements and model suitability determine the route. The gateway sends the request to a private small model, private large model or approved cloud API. LIVE REQUEST ROUTING Your applications Your gateway Bifrost · LiteLLM · other CLASSIFY DECISION GO LOCAL AI Your routing classifier Data boundary enforcement Workload + quality requirements Suitable model + cost GATEWAY SENDS THE REQUEST Private small Routine work Private large More demanding work Cloud / API Approved fallback

What happens
to a request.

Understand the request

The classifier model identifies the task, complexity and data sensitivity within the prompt or request.

Enforce data boundaries

The classifier model uses defined sensitive data classifications when deciding destinations.

Select the model

The classifier model chooses the lowest cost or carbon option that meets the quality requirements between the available models.

Return the decision

The classifier model sends the decision back to the gateway. The Routing Engine records the selected model, confidence, outcome and any fallbacks.

When a request
needs more care.

View difficult cases
Situation How routing handles it What feeds back
Sensitive data/content Restrict the request to approved private destinations. Block it when no compliant destination is available. The sensitivity decision and any blocked or failed route.
Unknown or low-confidence work Use the agreed safe route, escalate within policy, or block where no compliant route exists. Examples of the requests the classifier needs to learn, including wrong or unnecessary escalations.
A poor routing outcome Record failed requests, quality issues and routing failures for evaluation. Prepared and anonymised examples for the next client specific classifier fine-tune.

See the real impact of your AI use

The Monitoring Engine tracks spend, quality, carbon and routing by task, team and prompt.

Monitoring Engine: cost, quality, energy and carbon by workload Gateway logs, GPU and host telemetry, and grid data feed the Monitoring Engine. It tracks cost, quality, energy and estimated operational carbon, along with routing and escalation. Results are attributed by task, team, division and prompt. Changes feed back to assessment. THREE DATA SOURCES Gateway logs Usage + routing outcomes GPU / host telemetry Compute + energy Grid data Carbon intensity Monitoring Engine Task · team · division · prompt £ Cost Spend ✓ Quality Outcomes Wh Energy Telemetry CO₂e Carbon Estimate Routing + escalation Model selected · fallback + reason Changes return to Assessment Monitoring Engine: cost, quality, energy and carbon by workload Gateway logs, GPU and host telemetry, and grid data feed the Monitoring Engine. It tracks cost, quality, energy and estimated operational carbon, along with routing and escalation. Results are attributed by task, team, division and prompt. Changes feed back to assessment. DATA SOURCES Gateway + infrastructure data Gateway logs · GPU / host telemetry Grid carbon intensity Monitoring Engine Task · team · division · prompt £ Cost Spend ✓ Quality Outcomes Wh Energy Telemetry CO₂e Carbon Estimate Routing + escalation Model selected · fallback + reason Changes return to Assessment

Data to actionable information.

Collect

The Engine reads gateway logs, request costs, selected models and fallback reasons. It combines those records with quality signals, GPU and host telemetry, and live power grid data.

Attribute

The Engine breaks activity down by task, team, division and prompt. See where spend and compute are concentrated, rather than only the total provider bill.

Compare

The Engine checks task-level outcomes against their quality baseline. It watches workload mix, routing success, costs and infrastructure use as they change.

Trigger

The engine flags material changes to the Assessment Engine for a new recommendation, or supplies data for classifier updates when routing decisions need improvement.

What the platform tracks

View measures
Measure Recorded What you can act on
Cost Request, task and team spend alongside model and provider usage. The workloads driving the bill and whether a routing or deployment change reduces their cost.
Quality Task outcomes against the agreed baseline, including quality drift. Work that no longer meets the requirement, even when its model is cheaper.
Routing and escalation Selected models, failed requests, fallbacks and their reasons. Unknown tasks, wrong escalations and routing failures that need new examples or a policy review.
Energy Available GPU and host energy telemetry and location PUE. Where compute is being used, and evidence for comparing inference-capacity choices.
Carbon Derived from energy and grid carbon-intensity data or known site energy provenance. The operating environmental impact of the workload, with sources, coverage and assumptions visible.

How the platform can use gathered evidence.

Automated infrastructure recommendations.

Changed volumes, costs, model choices or capacity use trigger a fresh assessment. It recommends an updated model mix, routing approach or inference deployment, with the expected cost and impact.

See the recommendations

A classifier that learns from its mistakes.

Quality issues, routing failures and unfamiliar work feed a new training data set. Automatic fine-tuning and evaluation produce a candidate update for shadow testing and approval.

See how it improves

Find out what your AI
should be costing you.