See where your AI agents spend your time and money.

Your developers' agents make thousands of API calls you never see. Point them at one Ganox endpoint and see it all: which models they call, how many requests, what kind of work, where the caching goes, and what it costs. Then shape and cap it to run leaner and cut what you don't need.

What you get    one usage panel for all your agents · example from our own traffic

22,817
Requests
3.03B
Tokens processed
$4,110
Spend

Where the spend goes

Cache reads 51% Cache writes 26% Thinking 15% 8%
Cache reads $2,116 Cache writes $1,080 Thinking $605 Output + fresh input + other models $309
Opus 97% Fable 2% Haiku 1% Sonnet 0% 4 models · itemized to the cent

every agent → one endpoint → provider, on your key · 77% of this bill is cache, 3% is fresh input

What Ganox does · each lever labeled live, shadow, or idea

01 · the point

Visibility

One dashboard for every agent, readable across the whole team.

what you see
Models, in what mixLiveEvery model your agents call and how spend splits across them.
Requests & API callsLiveHow many calls, per agent and per developer, over any window.
Request typesLiveGrouped by intent: edits, file reads, command runs, research.
Caching opportunityLiveHow much of each session is cached versus re-billed.
Spend, to the centLiveBy developer, model, and cache, on real provider-billed cost.

02 · protect the cache

Caching

Keep the provider's prompt cache warm so long agent sessions are not re-billed for their own history.

levers
Deep cache anchorsLiveSpreads cache markers so 1000+-block agent prefixes stay a cheap read.
Prompt & reply compressionShadowTrim unneeded input and output while keeping the same answer. Measuring now.
Smarter cache reuseIdeaFirst-breakpoint caching, TTL tuning, prefix normalization, history summarization.

03 · shape the output

Replies

Steer the prose your agents get back, which also trims output cost. Prose only, code untouched, opt-in.

levers
Writing stylesLivePlain, terse, neutral, set per user.
Skills: caveman, i-have-adhdLiveOpt-in output styles for shorter, scannable replies.

04 · right-size the model

Models

Match the model to the request instead of paying top rates for everything. Under evaluation, not routing yet.

levers
Cost judgeShadowTags each request's difficulty and the cheapest model that could serve it.
Cheaper-model routingIn evalSend easy work to a cheaper model. Nothing routed yet.

05 · cap the spend

Spend control

A hard ceiling, because all traffic flows through one endpoint. Metered on real, provider-billed cost.

levers
Org & per-developer capsLiveWarn at 80%, hard pause at 100%, no overage. Alerts-only mode too.
Full itemizationLiveEvery call by model, request type, caching, and cost.

06 · what stays yours

Your key, your data

Requests pass through untouched unless you opt into a lever above.

guarantees
Key forwarded, never storedLiveYour provider key is passed through on each request, never written to disk.
Traffic not retained or soldLiveYour prompts and completions are not kept or resold.

One endpoint in front of every agent turns invisible spend into something you can see, shape, and cap.

Point your agents at Ganox with one base-URL change, keep your own key, and you can see which models, how many calls, what kind of work, and what it costs, across every developer. The levers (caching, reply styles, cheaper-model routing) are opt-in and labeled by status, so you always know what is live. The cap is a hard pause.

request an invite below · free during the invite-only beta

Get access

Request an invite.

We are opening access gradually. Tell us about your setup and we will be in touch.

  • Full visibility: every call by model, type, and cost.
  • Drop-in: one base-URL change, no code rewrite.
  • No surprise bills: warn at 80%, pause at 100%.
  • Your data stays yours: key never stored, traffic never sold.