Estate spend · June · estate: 12 workspaces $0.00 ▲ 18% MoM · $38 median / dev
Reconciled vs invoice +0% within ±1% tolerance
Workspaces · seats 0 · 0 4 providers metered

Your costs are
logged. Now make
them legible.

Share reports with clients, reconcile against provider invoices, and track spend across your whole team, from a hosted dashboard you don't have to run.

Get started See it live curl -fsSL https://get.haltonmeter.com/install.sh | sh
app.haltonmeter.com DEMO · INTERACTIVE

01 · Instruments · sample workspace

LLM cost attribution,
laid out for a quarterly review.

Charts like these are drawn straight from your log ledger — spend by project, daily rhythm, provider mix, reconciliation deltas. The same numbers your daemon prints, structured so finance can read them without a terminal.

$246.20 logged MTD · 6 projects · 4 providers · 3,412 requests · +0.30% variance (June) · workspace: meridian-studio · estate: 12 workspaces

Instrument 01

Per-project LLM cost.

Every request is tagged with a project slug at capture time. The daemon infers attribution automatically from the working directory. No manual labelling required.

$246.20 Total this month
6 Active projects
$8.21 Daily average
$82.40 Top project
Project Trend USD Share
contract-ai $82.40 33.5%
document-processor $61.25 24.9%
chat-support-bot $38.90 15.8%
invoice-assistant $24.70 10.0%
code-reviewer $21.15 8.6%
research-agent $17.80 7.2%
Total $246.20 100%

Daily API spend · July 2026 (1–15 Jul)

5 Jul10 Jul15 Jul
Peak: $42.15 Avg: $22.38

Instrument 02

Daily cost rhythm.

Spend doesn't arrive uniformly. Spikes signal batch runs, test cycles, or new agents being evaluated. The daily chart surfaces those rhythms so you can correlate cost with delivery milestones before the invoice arrives.

The daemon buffers locally and flushes to the cloud API every 60 seconds. These numbers are never more than a minute behind reality.

Instrument 03

How it lines up with the provider invoice.

The numbers your daemon computes from your LLM API calls: token cost tracking reconciled against the totals each provider charges you at month-end. Targeted variance is <1%. Anything over and the workspace flags it for review.

LOGGED · June 2026 $279.88
BILLED $279.04
VARIANCE +$0.84 / 0.30%
TOLERANCE ±1.00%

Every figure traces to a rate table you can inspect — the same one reconciliation uses. Full month-by-month reconciliation available on Business+.

Sonnet 5, GPT-5.6, and Grok 4.5 were priced within days of launch. New models never show up as $0.00.

Display spend in your team's currency — daily ECB-backed rates, and the rate used is recorded on every converted figure.

PROVIDER MIX · COST SHARE

What's running

TOTAL · JUNE 2026 $279.88 4 providers
Anthropic 65.2% $182.55
OpenAI 20.4% $57.10
Gemini 10.9% $30.51
xAI (Grok) 3.5% $9.72

ADVISORY · WORKSPACE: MERIDIAN-STUDIO

chat-support-bot

71% of requests run Opus 4.7 on prompts under 400 tokens.

Simulated on Haiku 4.5: −$18.40/mo at current volume.

Marked done 3 Jul · realised −$16.90 over the 14-day window.

Instrument 04

The meter that finds the savings.

Twelve optimization rules read your live traffic. Each advisory simulates the cheaper counterfactual, then tracks realised savings for 14 days after you act. Anomalies and forecasts surface before the invoice does.

Included at every tier — optimization advisories, anomalies, and forecasts ship on Solo, Team, and Business+.

02
How it sits in your stack · the request path

Your latency is your
provider's latency.

Halton Meter is a local pass-through, not a gateway. Your request goes straight to Anthropic (Claude), OpenAI, Google Gemini, and xAI (Grok). The meter reads it on the way past and computes cost from what it sees. There's no remote hop on your critical path, and nothing about how you call the model changes. If the daemon stops, your traffic falls straight through to the provider, untouched.

  • Runs on your machine. The proxy is local. Your call reaches the provider directly. It is not routed through Halton Labs infrastructure and back, so there is no remote hop to add.
  • Fail-open by design. If the daemon crashes or is shut off, requests pass through to the provider exactly as before. Observation is never a single point of failure for your work.
  • No SDK wrapper. We don’t intercept your client library or inject middleware. Claude Code, the Anthropic SDK, and everything else keep working unchanged.
Most cost tools sit in your request path. We sit beside it.
03
ARTIFACTS · CLIENT-READY

One click. PDF a client
would actually read.

Reports pull from your live ledger. No re-keying. Personal summaries for retros. Client invoices for billable AI work. Executive P&L for the leadership team. Generated, downloaded, shared with link expiry.

GENERATED IN 2.3s
Halton Labs
LLM COST REPORT
Period
June 2026
r-2026-06-01-pers
TOTAL
$279.88
June 2026
REQS
6,735
6 projects
RECON
0.30%
within tolerance
Model Reqs USD
Opus 4.7620$118.30
Sonnet 4.62,840$52.15
Haiku 4.51,905$12.10
GPT-5.5 Pro780$57.10
Gemini Flash480$30.51
Grok 4110$9.72
Download the sample PDF Clara GmbH · July 2026 · 157 KB

QUICK TEMPLATES

Personal summary POPULAR
Solo retros · 4 pages · 180 KB
Client invoice
Per-client billable AI · denominated in USD
Executive summary TEAM
Workspace P&L · MoM trends
Custom CSV
Pick rows, columns, period
Report wizard BUSINESS+
Compose your own — pick the sections the CFO asked for
Hosted, not local. Run reports against the team's full month. No waiting on your laptop.
Schedule monthly emails. The 1st of every month, your inbox gets last month's PDF.
Share links expire. Each generates a per-recipient URL with a configurable TTL and download audit.
04
Policies · detect, warn, or block

Record and Warn never touch the request path. Block stops the call locally.

Policies are defined in Cloud and enforced where they must be. Record, the default, detects, attributes, and logs. Warn alerts from Cloud and never touches the daemon. Block is enforced locally by the daemon, fail-safe, with no Cloud round-trip. If Cloud is unreachable, the block still applies. Record and Warn are broadly available; hard Block enforcement is Business+, opt-in per workspace.

Policy engine
  • Workspace spend cap resets monthly · $5,000 / mo
    WARN
  • Per-model cap Opus 4.7 · highest-cost model · $3,000 / mo
    WARN
  • Token-rate limit burst protection · 2.0M tok / min
    WARN
  • Model restrictions allow-list · Claude · GPT · Gemini
    WARN

Alert, never block.

Cloud emails the workspace owner and raises an in-app alert. The proxy is never touched. Traffic flows exactly as before.

142 requests would alert
$1,240 over-cap flagged

Recent alerts

  • contract-ai exceeded its $25 budget → #ai-spend (Slack) 3d ago
  • Opus rate spike · 2.4M tok/min 5d ago
  • Grok call · off allow-list 6d ago

Budget breaches land in the right Slack channel — scoped per workspace or org. Email is the default; Slack channels are the additional sink.

05
Coverage · one proxy, three altitudes

One proxy.
Three altitudes.

The same local daemon feeds every tier. You climb from a single developer's cost view to a multi-workspace estate without changing a line of how you call the model.

Altitude 01

Solo

One developer. The full cost view.

  • Overview KPIs, daily spend, Pareto
  • Projects detail, model split, CSV
  • Optimization advisories
  • Reports personal + client invoice
  • Settings keys, currency, rates
Altitude 02

Team

Adds collaboration across one shared workspace.

Everything in Solo, plus

  • Traces full request receipts
  • Teams budgets, rosters, activity
  • Members per-person spend
  • Machines live sync health — active, stale, offline at a glance
  • Alerts budget breaches land in the right Slack channel
  • Clusters client invoicing, margin at a glance
  • Reports executive summary PDF

06 · Pricing · open daemon, hosted Cloud

Daemon stays free.
Cloud makes it legible.

The daemon stays free. Cloud adds hosted persistence, team access, PDF reports, and provider reconciliation. Cancel any time; your daemon keeps working.

Solo

$16

per month · billed annually ($192/yr)

save 16% vs monthly

One developer. Full cost visibility across Anthropic (Claude), OpenAI, Google Gemini, and xAI (Grok). 1 seat.

What's included

  • 1 seat · 2 machines · 7-day traces
  • Overview
  • Projects (list + detail)
  • Reports workbench
  • Reports: Personal summary
  • Reports: Client invoice
  • Optimisation (recommendations)
  • Settings (account, workspace, API, billing)
Get Solo →

Business+

By enquiry

annual · invoiced

Unlimited seats and the full governance layer: provider reconciliation, audit, cost allocation, model policies, SSO, and a multi-workspace estate. Scoped and priced to your org.

What's included

  • Everything in Team
  • Reconciliation & cost
  • Month-by-month reconciliation against provider invoices
  • Cost centres + chargeback allocation
  • Executive P&L dashboard
  • Governance & audit
  • Append-only audit log, with SIEM export
  • Model policies: Record / Warn / Block. Hard Block enforced locally by the daemon, fail-safe
  • Your negotiated provider rates — reports show contracted cost, not list price
  • Identity & estate
  • SSO: SAML 2.0 with SCIM provisioning
  • Org HQ: multi-workspace estate
  • Custom roles — down to who may run the daemon
  • Per-developer budgets (warn-only)
Get in touch →

Daemon

The daemon runs standalone: local-first LLM observability, no Cloud account required.

haltonmeter.com →
$ curl -fsSL https://get.haltonmeter.com/install.sh | sh $ uvx halton-meter

On-premises deployment? Full data residency, your VPC, your keys.

View enterprise options →

Every plan works with the free daemon. Your data is yours; export anytime.

FAQs →

07 · Enterprise · self-hosted Cloud

When the data can't leave
the building.

Some workloads come with data-residency requirements that no public SaaS can satisfy. Halton Meter ships as a self-hosted distribution: the same backend and dashboard, running in your VPC under your keys, with no traffic routed through Halton Labs infrastructure.

  • Data residency Your VPC, your keys

    Deploy via Helm chart or Terraform module. Postgres, the API server, and the dashboard all run inside your perimeter. No outbound data. No shared tenant infrastructure.

  • Identity SAML, SCIM, audit log

    Connect Okta, Azure AD, or JumpCloud via SAML 2.0. Provision and deprovision teams automatically via SCIM. Every login, export, and policy change lands in an append-only audit log.

  • Support Dedicated channel, SLA

    Shared Slack channel with direct access to the engineering team. 99.9% uptime SLA on the hosted components. Annual security review with findings report.

Architecture · self-hosted deployment

Daemon Developer machine · mitmproxy
Local
Cloud API FastAPI · Postgres · your VPC
Your VPC
Dashboard Next.js · internal DNS
Your VPC
Reports Finance · clients · audit
Stays internal

08 · Questions

Frequently asked.

09 · The handshake

Pair your daemon.
Watch the readings populate.

The daemon keeps your readings. Cloud makes them legible to everyone else.

  1. Paste one code. halton-meter cloud connect prints XXXX-XXXX
  2. Approve in the browser. enter the code on the /connect screen, one click
  3. The first captured request appears. with its cost, attributed to its project

curl -fsSL https://get.haltonmeter.com/install.sh | sh

Windows · irm https://get.haltonmeter.com/install.ps1 | iex

or uvx halton-meter && halton-meter init

bg: Halton(2,3) — the sequence we're named for.