EARLY ACCESS Fuel-fraud intelligence · large fleets

Fleet fuel fraud,
proven — not guessed.

Your own personal detective — backed by science and forensic scrutiny.

Fuel Bench learns what normal looks like for every vehicle, watches millions of transactions for the moments that break the pattern, and then explains itself — citing the exact record behind every finding.

Most tools hand you a red flag and a shrug. Fuel Bench hands you a case file — assembled by a detective that never sleeps, never guesses, and never leaves a claim unsourced.

Request early access See how it works
Self-hosted · CPU-only · your data never leaves your infrastructure
CASE FILE #FB-2481·GP
ANALYSING
01 VEHICLE ADG 4A2·1 — 2.0 TDI row 8842
02 ODOMETER +410 km/day ok
03 FILL 94.0 L / tank cap 70 L flag
04 CARD GP 14:02 → FS 14:31 watch
05 PRICE R 22.14 / L ok
VERDICT INCONCLUSIVE
CONFIDENCE 40 · capped
CONFIRMED — locked until precision is proven
The problem

A loss that hides in plain sight.

A national fleet spends heavily on fuel every year — and some of that spend is theft. Each instance hides inside a transaction that looks entirely ordinary: a fill that's a little large, a card used a province away, an “oil purchase” that isn't oil.

TOO QUIET

Static thresholds miss the patient thief and drown you in false alarms.

TOO LOUD

Black-box “AI” accuses an innocent driver — and cannot show you why.

The second failure is the costly one. A wrong accusation can end a career — and start a lawsuit. Fuel Bench is built around that constraint.

FLAGGED
one anomaly · hidden among the ordinary
Our mandate · precision over recall

Admissible, not just impressive.

“Refusal is a first-class verdict — not a failure.”

Fuel Bench would rather tell you the evidence is too thin than hand you a confident accusation it cannot defend. Three mechanisms enforce that discipline in code — not in a policy document.

/ 03

It earns its certainty.

The highest verdict is disabled by default and fails closed.

CONFIRMED unlocks per tenant only after a calibration harness replays hundreds of already-adjudicated cases and proves at least 95% precision. No passing report on file, and the verdict is quietly downgraded — on every single request.

HOVER TO READ HOW →
/ 03

Every figure is real.

The model cannot invent a number.

A scrubber checks every value in the AI's prose against the verbatim cells of the records it retrieved. One unrecognised figure forces the verdict down to inconclusive. No rounding, no estimating, no embellishment.

HOVER TO READ HOW →
/ 03

Nothing is quietly rewritten.

Every answer is written to an append-only, owner-proof record.

A seven-year, trigger-enforced audit store rejects edits and deletions outright — even the database owner cannot alter a past answer. Only feedback columns can be written.

HOVER TO READ HOW →

The difference between an AI that is impressive and one that is admissible.

How it works

Three layers that check each other.

Each layer catches what the others can't — and none of them gets the final word alone. Select a layer to see how it works.

LAYER 01
Physics & law
Deterministic rules that don't need a model to be right.
LAYER 02
Learned baselines
Every vehicle gets its own definition of normal.
LAYER 03
Interpretation
A self-hosted detective reads the evidence and writes the case file.
LAYER 01 · PHYSICS & LAW

A tank that holds 70 litres did not just take 94. A card swiped in Gauteng at 14:02 cannot be swiped in the Free State at 14:31 — Fuel Bench computes the distance between the two stations from their coordinates and the time between swipes, and if the implied speed is impossible, that isn't a probability. It's arithmetic.

TANK 70 L FILL 94 L
GP 14:02 → FS 14:31 · Δ 29 min

Rules span fuel, oil and toll — including the oil and toll transactions most platforms never inspect — each carrying its own risk score and a severity band from Watch to Critical, with a recommended action attached.

LAYER 02 · LEARNED BASELINES

A blanket threshold is a blunt instrument. A loaded 4×4 on a mountain pass and a sedan on the highway don't share a fuel curve. Fuel Bench resolves each vehicle's expected consumption through a cascade that always prefers the tightest trustworthy evidence — the vehicle's own history, then a peer cohort, then wider tiers — and only trusts a cohort once it clears a sample-size gate.

OUTLIER
expected ±σ
this vehicle's fills

Peers aren't matched on the badge — they're matched on the code that means the same engine, gearbox and trim. Baselines then adjust for altitude and vehicle age, and every one carries its own sigma — so the system knows the difference between a vehicle it understands and one it's still learning.

LAYER 03 · INTERPRETATION

Every alert gets a detective assigned to it — one that works the case with forensic scrutiny, not a hunch: it gathers evidence, weighs it against the science of learned baselines and statistical distributions, and argues the case against itself. When an alert fires it doesn't summarise — it investigates, running an agentic loop across purpose-built tools: pulling fill history, scoring the vehicle against its true peer cohort, tracing a card's station fingerprint, checking whether this vehicle has fired alerts like this before.

→ fill 94 L vs tank 70 L [row 8842·fill_l]
→ +34% vs peer cohort [ml.cohort_stats]
COUNTER-HYPOTHESIS REQUIRED — meter fault · twin tank · data error

It writes a narrative with ranked hypotheses — and the schema refuses to validate unless it has also argued a credible innocent explanation. Every claim points at a specific row and column. An accusation is never the system's to make: Fuel Bench builds the case; a person decides.

Data sovereignty

Your data never leaves your infrastructure.

The investigator is a self-hosted, quantised open model running on CPU only — no GPU, no cloud inference, no third-party AI API. Every token of reasoning about your fleet, your drivers and your losses is generated on hardware you control.

Natural-language questions are compiled to SQL by the same local model, then passed through an AST-level safety validator: schema allowlist, mandatory tenant scoping, read-only transactions, hard timeouts, and PII redaction on the way out.

For a government fleet, that isn't a feature. It's the precondition.

CPU-only inference
GPU for AI interpretations.
Self-hosted model
Runs entirely inside your estate.
No third-party AI API
Zero external inference calls.
Proven at scale

Built on real fleet history.

LIVE FIGURES · AS AT 2026-07-10
25M+
Fuel transactions analysed
176K+
Vehicles catalogued
74
Detection rules · 12 families
5 years of history fuel · oil · toll thousands of stations & merchants
How we're building it

Guardrails first. Confidence second.

Fuel Bench's intelligence ships in phases. Each phase unlocks only when the previous one proves itself.

NOW

Deterministic rules across fuel, oil and toll. Per-vehicle learned baselines with same-engine peer cohorts. Anomaly scoring. A self-hosted investigator with tool-calling, refusal, confidence caps, hallucination scrubbing and an immutable audit trail — with a calibration harness enforcing the confidence gate.

NEXT

Supervised models trained on adjudicated outcomes as the feedback loop matures. A second model that argues against every finding before a human sees it. Financial-impact and investigator-console dashboards.

LATER

Graph learning over the relationship network — earned only if the classical signal proves out. Hybrid keyword-plus-vector retrieval. Cryptographically signed answer receipts.

// We ship the guardrails before the confidence. That order is not an accident.

Request early access

Find the criminals who steal fuel.
Never wrongly accuse the innocent.

Your own personal detective — backed by science and forensic scrutiny.

Fuel Bench is in early access for government and enterprise fleets running fuel cards at national scale.

Please enter a valid work email address.
Thanks — we'll be in touch. Your details go no further than our inbox.

We'll only use this to talk about access.