Your own personal detective — backed by science and forensic scrutiny.
Fuel Bench learns what normal looks like for every vehicle, watches millions of transactions for the moments that break the pattern, and then explains itself — citing the exact record behind every finding.
Most tools hand you a red flag and a shrug. Fuel Bench hands you a case file — assembled by a detective that never sleeps, never guesses, and never leaves a claim unsourced.
A national fleet spends heavily on fuel every year — and some of that spend is theft. Each instance hides inside a transaction that looks entirely ordinary: a fill that's a little large, a card used a province away, an “oil purchase” that isn't oil.
Static thresholds miss the patient thief and drown you in false alarms.
Black-box “AI” accuses an innocent driver — and cannot show you why.
The second failure is the costly one. A wrong accusation can end a career — and start a lawsuit. Fuel Bench is built around that constraint.
“Refusal is a first-class verdict — not a failure.”
Fuel Bench would rather tell you the evidence is too thin than hand you a confident accusation it cannot defend. Three mechanisms enforce that discipline in code — not in a policy document.
The highest verdict is disabled by default and fails closed.
The model cannot invent a number.
Every answer is written to an append-only, owner-proof record.
The difference between an AI that is impressive and one that is admissible.
Each layer catches what the others can't — and none of them gets the final word alone. Select a layer to see how it works.
A tank that holds 70 litres did not just take 94. A card swiped in Gauteng at 14:02 cannot be swiped in the Free State at 14:31 — Fuel Bench computes the distance between the two stations from their coordinates and the time between swipes, and if the implied speed is impossible, that isn't a probability. It's arithmetic.
Rules span fuel, oil and toll — including the oil and toll transactions most platforms never inspect — each carrying its own risk score and a severity band from Watch to Critical, with a recommended action attached.
The investigator is a self-hosted, quantised open model running on CPU only — no GPU, no cloud inference, no third-party AI API. Every token of reasoning about your fleet, your drivers and your losses is generated on hardware you control.
Natural-language questions are compiled to SQL by the same local model, then passed through an AST-level safety validator: schema allowlist, mandatory tenant scoping, read-only transactions, hard timeouts, and PII redaction on the way out.
For a government fleet, that isn't a feature. It's the precondition.
Fuel Bench's intelligence ships in phases. Each phase unlocks only when the previous one proves itself.
Deterministic rules across fuel, oil and toll. Per-vehicle learned baselines with same-engine peer cohorts. Anomaly scoring. A self-hosted investigator with tool-calling, refusal, confidence caps, hallucination scrubbing and an immutable audit trail — with a calibration harness enforcing the confidence gate.
Supervised models trained on adjudicated outcomes as the feedback loop matures. A second model that argues against every finding before a human sees it. Financial-impact and investigator-console dashboards.
Graph learning over the relationship network — earned only if the classical signal proves out. Hybrid keyword-plus-vector retrieval. Cryptographically signed answer receipts.
// We ship the guardrails before the confidence. That order is not an accident.
Your own personal detective — backed by science and forensic scrutiny.
Fuel Bench is in early access for government and enterprise fleets running fuel cards at national scale.
We'll only use this to talk about access.