NEVDIS Registration
Intelligence Engine
How one monthly data purchase becomes the industry's only month-by-month record of Australia's heavy vehicle fleet.
Where this starts
Thanks for the call, Todd. As promised, this is ideas rather than War and Peace: what we heard, where we push back on the standard pitch, and the solution we build instead. Scope, timeline and pricing come after we have seen one real extract.
- The goal: a monthly HVIA report on new truck and trailer registrations, split by weight class, manufacturer and axle configuration.
- The source: the NEVDIS extract purchased through Austroads. Manufacturer, vehicle type, GVM, axles and more.
- The problem: manual entry at eight state registries makes the data dirty. Kenworth is a K in Victoria and a KW in Queensland.
- The catch: NEVDIS keeps no history. You can only ever buy the whole database as of right now.
- The field: other vendors propose a data lakehouse with hand-built logic layers and an LLM front end. You asked for a better way if we have one.
Because NEVDIS only ever sells now, whoever banks each monthly snapshot ends up owning something Austroads itself does not offer: a month-by-month history of the national heavy vehicle fleet. That history is not for sale at any price. Bank it from day one and, twelve months in, HVIA holds a dataset nobody else in the industry can reproduce. This is a moat, not just a report.
Playing devil's advocate
You asked for our honest view of the approach on the table. Here it is, applied to the other vendors' pitch and to our own first instincts.
A lakehouse is the wrong size
Lakehouses are enterprise plumbing designed for petabyte-scale streaming data from many sources. HVIA has one authoritative source arriving once a month. Build a lakehouse and you pay enterprise complexity, and enterprise invoices, forever, for a job a lean cloud warehouse does in weeks. The architecture should fit the data, not the fashion.
Counting rows is not the hard part. Identity is.
A new VIN in this month's snapshot is not automatically a new truck sold. It can be a re-registration, an interstate transfer or an import returning to the register. A report that cannot tell the difference inflates the market, and the first manufacturer whose own sales figures disagree says so publicly. The fix is VIN-level differencing plus classification, shown in Diagram 3.
An LLM front end on dirty data is confident nonsense
The other pitch leads with the chat interface. We put AI where it earns its keep: entity matching inside the cleansing engine. Everything else is deterministic automation, which is cheaper, faster and auditable. As we said on the call, not everything needs AI. A conversational layer is a genuinely good later add-on, grounded in a clean warehouse so it quotes real numbers.
- Month one has no baseline. The first change report lands in month two, so the first extract should be bought and banked as early as possible. Every month of delay is history lost for good.
- Licence terms decide the product. What Austroads permits HVIA to publish to members, and to media, needs confirming before the build, not after.
- Nobody knows how dirty the data really is. One real extract and one day of audit answers it, and that should happen before anyone quotes a fixed price.
The better way: a Registration Intelligence Engine
Four right-sized layers, fully automated end to end. Once live, the monthly report builds itself; a human touches it only to approve it before it goes to members.
Every component, including the raw archive, the alias dictionary and the warehouse, is built as an HVIA-owned asset: portable, documented and yours. Nothing is locked inside a vendor's platform.
Month-on-month truth, at VIN level
The report's credibility rests on one question: was this vehicle genuinely registered for the first time this month? We answer it per VIN, never per total. As a bonus, VINs that leave the register give HVIA fleet attrition and fleet age insights that a simple count can never see.
Cleansing that compounds
Deterministic rules handle every alias already learned. AI proposes matches for new variants with a confidence score. Only the genuinely ambiguous reaches a person, who confirms in about thirty seconds, and the engine remembers that answer forever.
What lands in inboxes each month
One branded report, drafted by the engine and approved by HVIA in minutes:
The report is the monthly deliverable. The asset is the history bank behind it: it grows every month and gets harder to compete with. Later, the same warehouse can power member benchmarking portals, anomaly alerts, forecasting, media-ready market stats and tiered data access for members and non-members. Each is a decision for another day, and none requires rework.
How this starts
We de-risk before we build. Everything below firms up after one day with one real extract.
One sample extract under NDA. We measure how dirty the data really is, confirm field coverage and confirm Austroads licence terms for publishing.
Ingestion, permanent archive, cleansing engine v1 and the first automated monthly report, end to end.
History warehouse, VIN-level differencing and classification, live dashboard and automated distribution.
Member benchmarking, anomaly alerts and conversational analytics, once the foundation has earned trust.
No pricing in this paper, deliberately. We would rather show you real numbers from your real data than guess in a document.
Next step
A 45-minute Google Meet to walk through these ideas and, ideally, look at a sample extract together. Connecting to databases, extracting and amalgamating data, cleansing it and producing automated reports is our bread and butter. We would love to build this with HVIA.
Ready when you are, Todd
One email locks in the Google Meet and we bring the whiteboard.
Lock in the Google Meet