Why AI Needs Reconciled Revenue Data: Lessons for Publishers

An earlier version of this article was originally published in AdExchanger.

Every publisher and media company is asking the same question right now: how do we make AI work for our business? The conversations are happening at every level, from product teams experimenting with generative tools to executives promising board members an “AI-first” strategy. But in the rush to adopt, there’s a foundational question most teams aren’t asking loudly enough: is your data actually ready for it?

AI doesn’t conjure insights from nothing. It works on what you give it. And if what you give it is a mess of inconsistent naming conventions, ungoverned exports, siloed systems, and unreconciled discrepancies, the AI will faithfully process all of that mess and hand you something that looks authoritative but is, at its core, unreliable.

This article outlines the critical aspects of publisher data readiness that media companies need to address before AI can deliver real value. These items will allow AI to support them in revenue reporting, client deliverables, and other operational decisions.

1. Data Hygiene: Garbage In = Garbage Out, Amplified!

The “garbage in, garbage out” principle is as old as computing. But AI makes it worse, not better, because the output looks clean even when it isn’t.

Traditional software fails loudly when the data is bad. A SQL query on a malformed table throws an error. A dashboard with missing fields leaves a blank cell. AI, on the other hand, will confidently fill those gaps with probabilistic inference. That means bad data doesn’t produce obvious errors anymore; it produces plausible-looking wrong answers. What AI can't replicate is the practitioner who has lived inside that data long enough to sense when something is off: when a figure is suspiciously tidy, or when a pattern contradicts known business reality. Not everyone in the organization will have that instinct, and most won't know when they're missing it. A well-governed data foundation is what keeps those users safer, limiting the AI's exposure to the inconsistencies and gaps that produce misleading answers.

2. Auditability and Explainability: Your Finance Team Will Ask

There is a version of AI adoption that creates two parallel worlds inside your organization: the world where the AI does interesting things with data, and the world where the finance team actually operates. If those two worlds don’t connect, and in most vibe-coded or hastily assembled AI implementations, they won’t, you’ve created extra work.

Finance, legal, and billing functions require auditability. They need to know exactly where a number came from, what logic produced it, and whether it will produce the same result if re-run tomorrow. An AI model, especially one that was assembled quickly without formal data governance, cannot guarantee that. When billing runs at month-end, your team will still need to normalize, validate, and reconcile data from scratch. The AI layer becomes a detour instead of a shortcut.

Some AI tools execute code in a temporary sandbox and discard both the code and the environment immediately after, leaving zero audit trail and creating a compliance nightmare for your CISO.

The bigger risk is fragmentation. Without a single source of truth, performance analytics data, billing data, performance analytics data, and client-facing reporting can all come from different pipelines with different logic. Every team is right in its own system, and no one agrees at the executive level. That is not a new problem created with AI, but an underlying data architecture problem that AI makes visible and urgent.

3. Temporal Consistency: The Same Question Should Get the Same Answer

Here is a simple test for your AI setup: ask it how much revenue your business generated last month. Then ask it again tomorrow, with no changes to the underlying data.

If you get different answers, you have a problem.

LLM-based systems, by design, introduce variance at the inference layer. Even when the data hasn’t changed, the model can produce different outputs on different runs. For internal productivity tasks, drafting emails or summarizing documents, that variance is tolerable. For revenue reporting, pacing, billing reconciliation, or client analytics, it is not.

One proposed solution is to use AI to generate deterministic code that then runs on top of the data. This can work in theory, but it introduces a different problem: you end up with business-critical logic that lives in generated code no one on your team fully understands, maintains, or can audit when something breaks. That is a serious operational and compliance risk, particularly in regulated or high-value client environments.

4. Compliance: AI Still Has to Pass Procurement

For anyone who has sat through a procurement or vendor security review at a large media company, this section will be familiar. Enterprise organizations have formal requirements for any software that touches business-critical data. Certifications and standards compliance, documented data handling processes, access controls, retention policies, and clear vendor accountability are table stakes.

A vibe-coded internal tool or ad hoc pipeline assembled with a few AI prompts and a CSV export will not pass that review. If the output of that pipeline is used to produce numbers that go into billing, reporting, or client deliverables, the compliance team’s answer will be “no”, and it will stay “no” until the foundation is rebuilt properly.

They’re not luddites. But compliance requirements exist for real reasons, and AI adoption in enterprise contexts has to happen within those constraints, not around them.

5. Normalization Across Sources: AI Can’t Reconcile What You Haven’t Aligned

Publisher advertising data doesn’t live in one place. It lives across ad servers, order management systems, SSPs, verification vendors, audience platforms, and bespoke internal tools, each with its own naming conventions, metrics definitions, time zones, and attribution logic.

AI cannot reconcile those differences on its own. It will map what it can, make assumptions about the rest, and produce output that looks unified but contains hidden discrepancies. An AI that smooths over a 10% variance between what your ad server reports and what your SSP invoices is making a business decision you likely did not intend to empower it to make. The normalization and integration layer has to be built deliberately, with human oversight, before AI-powered analytics can be trusted. That work isn’t glamorous; it’s taxonomy alignment, field mapping, discrepancy thresholds, and data contracts with your tech partners. But it’s the work that makes everything downstream reliable.

Getting Ready: A Practical Checklist

Before investing time, resources, or money in AI tooling for your data and analytics workflows, make sure the foundation is in place:

  • Normalize and aggregate advertising data from all sources into a single, governed pipeline before routing it to any AI system.

  • Ensure your data is stored in a format and location that AI systems can access reliably, whether through an MCP server, a structured database, or a well-documented API layer.

  • Establish data quality standards and implement automated checks upstream, so the AI never sees bad data in the first place.

  • Work with your BI, finance, and legal teams to define what “audit-ready” means for your organization and build to that standard from day one.

  • Document your data logic explicitly, not just for compliance, but so that any AI-generated analysis can be traced back to a verifiable methodology.

The promise of AI in advertising and media is real. AI is incredibly valuable, and we should collectively embrace it, but the companies that will realize it aren’t the ones who moved fastest; they’re the ones who built on solid data foundations. AI amplifies what’s underneath it. Make sure what’s underneath it is worth amplifying.


See how CBC & Radio-Canada Media Solutions modernized its data foundation with Burt.

Previous
Previous

Open Auction Benchmarks: June 2026

Next
Next

Open Auction Benchmarks: May 2026