---
title: "Everyone agreed AI needs accuracy. Nobody showed the scorecard."
date: "2026-09-19"
summary: "Fierce Pharma Week named the right problem: accuracy in AI answers. Naming isn't measuring — here are four questions that actually score it for a brand team."
tags: [aeo, geo, measurement]
---

I once sat inside an account reporting a 76% paid search conversion rate. Nobody had questioned it. Month after month it went into the deck, and month after month it got approved, because a number that good does not invite scrutiny.

The inflation was ordinary. Imported micro-events counted as conversions. Locator taps double-counted. Smart Bidding pointed at the resulting noise, which meant the algorithm spent real money optimizing toward a fiction.

Data that flatters does not get audited. Data that disappoints does. That asymmetry is the most expensive thing in digital marketing, and it has just moved to a new layer.

MM+M's preview of Fierce Pharma Week said the industry was done with AI as spectacle. The agenda would be accuracy, measurement, and whether brands are completely represented in the answers large language models give patients and clinicians. That is the right conversation. It is also the one the industry has not operationalized.

What shipped around the same week was more machinery. DeepIntent opened Cora, an agentic planning and activation layer where brands and agencies share a workspace. The Trade Desk pushed Crossix and IQVIA signals deeper into the buy path, with Gross NBRx optimization coming later this year. Doceree launched an Intent Powered HCP Suite promising clinical intent across CTV, LinkedIn, and programmatic. Assembled Intelligence named Andrea Palmer, formerly of Publicis Health Media, as CEO and waved GEO and answer-engine omnichannel as part of the story. Genentech's CMO said new marketer measurement tools arrive in December.

None of it is unserious. All of it is the same pattern: more tools that report on themselves, and no independent scorecard for the layer that now sits upstream of the website.

## The agency gap

Every holding company now has an AI practice, a chief AI officer, and a deck that opens with "AI-first." Almost none of them have an AI line item in a scope of work.

Ask your AOR a direct question: who on this team owns whether our brand is cited accurately in AI answers? In most rooms the answer is a pause, then a name from the SEO pod, then a follow-up meeting. The people who wrote the AI strategy deck are frequently not the people who could tell you whether the brand site is crawlable by the engines doing the citing.

That is not incompetence. It is incentive. Agencies are compensated for production, and the answer layer has nothing to produce. No asset, no MLR job number, no trafficking sheet. A channel that generates no deliverable does not get staffed, and what does not get staffed does not get measured.

## The compliance problem nobody has a code for

Presence without on-label accuracy is not an AEO win. It is a compliance incident with distribution.

When a caregiver triggers an AI Overview on an unbranded disease-state query and the engine lifts your efficacy language while dropping the boxed warning, that extract is in front of a patient. It carries your brand. It did not go through MLR, because there is no review code for it. Your team reviewed the page. The engine published the paraphrase.

Every other channel has an owner and a workflow. The answer layer has neither. It is the only surface where brand language reaches patients and physicians with nobody accountable for what survives summarization.

It is also not one surface. Consumer ChatGPT, clinician-facing deployments, and enterprise healthcare instances return different answers to the same question. A brand can look fine in an AI Overview and be invisible to a physician asking a clinical question. Treating GEO as a single line item is how you get a false green on a dashboard.

## Four questions worth an afternoon

**On our priority unbranded queries, who gets cited today in Google AI Overviews?**

Pull your top 20 unbranded disease-state queries and run each one logged out, capturing the cited domains in the AI Overview. It is common for that citation set to be dominated by third-party health publishers, advocacy organizations, and payer sites, with the manufacturer's own site absent entirely.

**Who gets cited in consumer ChatGPT versus clinician-facing deployments?**

Ask the same clinical question in consumer ChatGPT and in a clinician-facing or enterprise healthcare deployment, then compare the sources each one names. Divergence is the finding: a brand present in one and absent in the other has a surface-specific gap rather than a general GEO problem, and the remediation differs accordingly.

**Where we appear in AI answers, is the extract on-label?**

Read the generated answer the way MLR would read a promotional piece, checking whether the efficacy claim arrived with its associated risk language and whether the indication is stated accurately. An extract that carries the benefit without the boxed warning is a labeling exposure regardless of whether your own page was compliant.

**Are our paid programs optimizing toward real high-value actions?**

Audit the conversion actions marked primary in GA4 and imported into the bidding platform, and confirm each one maps to a behavior with real commercial meaning such as a treatment center search, a rep request, or a support program enrollment. Scroll depth, video starts, and duplicate locator taps inflate the conversion rate and teach Smart Bidding to buy the wrong traffic.

## The independent layer

Same diagnostic shape as any measurement audit. The expensive problems rarely look like problems on a stage.

This is not an argument for firing your agency. It is an argument for one scorecard your agency does not own: whether the brand is present in the answers patients and clinicians actually read, whether that presence is accurate, and whether the dollars chasing those audiences point at outcomes that mean something.

Presence in AI answers is being allocated right now, query by query, while most brand teams still report rankings and impressions. The teams that treat the answer layer as a channel this year, scored rather than sloganeered, will be very hard to displace from it later.

I was not in the room. The scorecard matters more than the room.
