← All posts
Measurement

Citation share is not accuracy. June gave us both numbers side by side.

A PR firm says AI engines cite Eli Lilly in 12.5% of pharma answers. A peer-reviewed study says those answers are fully accurate with reliable references 19% of the time. Only one of these numbers should change what you do.

VizLoop Measurement notes · 2026-09-02

Two documents about AI answers and prescription drugs circulated in June 2026, and the distance between them is the clearest picture available of where this field actually is.

The first is the "Pharma / Rx AI Visibility Index," published June 19 by 5W AI Communications and picked up by Fierce Pharma, Yahoo Finance, and Morningstar. The second is a study in the June 15 issue of the American Journal of Health-System Pharmacy by researchers at the University of Michigan. The index measures which drugmakers AI engines mention. The study measures whether what the engines say is true and complete. If your company is discussing one of these documents and not the other, it is almost certainly discussing the wrong one.

What the index says, and what it is

5W's index ranks the top 25 drugmakers by estimated "AI citation share" across ChatGPT, Claude, Perplexity, Gemini, and Google AI Overviews. The method, per the firm's own release: more than 60 tracked patient and consumer prompts spanning metabolic, oncology, immunology, vaccine, and specialty categories, each run five times per engine in clean sessions. The headline results: Eli Lilly at 12.5% and Novo Nordisk at 11.5%, with Pfizer at 8.5%, J&J at 7.0%, and Merck at 6.0%. The framing device is spend: US pharma put an estimated $8 billion into direct-to-consumer advertising in 2024, and 5W's founder is quoted observing that the engines name the companies whose drugs patients research, not the companies running the most prime-time spots.

Now the part that has to be said before anyone puts those percentages in a slide. This is a self-published ranking from a PR firm that sells AI-visibility services. The prompt set is self-selected, the share figures are estimates over that self-selected set, and the methodology is not independently reproducible. Fierce Pharma repeating the numbers does not upgrade them; syndication is distribution, not verification. A second such ranking, the Worldcom "Healthcare Monitor Pharmaceuticals 2026," is circulating with the same character and warrants the same treatment.

None of this makes the index dishonest. It makes it marketing, and it should be read the way any vendor benchmark is read: as an argument for the vendor's category, illustrated with the vendor's own data.

What the study says, and what it is

The AJHP paper is a different kind of document. Frazer, Yoon, and Durant compared ChatGPT and Gemini against drug-information pharmacists at University of Michigan Health on 300 real queries received by the drug-information service. The results, quoted directly: both assistants were completely accurate with reliable references provided 19% of the time. Fifteen percent of ChatGPT's answers were completely inaccurate and included fake or unreliable references; 5% of Gemini's were. The remainder, 66% for ChatGPT and 76% for Gemini, were partially correct and/or incompletely sourced.

Two caveats, both of which we would rather state than have a reader discover. The paper published online in December 2025 and carries a June 2026 print date, so its appearance in June coverage is an artifact of publishing calendars. And the queries and model versions date to July 2022 through June 2023, so the specific percentages describe models that have since been replaced.

What survives the caveats is the distribution. Complete inaccuracy was the rare failure. The dominant outcome, roughly two-thirds to three-quarters of answers, was the partial answer: correct as far as it went, incomplete in what it included or what it cited. Fabricated references appeared in a measurable share. Newer models have shifted these numbers, but no published evidence shows the distribution inverting, and anyone claiming it has should be asked for their data with the same skepticism applied above.

The two numbers measure different worlds

Put the documents together and the mismatch is the finding.

Citation share asks: when a patient asks an AI engine about a condition or a treatment, does your company's name appear? It is a marketing metric, and by that logic the June index is good news for Lilly and Novo Nordisk and a budget argument for everyone else.

The AJHP distribution asks: when the answer names your drug, what did it leave out? An answer can cite your product in every session the index tracks and still omit the boxed warning, the contraindication, or the interaction that the label requires and a pharmacist would have included. On the study's numbers, that incomplete answer is not the edge case. It is the most common thing the engines produce.

Which means citation share is not merely a different metric. For a regulated product, high citation share with unmonitored answers is exposure, not achievement. Every additional answer that names your drug is another answer that may be carrying a partial version of your label to a patient, with no promotional review anywhere in the loop, and the index's own premise, that these mentions are earned by patient research volume rather than by anything the company approved, is precisely the problem. The mentions happen without you.

What to do with each document

The index will surface internally; it is built to. When it does, the right response is not to debate the percentages, which cannot be verified, but to redirect the question they raise. Whether the company is mentioned is not actionable for a regulated product. What the mentions say against the label is.

The study earns a different use: it is the current best peer-reviewed estimate of the answer-quality baseline, suitable for citing to compliance and medical affairs with its caveats attached, and a standing reason why "we rank well" and "we are fine" are different sentences.

The practical program that follows is not complicated. Take the questions patients and clinicians actually ask about your products. Capture what the engines answer, verbatim and dated. Score each answer against the current label, weighting the elements the label itself treats as most serious. Route what is missing to someone whose name is on it, and re-measure on a schedule, because retrieval changes monthly. That is monitoring, not visibility. June 2026 supplied the clearest one-month argument yet for knowing the difference: one document celebrating that the engines talk about pharma, and one documenting, with a method anyone can check, how rarely they say everything the label requires.

This is one of three deep dives on the month. How the engines' own retrieval changes move these answers with no review in the loop is in The answers changed in June and nobody reviewed them; the full June record is in the June 2026 roundup.

AI visibility measurement evidence pharma
Read the latest posts

Email subscriptions are paused. All posts remain available on the blog.

Related posts