Chapter 04

Field evidence: what multi-model audits actually reveal

Between 2025 and 2026, GEOMed360 conducted a structured audit programme across a portfolio of assets spanning oncology, cardiometabolic disease, and interstitial lung disease — running standardized HCP- and patient-persona prompt sets across five general-purpose large language models. The findings below are reported at the pattern level; they replicated across assets and therapeutic areas.

Finding 1 — The citation mix is dominated by sources the manufacturer does not control

Across audits, third-party clinical aggregators and reference platforms consistently captured the largest citation share, followed by regulatory repositories and open-access journals. Manufacturer-owned web properties captured under 10 percent of citations for the launch-phase oncology asset — despite being the most accurate and current source available — and gated HCP portal content captured exactly zero, across every model and every prompt.19 The manufacturer's most controlled channel is its least visible one.

Exhibit 4
Who actually gets cited when AI answers questions about a medicine
Citation share by source type showing third-party aggregators dominating and manufacturer sites under 10 percent
Source: GEOMed360 multi-model audit of a launch-phase targeted oncology therapy, 2025–2026; five general-purpose LLMs, standardized HCP- and patient-persona prompt sets; n = all logged citations. Percentages rounded; asset and sources anonymized.

Finding 2 — Format determines visibility more than authority does

The audits repeatedly surfaced a paradox: the most scientifically authoritative content generated the fewest citations when its format was machine-hostile. Conference posters carrying the most current subgroup data for an investigational bispecific agent produced effectively zero citations because they exist as images. Paywalled pivotal publications were displaced by derivative news coverage. Regional health-technology-assessment reports — scanned PDFs — were invisible even to models specifically prompted about access questions. Meanwhile structurally clean third-party pages with older, thinner data were cited constantly.19 Authority that machines cannot parse is authority that does not exist.

Exhibit 5
Machine-readability, not scientific weight, gates the citation
Index of citation frequency by content format, structured HTML highest, image-based and gated formats near zero
Source: GEOMed360 analysis across portfolio audits, 2025–2026. Index synthesizes observed citation frequency by source format, normalized to structured HTML = 88; illustrative of consistent rank-ordering rather than precise measurement.

Finding 3 — The errors are systematic, not random

Four recurring, correctable error patterns accounted for the large majority of clinically meaningful inaccuracies:

  • Version confusion. Models cited superseded label versions after supplements had been approved — sometimes months after — because the older version remained more prominently indexed and nothing signalled succession. The clinical consequence: outdated dose-modification guidance presented as current.19
  • Subgroup conflation. Pooled response rates were attributed to specific subpopulations (and vice versa) whenever pooled and subgroup data were not structurally separated with explicit population definitions and n values in the source content.19
  • Silence misread as risk — or as safety. Where content did not explicitly state that no dose adjustment is required for a given population, or that a studied interaction is not clinically significant, models inferred answers from drug-class patterns — in both directions. Explicit negative statements ('not required', 'not clinically significant') proved to be among the highest-value content additions per word.19
  • Audience blending. Patient-facing sources appeared in HCP-persona answers and clinical sources in patient-persona answers whenever audience was not machine-legible in metadata and body text — degrading precision for the clinician and comprehension for the patient.19

Finding 4 — The two personas fail differently, and patient answers fail more

Scored across five accuracy dimensions, HCP-persona answers consistently outperformed patient-persona answers on the same underlying clinical questions, with the widest gaps in safety framing and special-population guidance — exactly the dimensions with the highest potential for patient harm. The gap is structural: the clinical evidence base is written for professionals, and the plain-language layer that would let machines answer patients accurately is the layer pharmaceutical content estates most often lack.19

Exhibit 6
The dual-persona gap: patient-facing answers trail on the dimensions that matter most
Accuracy comparison between HCP-persona and patient-persona AI answers across five dimensions
Source: GEOMed360 multi-model audit programme, 2025–2026; mean accuracy scores (5-point scale) across five general-purpose LLMs and standardized prompt sets; anonymized and aggregated across assets.

Finding 5 — Model heterogeneity makes single-platform strategies untenable

The five audited models differed materially in source preference (some favouring regulatory repositories, others news and reference aggregators), citation transparency, and error profile. An asset could score well on one platform and poorly on another for identical questions. Optimizing for a single engine — the instinct inherited from a Google-centric decade — leaves the majority of AI-mediated decision moments unmanaged.19 A related technical finding: broken, redirected, or truncated URLs in otherwise accurate content zeroed out its citation value; URL integrity is a prerequisite, not a refinement.19

References cited in this chapter

Numbering follows the full GEOMed360 whitepaper, Winning the Answer.

  1. 19.GEOMed360 analysis: multi-model, dual-persona audit programme across a pharmaceutical portfolio spanning oncology, cardiometabolic disease, and interstitial lung disease, 2025–2026 (see the methodology note in Measuring what matters).