CONTROVERTIST

Examination 018

Can AI improve clinical attrition, decision timing, and trial execution enough to reduce inflation-adjusted expected cost per approved therapeutic by 30% against a defined no-AI 2035 counterfactual, r

The thesis

AI will reduce the cost of discovering and bringing a new drug to market by at least 30% by 2035.

Examined: August 2026

Evidence current through: August 2026

01

Independent examination

An examination of the thesis, not a recommendation.

Thesis under examination

Can AI improve clinical attrition, decision timing, and trial execution enough to reduce inflation-adjusted expected cost per approved therapeutic by 30% against a defined no-AI 2035 counterfactual, rather than merely making discovery tasks cheaper?

Current read

Verdict: too early to test as an industry-wide numerical forecast, and currently underspecified rather than supported or falsified. The investigation shifts the thesis from AI productivity in discovery to whether AI changes the costly distribution of failures, especially late clinical failures, under a consistent cost-per-approval denominator. The strongest non-obvious finding is that cheaper prediction and molecule generation could increase total experimentation without lowering cost per marketed drug; portfolio behavior is therefore part of the causal mechanism, not an accounting footnote. Public evidence confirms that AI-associated candidates can reach clinical trials, but trial entry does not establish better transition probabilities, approvals, or lower end-to-end cost. The conclusion would change with mature, matched portfolio evidence showing AI-attributable improvements in phase transitions, earlier termination of doomed programs, development time, and fully loaded cost under a common accounting method.

Decisive unknown

The decisive unknown is the AI-attributable change in expected clinical loss per eventual approval: whether comparable AI-assisted programs fail less often, fail earlier, or progress faster than the no-AI counterfactual. That cannot presently be resolved without matched program histories and common cost accounting.

Strongest counterargument

Drug development may be bottlenecked not by search over molecules but by biological uncertainty, heterogeneous human response, long safety observation periods, manufacturing validation, and evidentiary requirements. If AI mainly lowers upstream search costs while expanding the number of candidates entering expensive trials, it could leave expected cost per approval unchanged or even raise aggregate R&D spending.

What would change our view

Matched clinical phase-transition and approval probabilities — Persistent, risk-adjusted improvement across complete AI-assisted cohorts would directly lower expected cost per approval; no improvement would confine much of the value to upstream efficiency.

02

Evidence

The research foundation, before any interpretation. Inference is never presented as fact.

  • Established

    Drug development comprises target identification, hit and lead work, preclinical testing, phased clinical trials, regulatory review, and manufacturing preparation; AI has different potential leverage at each stage.

    Verified in FDA, NIH, and standard drug-development descriptions cited by the grounded brief.

  • Established

    Expected cost per approved drug depends materially on failed programs and elapsed development time, not only on spending for the successful candidate.

    Verified in pharmaceutical R&D cost methodology; estimates differ according to attrition allocation and capitalization.

  • Established

    Reducing the cost of an experiment or discovery task does not by itself demonstrate a reduction in probability-adjusted cost per approval.

    This follows directly from the distinction between task cost, portfolio attrition, development duration, and the denominator of approved products.

  • Established

    Clinical development and late-stage failures are major components of expected end-to-end cost, although their exact shares vary by modality and accounting method.

    Verified by the published phase-duration, attrition, and cost literature summarized in the grounded brief; the relevant peer-reviewed decompositions remain a retrieval coverage gap in this run.

  • Claimed

    AI-focused biotechnology companies and pharmaceutical partners state that machine learning improves target selection, molecule generation, structure prediction, screening, trial design, and patient selection.

    Company-stated productivity measures are not uniform and generally lack controlled no-AI counterfactuals.

  • Established

    Candidates publicly attributed by sponsors to AI-enabled discovery have entered registered human trials.

    Trial registration is independently checkable, but AI attribution usually comes from the sponsor and trial entry does not establish approval, superior clinical success, or lower total cost.

  • Unknown

    Whether AI-associated candidates have higher clinical phase-transition or approval probabilities than comparable conventional candidates remains unresolved.

    The available cohorts are young, definitions of AI association differ, and sponsor selection, survivorship, indication mix, and calendar-time effects complicate comparisons.

  • Unknown

    No consistently structured public dataset has yet been established in this investigation that joins program-level AI use, spending, elapsed time, failures, termination reasons, and approvals.

    This is partly a substantive data problem and partly a retrieval limitation; relevant trial registries, peer-reviewed comparisons, and sponsor filings still require systematic checking.

  • Unknown

    The proposed 30% reduction has no determinate denominator because it does not specify out-of-pocket cost, capitalized cost, cost per candidate, cost per approval, or industry expenditure.

    The claim also lacks a baseline year, geography, modality scope, inflation convention, and explicit no-AI counterfactual.

  • Unknown

    It is not known whether productivity gains would be retained as lower cost per approval or reinvested in larger portfolios, harder targets, biomarker programs, and richer clinical evidence.

    Observed R&D expenditure could rise even if specific activities become cheaper, so expenditure alone cannot identify AI productivity.

  • Inferred

    A 30% end-to-end reduction is arithmetically possible only if AI substantially affects a large cost share or produces unusually large gains in a smaller share, including through attrition and time.

    This is a decomposition constraint, not empirical evidence that either condition will occur.

  • Unknown

    Current public evidence has not yet been updated here to establish whether any fully approved drugs have independently auditable, end-to-end AI-attributable cost histories.

    The grounded finding was stale as of approximately mid-2024 and must be refreshed using current FDA approval records, trial registries, peer-reviewed studies, and sponsor disclosures. The retrieved FDA materials supplied in this run do not discriminate the AI-cost thesis and therefore do not resolve this point.

03

Thesis stress test

The strongest available case on each side, argued at full strength.

What supports the thesis

  • Interpretation

    AI could lower expected loss by identifying weak targets or molecules before expensive clinical commitments.

    Better prospective prediction could shift failures into discovery or preclinical stages, where they are cheaper. The weakest link is evidence that benchmark accuracy transfers to prospective human outcomes and changes actual termination decisions.

  • Interpretation

    AI could improve candidate quality and thereby raise clinical phase-transition probabilities.

    Joint optimization of potency, selectivity, toxicity, pharmacokinetics, and developability could produce better entrants to the clinic. Public trial entry supports feasibility, but mature comparative success and approval cohorts are not yet established.

  • Interpretation

    AI could shorten development and reduce capitalized cost even when direct spending changes little.

    Faster design cycles, enrollment forecasting, protocol optimization, document preparation, and earlier decisions could reduce elapsed time. The claim requires program-level dates and a specified cost-of-capital method to distinguish real acceleration from indication or sponsor differences.

  • Interpretation

    AI could reduce operational clinical costs through better site selection, patient identification, and data-quality monitoring.

    These mechanisms reach beyond molecule discovery and therefore address a larger portion of the cost base. Their weakest link is attribution: ordinary analytics, decentralized-trial tools, biomarker advances, and organizational redesign may produce similar effects.

What challenges the thesis

  • Contradiction

    Human biological uncertainty may remain irreducible until sufficiently powered clinical evidence is collected.

    Models trained on incomplete, biased, or non-comparable biological data may improve ranking without materially reducing unexpected safety failures or failures of efficacy in humans.

  • Contradiction

    Clinical and regulatory requirements impose time and cost floors that faster molecular design cannot remove.

    Enrollment, treatment, follow-up, endpoint maturation, manufacturing validation, and regulatory review often depend on disease biology and evidentiary standards rather than computational speed.

  • Contradiction

    AI may expand the option set faster than it improves selection discipline.

    Generating more plausible candidates can increase downstream experiments and trials unless governance becomes better at stopping weak programs; lower unit costs can therefore trigger a rebound in activity.

  • Contradiction

    Industry-level averaging may conceal incompatible cost structures.

    Small molecules, biologics, vaccines, and cell or gene therapies differ in discovery leverage, manufacturing burdens, safety risks, and trial design. A single 30% estimate may be an artefact of aggregation.

  • Contradiction

    Selection bias can make early AI portfolios appear unusually productive.

    Sponsors may label only favored programs as AI-associated, choose tractable targets, partner out stronger assets, or report milestone successes more visibly than discontinuations. A credible comparison requires preregistered definitions and complete cohorts.

04

Interdisciplinary examination

What each discipline sees that the original framing of the question does not.

Health economics x survival analysis

The relevant object is a censored portfolio of programs moving through states, not an average successful drug. Multi-state survival models can separate phase-transition probabilities, time in state, competing termination risks, and approval while avoiding the mistake of treating immature pipelines as completed successes.

Mechanisms it reveals

  • Expected cost must include spending on failed programs allocated across eventual approvals.
  • Right-censoring is severe because many AI-associated clinical programs have not had time to succeed or fail.
  • Capitalized cost depends on when spending occurs, not merely how much is spent.
  • Calendar time, therapeutic area, sponsor size, and modality are confounders in AI-versus-conventional comparisons.

Questions this lens makes unavoidable

  • What is the preregistered time origin: target nomination, project initiation, candidate nomination, IND filing, or first patient dosed?
  • Do AI-associated programs have different cause-specific hazards of technical, safety, commercial, and strategic termination?
  • How sensitive is the 30% result to the discount rate and treatment of shared platform costs?
Decision theory x experimental science

The economic value of AI lies less in prediction accuracy than in changing consequential go, redesign, and stop decisions. A model can perform well on retrospective benchmarks yet destroy value if errors are concentrated near high-cost commitment thresholds or if teams ignore its recommendations.

Mechanisms it reveals

  • Calibration near decision thresholds matters more than aggregate benchmark accuracy.
  • False confidence can delay termination and make an apparently better predictor economically harmful.
  • Prospective shadow-mode studies can compare model recommendations with actual decisions before allowing the model to control capital allocation.
  • Value of information depends on whether an AI output changes an experiment or portfolio decision.

Questions this lens makes unavoidable

  • Which specific decisions does the system alter, and how often does management follow it?
  • What are the costs of false continuation and false termination at each stage?
  • Are claimed gains measured prospectively against locked baselines or reconstructed after outcomes are known?
Operations research x portfolio governance

Local acceleration can congest the global system. If generative tools increase candidate throughput while animal facilities, CMC capacity, trial sites, specialist review, or management attention remain fixed, value migrates to the bottleneck rather than appearing as proportional end-to-end savings.

Mechanisms it reveals

  • Cheaper candidate generation can create queues at validation, toxicology, manufacturing, and clinical operations.
  • The binding constraint may shift over time as one stage becomes more productive.
  • Portfolio optimization requires comparing marginal expected value across programs, not maximizing the number of candidates.
  • Termination governance determines whether expanded search becomes learning or waste.

Questions this lens makes unavoidable

  • Which resource becomes capacity-constrained after discovery throughput rises?
  • Does AI reduce queue time at the bottleneck or merely feed it faster?
  • Are portfolio committees terminating a larger fraction of weak programs before expensive commitments?
Causal inference x technology attribution

AI adoption is bundled with better assays, richer datasets, biomarker strategies, cloud infrastructure, external partnerships, and organizational redesign. Attribution therefore requires more than comparing firms that call themselves AI-enabled with historical industry averages.

Mechanisms it reveals

  • A qualifying-AI rule must be fixed before outcomes are observed.
  • Matched controls must account for indication difficulty, target novelty, modality, sponsor quality, and start year.
  • Within-sponsor staggered adoption may provide stronger comparisons than cross-company branding categories.
  • Negative controls can test whether apparent gains reflect general sponsor execution rather than AI-sensitive activities.

Questions this lens makes unavoidable

  • What observable treatment defines meaningful AI exposure at the program level?
  • Can staggered deployment, randomized workflow trials, or threshold rules provide a credible source of causal variation?
  • Which co-interventions must be measured to prevent AI from receiving credit for broader modernization?
Regulatory science x evidence production

Regulators approve evidence, not computational novelty. AI can reduce cost only where its outputs are accepted within validated workflows or help sponsors produce more reliable clinical and manufacturing evidence; otherwise additional validation may offset savings.

Mechanisms it reveals

  • Model provenance, data integrity, version control, and context-of-use validation can add compliance work.
  • Patient-selection models may reduce sample size but narrow the label or require companion diagnostics.
  • Novel combinations may require evidence isolating each component's contribution, limiting shortcuts from computational predictions.
  • Regulatory acceptance may vary by whether AI supports internal decisions, trial operations, endpoints, or product-critical manufacturing controls.

Questions this lens makes unavoidable

  • Which AI uses require regulator-facing validation rather than remaining internal sponsor tools?
  • Do AI-selected populations shorten trials after accounting for screening failures and diagnostic development?
  • Will evidentiary requirements by 2035 absorb, preserve, or amplify operational savings?
05

Hidden assumptions

Assumptions embedded in the original question, and what follows if they do not hold.

AI is a stable, separable intervention.

AI ranges from routine predictive analytics to generative chemistry, multimodal biological models, trial optimization, and regulatory automation, usually bundled with non-AI changes.

If it is false

A single industry-wide treatment effect becomes uninterpretable; the forecast must be decomposed by use case and attribution rule.

The present cost base is the correct comparison.

The causal benchmark is the 2035 cost trajectory without AI, which may already improve through assays, biomarkers, automation, platform trials, and accumulated biological knowledge or worsen through harder targets and richer evidence demands.

If it is false

A measured decline from today's cost could overstate AI's contribution, while rising nominal R&D spending could conceal genuine AI savings.

Savings will appear as lower spending.

Firms may reinvest productivity gains in more programs, more complex indications, additional evidence, or increased probability of technical success.

If it is false

Aggregate expenditure ceases to be a valid test; expected cost, output, and portfolio risk must be examined jointly.

One percentage applies meaningfully across human therapeutics.

The addressable cost share and binding uncertainty differ radically across small molecules, antibodies, vaccines, and cell or gene therapies.

If it is false

The broad claim may need replacement by modality-specific forecasts with distinct mechanisms and ceilings.

More accurate models necessarily improve decisions.

Economic value depends on calibration, error costs, workflow adoption, and whether decision-makers act on model outputs.

If it is false

Technical progress can coexist with negligible or negative cost effects.

06

Hidden connections

What this question resembles outside its obvious domain.

The rebound effect in scientific search

This resembles the Jevons paradox: making a unit of search cheaper can increase total use of search rather than reduce total resource consumption. In drug R&D, the key empirical question becomes whether additional candidates improve the approval denominator or merely consume downstream clinical capacity.

Prediction is not the scarce complement

The thesis resembles productivity puzzles in which information technology becomes valuable only after complementary organizational redesign. AI predictions cannot lower expected cost if incentives reward program continuation, committees distrust model-based stops, or budgets are allocated by candidate count rather than expected portfolio value.

Drug development as reliability engineering

A development program is closer to a series system than a single optimization problem: target validity, candidate properties, safety, efficacy, manufacturing, and evidence generation must all work. Improving one already-reliable component produces little system-level gain if another component dominates failure, suggesting that AI value should be measured through changes in the full failure-mode distribution.

The denominator is a governance choice

Cost per approval is not merely an accounting output; it reflects which experiments an organization permits itself to run and when it stops them. AI may therefore create more value through disciplined abandonment than through generating successful molecules, an effect conventional discovery benchmarks rarely capture.

08

What would change the thesis

Unresolved variables, ranked by how much the conclusion moves when they resolve.

  • High impact

    Matched clinical phase-transition and approval probabilities

    Persistent, risk-adjusted improvement across complete AI-assisted cohorts would directly lower expected cost per approval; no improvement would confine much of the value to upstream efficiency.

  • High impact

    Timing and cost of terminated programs

    Evidence that AI moves failures materially earlier would support savings even without higher approval rates; unchanged or later termination would weaken the thesis.

  • High impact

    Fully loaded expected cost per approval under one accounting standard

    This is the outcome the forecast implicitly concerns. It must include failed programs, shared infrastructure, external payments, time, and an explicit capitalization rule.

  • High impact

    Share of the end-to-end cost base causally affected by AI

    A small addressable share creates an arithmetic ceiling below 30% unless AI also improves attrition or duration.

  • Medium impact

    Elapsed time from program initiation to regulatory decision

    A sustained reduction would lower capitalized cost, but its magnitude depends on the baseline, cost of capital, and whether shorter timelines reflect AI rather than indication selection.

  • Medium impact

    Portfolio response to cheaper experimentation

    If savings are reinvested into more or harder programs, cost per approval may fall while total expenditure rises; if candidate proliferation lowers selection quality, even unit economics may fail to improve.

  • Medium impact

    Modality and disease-area mix

    Concentration in AI-responsive small-molecule workflows would strengthen a scoped version of the claim but not an all-therapeutics industry claim.

09

Questions to ask before proceeding

Each one resolves an uncertainty that materially affects the thesis.

  1. 01What exact numerator, denominator, baseline year, geography, currency basis, inflation index, and capitalization convention define the 30% claim?
  2. 02What operational rule qualifies a program as AI-assisted, and at what date must that status be assigned to prevent retrospective relabeling?
  3. 03For each modality, what fraction of expected cost lies in stages AI can plausibly affect by 2035?
  4. 04Do complete AI-associated cohorts show higher risk-adjusted phase-transition probabilities after controlling for indication, target novelty, modality, sponsor, and calendar year?
  5. 05When AI-associated programs fail, do they terminate earlier and after less cumulative spending than matched conventional programs?
  6. 06How much of any observed timeline improvement remains after separating AI from indication selection, biomarkers, assay advances, outsourcing, and organizational redesign?
  7. 07Does AI adoption reduce trial enrollment time, protocol amendments, screen-failure rates, site underperformance, or monitoring cost in prospective comparisons?
  8. 08Are productivity gains retained as lower expected cost per approval, converted into higher approval output, or absorbed by larger and riskier portfolios?
  9. 09Which discovery, clinical, manufacturing, and regulatory bottlenecks become binding as candidate-generation capacity expands?
  10. 10What approvals, discontinuations, and phase transitions have occurred since mid-2024 among candidates publicly described as AI-discovered or AI-designed?
10

Research roadmap

What to investigate, what evidence to obtain, and how to verify it.

1

Define the estimand and counterfactual

Convert the slogan into one falsifiable primary claim plus modality-specific secondary claims.

  • Fix geography, baseline year, 2035 price basis, inflation index, cost boundary, capitalization rule, and no-AI counterfactual.
  • Choose cost per approved therapeutic as the primary denominator and separately report out-of-pocket cost, capitalized cost, time, and approval output.
  • Write a prospective rule for qualifying AI exposure and assigning hybrid programs.

SignalThe thesis strengthens if a plausible decomposition permits a 30% reduction without relying on undefined spillovers; it weakens if the selected denominator or counterfactual makes the number internally inconsistent.

2

Public evidence still retrievable: construct the program universe

Create a dated, complete cohort of publicly identified AI-associated therapeutic programs and conventional comparators.

  • Search ClinicalTrials.gov and other relevant official trial registries for sponsor, intervention, phase, dates, status changes, and termination reasons.
  • Check current FDA approval databases and regulatory documents for approvals or review outcomes involving sponsor-attributed AI-associated candidates.
  • Archive dated sponsor pipeline pages, press releases, presentations, and partnership announcements while recording that AI attribution is company-stated unless independently substantiated.

SignalA growing mature cohort with traceable outcomes makes the claim testable; sparse, selectively labeled, or mostly preclinical cohorts keep it too early to test.

3

Current retrieval coverage gap: recover comparative clinical evidence

Obtain public peer-reviewed analyses of phase transitions, attrition, timelines, and approval outcomes.

  • Search PubMed, Crossref, Web of Science, and journal databases for comparative AI-associated versus conventional pipeline studies.
  • Extract cohort definitions, censoring dates, control construction, conflicts of interest, therapeutic-area adjustment, and survival methods.
  • Recalculate results where data permit using multi-state survival models and competing risks rather than naive proportions.

SignalRisk-adjusted gains that persist under complete cohorts and censoring corrections strengthen the thesis; disappearance after adjustment indicates selection or maturity bias.

4

Current retrieval coverage gap: decompose the cost base

Estimate the maximum end-to-end savings compatible with published cost shares and phase-specific mechanisms.

  • Retrieve peer-reviewed drug R&D cost studies that separate discovery, preclinical, clinical phases, failures, duration, and cost of capital.
  • Normalize estimates to one currency year while preserving alternative accounting assumptions.
  • Build modality-specific sensitivity models linking AI effects on task costs, transition probabilities, termination timing, and duration to expected cost per approval.

SignalThe thesis strengthens if conservative assumptions across several credible cost decompositions exceed 30%; it weakens if 30% requires near-elimination of costs AI cannot plausibly influence.

5

Current retrieval coverage gap: test financial and operating claims

Compare company-stated productivity with disclosed spending, pipeline throughput, failures, and obligations.

  • Retrieve SEC EDGAR filings and official investor disclosures for public AI-drug-discovery companies, including R&D expense, acquired in-process R&D, collaboration payments, milestone obligations, headcount, and pipeline counts.
  • Build program histories from dated official sponsor disclosures, including discontinuations and partnered assets rather than successes alone.
  • Distinguish platform expenditure from asset-level expenditure and cash cost from contingent economics.

SignalStable or improving cost per risk-adjusted clinical asset alongside complete failure reporting supports operational leverage; rising throughput without better transitions suggests candidate proliferation.

6

Genuinely private evidence requiring diligence

Resolve AI attribution and fully loaded cost using matched internal portfolio data.

  • Request program-level ledgers for all AI-assisted and conventional programs initiated during 2018-2028, quarterly through the latest available date, including internal labor, compute, laboratory work, CRO spending, CMC, clinical operations, allocated overhead, collaboration payments, and capitalized-time assumptions.
  • Request a frozen program roster with AI tools used, deployment dates, decision logs, candidate nominations, phase entries, holds, terminations, reasons, and approvals for the same window.
  • Request matched-cohort specifications by modality, indication, target novelty, sponsor team, and start year, plus audit access to verify that failed and discontinued programs were not excluded.
  • Request model recommendation logs and governance records showing whether AI changed go, stop, experiment, protocol, or enrollment decisions.

SignalAuditable evidence of lower fully loaded expected cost, earlier failure, or improved transitions under common accounting would materially strengthen the thesis; benefits confined to retrospective benchmarks would weaken it.

7

Integrate evidence into a falsifiable 2035 forecast

Produce a forecast distribution rather than a single unsupported percentage.

  • Estimate separate causal effects for discovery cost, clinical operating cost, attrition, failure timing, and duration.
  • Model adoption diffusion, validation burden, bottleneck migration, and portfolio reinvestment under explicit scenarios.
  • Report modality-weighted median, downside, and upside outcomes against the fixed no-AI counterfactual.
  • Define annual update triggers based on cohort maturity, approvals, discontinuations, regulatory acceptance, and audited cost evidence.

SignalThe original thesis becomes supportable only if the probability-weighted industry estimate clears 30% under defensible adoption and attribution assumptions, not merely in an optimistic discovery scenario.

Investment implications

What this examination could mean for investors.

  • If AI-driven discovery induces a Jevons paradox, firms may increase their total volume of clinical experimentation, potentially leading to higher aggregate capital expenditure despite lower marginal costs per candidate.

    This shift in portfolio behavior creates an exposure to capital allocation risks where companies may prioritize total output over the cost-per-approval denominator.

  • Pricing power and long-term margins could erode if AI-enabled patient selection leads to narrower trial labels or increased reliance on expensive companion diagnostics.

    The shift from broad population medicine to granular patient segments relies on the assumption that AI can optimize therapeutic efficacy without significantly inflating regulatory compliance and diagnostic costs.

  • The efficacy of AI as a moat depends on whether firms use it for disciplined abandonment of doomed programs rather than merely accelerating the entry of new molecules into clinical trials.

    Value capture is conditional on organizational governance and the ability to exit failed programs early, as technological capability alone cannot overcome irreducible biological uncertainty.

  • Industry structure could shift toward higher operational complexity if AI productivity gains in molecular generation create downstream bottlenecks in toxicology, manufacturing, and clinical trial validation.

    The constraint on value shifts from the search phase to the capacity of clinical operations and regulatory infrastructure to process an increased volume of AI-generated candidates.

Consequences to examine, drawn from the research above. Not investment advice and not a recommendation regarding any security.

Have a thesis of your own?

Examine it →