The Market Bets on Your Readout: What Biopharma Prediction Markets Get Right, and What Nobody Can Prove Yet

Black and gold playing cards, poker chips, and dice on a dark table.
By Ignacio Sancho-Martinez, PhD | 5 October 2026

Section 3 explains the forecast-scoring methods. Readers interested mainly in their use in dealmaking and investment can continue from section 2 to section 4.

On July 16, 2026, Kalshi opened a pilot of event contracts on clinical-trial endpoints and FDA approvals, using AppliedXL to help determine how the contracts resolve. One launch contract asks whether AriBio's AR1001 meets the registered primary endpoint of its Phase 3 POLARIS-AD trial in early Alzheimer's disease. Another asks whether the FDA approves Gilead and Arcellx's anito-cel for relapsed or refractory multiple myeloma. The resolution criteria are fixed before trading begins and refer to public records, such as the endpoint registered on ClinicalTrials.gov or an FDA approval letter. [1]

These prices give deal teams and investors another estimate to consider before a readout. The difficulty is deciding how much weight to give it. Trading remains limited, and the public sources reviewed for this article contain no calibration record showing whether the quoted probabilities match subsequent outcomes. For now, I would compare a market quote with an asset-specific probability-of-success estimate and examine the trading behind it before using it in a decision. The confidentiality risks deserve attention even while the forecasting evidence is incomplete.

1. Which historical success rate is relevant?

Wong, Siah and Lo analyzed 406,038 clinical-trial entries associated with 21,143 compounds from 2000 to 2015. By reconstructing development paths, they estimated that 13.8% of industry-sponsored programs entering Phase 1 would reach approval. The oncology estimate was 3.4%. These are probabilities for a program progressing through development, rather than the chances that an individual trial succeeds. [2]

The estimates also varied within the study. For oncology, the rate fell to 1.7% in 2012 before recovering to 8.3% in 2015. The authors later issued a corrigendum covering specific phase-transition tables, so phase-by-phase figures need to come from the corrected journal version. [2][3]

BIO/QLS reported an overall Phase 1-to-approval estimate of 7.9% for 2011 to 2020; Citeline's April 2024 summary of Biomedtracker's 2014–2023 analysis reports 6.7%. The periods overlap and the methods differ, so the estimates cannot establish a like-for-like decline in success rates. [4][5]

Citeline's April 2024 summary of Biomedtracker's 2014–2023 analysis reports transition rates of 47% from Phase 1 to Phase 2, 28% from Phase 2 to Phase 3, 55% from Phase 3 to filing, and 92% from filing to approval. These figures show why the development stage matters when choosing a benchmark: a program that has reached filing has already survived the earlier stages. [5]

Narrower studies describe quite different groups. In lymphoma, 6.1% of agents entering Phase 1 reached approval, with a mean 7.9 years from first Phase 1 initiation to approval [6]. A study of venture-backed biopharma companies reports 14.1% overall and 8.7% for cancer programs [7]. Among Phase 3 trials in genitourinary oncology, 51.3% led to approval [8]. A study of US first-in-human anticancer trials from 2007 to 2017 reports a 10.3% probability of approval [9].

Patient selection can change the comparison too. Wong and colleagues report a path-level probability of success of 10.7% for oncology programs using biomarkers for patient selection, compared with 1.6% for programs without that selection. That is an association within the study, not evidence that adding a biomarker would give any particular program the same improvement. [2]

Before using any of these rates in a valuation, ask what the study counted and where its observation began. An agent entering Phase 1 faces a different set of remaining risks from a program already in Phase 3. A historical rate is useful only to the extent that its population and outcome resemble the asset and decision being assessed.

Two panels showing overall Phase 1-to-approval estimates from three reports and rates for cohorts with different phases, populations, and units of analysis.
Figure 1. Panel A shows overall Phase 1-to-approval estimates: Wong, Siah and Lo (13.8%), BIO/QLS 2021 (7.9%), and Citeline/Biomedtracker 2024 (6.7%). Their study periods and methods differ, so the estimates are not a controlled time series. Panel B groups reported rates by population, phase, and unit of analysis; these values are not directly comparable. The 2024 rate is reported in Citeline's April 2024 summary of the 2014–2023 analysis. [2][4][5][6][7][8][9]

2. How much trading is behind the price?

The AppliedXL white paper accompanying Kalshi's launch reports lifetime trading volumes of $3,000 to $30,000 for most biopharma contracts. It cautions that “greater liquidity is still needed before making definitive statements about market accuracy.” With few orders available, a relatively small trade can move the displayed price, and other traders may be slow to move it back. Completed trading volume tells us how much has changed hands. Liquidity concerns how much could be traded now near the quoted price. [10]

Fierce Biotech reported less than $50,000 across Kalshi’s biotech-related contracts on May 21, 2026 [11]. On July 28, Crypto Briefing reported more than $100,000 across roughly 13 newly launched markets tied to companies including Sanofi, Eli Lilly, and Gilead [12]. These are press-reported aggregates for different groups of contracts and dates, not a like-for-like measure of growth. Neither report provides a per-contract breakdown.

In Polymarket records retrieved through its Gamma API on October 5, 2026 at 13:22 UTC, the retatrutide market showed $593,832 in cumulative volume and the skin-cancer-vaccine contract showed $133,701. These are October 5 totals, not September 23 snapshots. The latest CLOB price observations returned for September 23 were 2.95% for retatrutide at 23:00:31 UTC and 65.5% for the vaccine at 23:00:29 UTC [13][14][15][16]. PharmaTher's ketamine approval contract had $6,505 traded before resolving YES on August 9, 2025. It was the only resolved real-money approval contract with a documented public outcome located in this review [17]. That finding is limited to the reviewed material, rather than a claim that no other contracts have resolved.

Endpoint Arena offers another forecast setting for clinical-trial outcomes. Fierce’s May report described a paper-trading pilot with contracts tied to trials at Palisade Bio, argenx, and Jazz Pharmaceuticals [11]. Its current Season 8 documentation says only its AI models can trade and the public API is read-only [18]. Neither setup provides public real-money trading volume comparable to Kalshi or Polymarket.

The much larger figures reported for prediction markets overall cover different activity. Crypto Briefing reported $44.8 billion in combined Kalshi and Polymarket trading volume across all events in June 2026 [19]. That unaudited monthly total includes sports and other non-biopharma events. It cannot tell us how easily an investor could trade a particular clinical-trial contract.

Log-scale comparison of reported biopharma contract volumes, with unaudited sector-wide totals shown separately.
Figure 2. Biopharma observations include PharmaTher’s ketamine contract, Polymarket’s skin-cancer-vaccine and retatrutide markets, and press-reported Kalshi aggregates. Polymarket volumes are cumulative Gamma API totals retrieved October 5, 2026 at 13:22 UTC. The separate $44.8 billion context bar is a reported June 2026 total across all events, not a biopharma figure. These observations cover different periods and populations. The AppliedXL white paper cautions that “greater liquidity is still needed before making definitive statements about market accuracy.” [10][11][12][13][14][17][19]

To assess a quote, a team needs more than the volume displayed beside it. Bid-ask spreads, available orders, and the timing of recent trades would help establish whether the price reflects active trading or an old transaction. Thin trading gives us a reason to be cautious about the quote; determining whether it was a good forecast requires a record of predictions and outcomes.

3. How to test a probability forecast

The same scoring methods apply to internal probability models and to prices on binary event contracts.

A forecast system is calibrated when events assigned a probability near 70% occur about 70% of the time across a sufficiently large set of comparable forecasts. Testing it requires probabilities recorded before the outcomes, fixed criteria for resolving each event, and enough observations to estimate a pattern. The contracts reviewed here generally name public resolution documents. The missing evidence is a sufficiently large record linking pre-resolution prices to resolved outcomes.

For one binary event, the Brier score is BS = (p - o)², where p is the forecast probability and o is 1 if the event occurs or 0 if it does not. Lower scores are better. The log score uses the negative logarithm of the probability assigned to the outcome that occurred, penalizing a confident miss more heavily. [20]

Across many forecasts, calibration-in-the-large tests for systematic over- or underprediction. Calibration slope assesses whether the predicted probabilities are too extreme or too compressed relative to observed outcomes. Discrimination asks a different question: do the forecasts tend to assign higher probabilities to events that occur than to those that do not? A useful evaluation needs to examine both calibration and discrimination. [21][22]

PharmaTher's ketamine contract illustrates the scoring calculation, although we found no public pre-resolution price for it. The observed outcome was YES [17]. If someone had assigned a probability of 0.25, the Brier score would be (0.25 - 1)² = 0.5625. A forecast of 0.50 would score 0.25, and a forecast of 0.90 would score 0.01.

These are hypothetical teaching values, not historical market prices. The 0.90 forecast scores best for this outcome, but one result cannot establish whether the forecaster is well calibrated. The public price record reviewed here is also insufficient to calculate a pooled Brier score for the markets.

Contract design determines what would count as the outcome in that evaluation. The launch materials describe templates for trial endpoints, approvals, advisory committee votes, approval timing, drug shortages, label expansions, and emergency use authorizations. They identify resolution sources and flag distinctions such as accelerated versus full approval. Kalshi's CFTC product certification for FDA-approval-date contracts dates to June 11, 2025, before the pilot launch. [10][23]

One example in the white paper makes the importance of wording clear: a trial-endpoint contract resolved YES while the separate approval contract resolved NO. The trial had met its primary endpoint, but that was not the same event as approval [10]. Evaluating either forecast against the other contract's outcome would give the wrong score.

Taxonomy of biopharma event-contract templates, resolution sources, examples, and associated platforms.
Figure 3. Seven biopharma event-contract templates, with their resolution sources, examples, and platforms. The table describes contract design rather than a census of active markets; the source materials do not name a public example for every template. Endpoint Arena examples refer to the paper-trading pilot described in May 2026. The registry identifies CDR-SB change from baseline to week 52 as a POLARIS-AD primary outcome. [1][10][11][18][23][24]

4. Compare forecasts of the same event

On September 23, 2026, Polymarket showed 2.95% for retatrutide receiving FDA approval by the end of 2026 in the latest returned CLOB observation at 23:00:31 UTC. It is tempting to compare that with Wong, Siah and Lo's 13.8% overall Phase 1-to-approval estimate and conclude that the market is pessimistic. But the contract asks about one asset that has already advanced through development, with only a few months left before its deadline. The historical estimate starts at Phase 1 and follows programs through to approval. [2][13][15]

The same problem arises with the 65.5% latest returned CLOB observation at 23:00:29 UTC on September 23 for a skin-cancer vaccine receiving approval by December 31, 2027. Comparing it with the 3.4% oncology-wide Phase 1 estimate ignores the particular asset's development stage and the contract's deadline. Neither comparison, on its own, tells us whether the market price is too high or too low. [2][14][16]

A later-stage benchmark would address some of that mismatch. The genitourinary oncology study, for example, reports approval following 51.3% of Phase 3 trials [8]. It is closer in development stage, but the therapy area still does not match retatrutide or the vaccine. Nor can the difference between the lymphoma and venture-backed studies be assigned to the denominator alone: their populations and study designs differ too [6][7].

The sources reviewed here contain no verified obesity-specific Phase 3-to-approval rate matched to retatrutide's contract and deadline. Without that comparison, we cannot use the broad historical rates to judge the 2.95% quote. For anito-cel, the launch materials do not provide a public price, and a Phase 3 asset would in any case need a later-stage benchmark than the oncology Phase 1 estimate. [1][2][13]

A valuation team should first write down the event it is forecasting. Is the question eventual approval, approval by a particular date, or a trial meeting a specified endpoint? Then choose a historical cohort with the closest therapy area, development stage, outcome definition, and unit of analysis. Only after those choices does the difference between the internal estimate and a market quote become interpretable.

Where no close cohort exists, show the nearest available benchmarks separately and explain their limits. Keep the contract wording, deadline, dated quote, and available trading information alongside the model assumptions. When the outcome arrives, that record will allow the team to assess what it actually predicted.

Dated Polymarket probability snapshots beside broad, unmatched historical base rates, with hypothetical Brier scores for one resolved contract.
Figure 4. Panel A compares latest returned CLOB price observations for September 23, 2026 (2.95% for retatrutide at 23:00:31 UTC and 65.5% for the skin-cancer vaccine at 23:00:29 UTC) with broad Phase 1-to-approval estimates to show why the pairs are not calibration tests: they differ in phase, population, and deadline. Panel B gives illustrative Brier scores for hypothetical pre-resolution prices of 0.25, 0.50, and 0.90 on the resolved PharmaTher ketamine contract: 0.5625, 0.25, and 0.01. These prices are hypothetical; the scores use one outcome and are not a market-accuracy statistic. [2][13][14][15][16][17]

5. What the published research can tell us

The abstract of a 2026 SSRN working paper by Mir and Qadir describes an FDA-approval model using phase success rates, company size, advisory committee votes, trial endpoint strength, and historical approval rates. The authors evaluate it using Brier scores, ROC curves, and realized trading profitability. They report an advantage over options-implied probabilities and test a trading rule triggered when the two estimates diverge beyond a threshold. These are the authors’ reported findings, available in the publisher-deposited abstract. [25]

This is a preprint, and its comparator is options-implied probability. It suggests that clinical and company information can improve forecasts derived from options prices in the studied setting. It does not establish how well Kalshi or Polymarket event contracts forecast the same decisions. Testing those markets would require their own price histories.

Hamill and colleagues examine equity-market reactions to FDA approval announcements. They find that reactions concentrate on the event day once the event window is specified appropriately. That is consistent with approval announcements bringing new information to equity markets. It does not measure whether pre-announcement event-contract probabilities were calibrated. Equity returns, options-implied probabilities, and event-contract prices need to be evaluated on their own terms. [26]

We found no verifiable historical price-and-outcome dataset for FDA-approval contracts on Intrade in the public sources reviewed. Without such a record, no conclusion about their calibration follows. This is a limit of the review, not proof that no such contracts or research existed.

A study of current markets could archive timestamped prices before resolution, apply the published outcome rules, and compare forecast probabilities with observed frequencies over a substantial sample. Until there is such a record, the studies above offer methods and related findings, but no direct verdict on the accuracy of these event contracts.

6. Confidential information creates a separate risk

Employees, investigators, contractors, and trial participants may learn something about a clinical milestone before the public does. A contract tied to that milestone gives the information a direct trading use. Counsel writing in Bloomberg Law identify potential insider-trading and manipulation risks, including attempts to affect trial timing, design, or reported results to benefit an open position. These are possible risks, not findings that particular traders have done so. [27][28]

Jerrob Duffy of Fried Frank discusses how a breach of a duty of trust or confidentiality may support an insider-trading theory. Debevoise & Plimpton describe the risk that pharmaceutical employees could try to influence trial conduct or reporting. Their June commentary predates later CFTC insider-trading enforcement actions reported by Bloomberg Law on September 16, 2026. Companies therefore have reason to seek advice on how that authority is being applied. [27][28]

Platforms have begun responding. Polymarket partnered with Chainalysis in April 2026, and Kalshi introduced integrity measures in June, including whistleblower channels and employment verification. Kalshi lists biotech contracts after enrollment closes and focuses on late-stage trials. That timing narrows some opportunities to trade on information, but confidential findings can still emerge while a trial is underway and before results or regulatory decisions become public. [1][29][30]

Pillsbury partner David Oliwenstein recommends reviewing policies, assessing company-specific exposure, and naming prediction markets directly in employee training. His article identifies clinical-trial outcomes and FDA drug-approval timelines as potential contract topics. These concerns are separate from any evaluation of predictive accuracy. [31]

7. What teams can do now

For BD and IR teams: if a market quote enters a discussion, keep its date, contract wording, volume, and any available bid-ask or depth information. Use a large disagreement with your internal forecast as a reason to examine the assumptions in both. A screenshot without that context is insufficient grounds to change a valuation.

For founders and clinical leaders: check that registry entries and endpoint definitions accurately describe the planned analysis. Contracts may use those public records to resolve a bet. Continue to make updates through normal scientific and operational governance; external trading should not determine trial design or reporting.

For investors: ask which event the quoted probability refers to, when the price was observed, and what historical cohort supports the comparison. Request the model assumptions behind any claim that the market confirms or challenges a valuation. An options-market study cannot substitute for an evaluation of event-contract forecasts. [13][25]

For compliance officers: identify who receives confidential clinical or regulatory information and review how trading and disclosure policies apply to them. Name prediction markets in training, provide a way to escalate questions, and obtain counsel's assessment of the company's specific exposure. [27][31][32]

Need a probability estimate for a biopharma decision? We build probability-of-success models using relevant clinical evidence and documented assumptions.

CONTACT US NOW

Sources

  1. [1] Kalshi News. “Kalshi launches biotech prediction markets pilot program.” https://news.kalshi.com/p/kalshi-biotech-prediction-markets
  2. [2] Wong CH, Siah KW, Lo AW. “Estimation of clinical trial success rates and related parameters.” Biostatistics, 2019. https://pmc.ncbi.nlm.nih.gov/articles/PMC6409418/
  3. [3] Corrigendum: “Estimation of clinical trial success rates and related parameters.” Biostatistics, 2019. https://pmc.ncbi.nlm.nih.gov/articles/PMC6409416/
  4. [4] BIO, QLS Advisors, Informa Pharma Intelligence. “Clinical Development Success Rates and Contributing Factors 2011–2020.” https://go.bio.org/rs/490-EHZ-999/images/ClinicalDevelopmentSuccessRates2011_2020.pdf
  5. [5] Citeline. “Why Are Clinical Development Success Rates Falling?” White paper, April 2024. https://www.citeline.com/-/media/8f1a3eb60cde4827bb513ea6f78316e2
  6. [6] Luo Z, et al. “Clinical trial success rate in lymphoma: fate of trials and agents from 2000 to 2019.” Blood Advances, 2025. https://pmc.ncbi.nlm.nih.gov/articles/PMC12964034/
  7. [7] Kang SY, Liu M, Huang SS. “Success Rates of Venture Capital Investment in Biopharmaceutical Development.” Medical Care, 2026;64(7):413–418. Publisher abstract (Ovid / Wolters Kluwer). https://www.ovid.com/jnls/lww-medicalcare/abstract/10.1097/mlr.0000000000002325~success-rates-of-venture-capital-investment-in
  8. [8] Steidle M, et al. “Study Endpoint Fulfillment and FDA Approval Success in Phase III Clinical Trials of Genitourinary Malignancies.” Clinical Genitourinary Cancer, 2025;23(6):102436. NCBI bibliographic record and abstract. https://eutils.ncbi.nlm.nih.gov/entrez/eutils/efetch.fcgi?db=pubmed&id=41076989&rettype=abstract&retmode=text
  9. [9] Mukaida A, Maeda H. “A cross-sectional study on the first-in-human trials of anticancer drugs in Japan and the United States and the probability of approval.” International Journal of Clinical Oncology, 2025. https://pmc.ncbi.nlm.nih.gov/articles/PMC12474682/
  10. [10] AppliedXL and Kalshi. “Biopharma’s Public Probability: The State and Future of Prediction Markets in Drug Development.” https://www.appliedxl.com/research/biopharma-public-probability-report
  11. [11] Fierce Biotech. “Betting on biotech: Prediction markets set sights on clinical trials.” May 21, 2026. https://www.fiercebiotech.com/biotech/betting-biotech-prediction-markets-set-sights-clinical-trials
  12. [12] Crypto Briefing. “Kalshi and Polymarket Now Let You Bet on Drug Approvals, and Not Everyone Is Thrilled About It.” July 28, 2026. https://cryptobriefing.com/kalshi-polymarket-drug-approval-betting/
  13. [13] Polymarket Gamma API. Retatrutide contract rules and cumulative market volume, retrieved October 5, 2026 at 13:22 UTC. https://gamma-api.polymarket.com/events?slug=fda-approves-retatrutide-this-year
  14. [14] Polymarket Gamma API. Skin-cancer-vaccine contract rules and cumulative volume for the December 31, 2027 market, retrieved October 5, 2026 at 13:22 UTC. https://gamma-api.polymarket.com/events?slug=skin-cancer-vaccine-fda-approved-by-december-31-2027
  15. [15] Polymarket CLOB. Retatrutide YES-token price history for September 23, 2026 (fidelity 60 minutes; timestamped points). CLOB API price history, retatrutide YES token, September 23, 2026
  16. [16] Polymarket CLOB. Skin-cancer-vaccine YES-token price history for September 23, 2026 (fidelity 60 minutes; timestamped points). CLOB API price history, skin-cancer-vaccine YES token, September 23, 2026
  17. [17] Polymarket Gamma API. “FDA approves PharmaTher Holdings’ Ketamine?” Contract rules, final volume and resolution; approval deadline August 31, 2025. https://gamma-api.polymarket.com/events?slug=fda-approves-pharmather-holdings-ketamine-by-august-31
  18. [18] Endpoint Arena. Season 8 read-only API documentation. https://endpointarena.com/api
  19. [19] Crypto Briefing. “World Cup Prediction Market Volumes Soar to Record Highs.” July 4, 2026. https://cryptobriefing.com/world-cup-prediction-market-volumes-record-highs/
  20. [20] Brier GW. “Verification of forecasts expressed in terms of probability.” Monthly Weather Review, 1950;78(1):1–3. Publisher article record. https://journals.ametsoc.org/view/journals/mwre/78/1/1520-0493_1950_078_0001_vofeit_2_0_co_2.xml
  21. [21] Van Calster B, et al. “Calibration: the Achilles heel of predictive analytics.” BMC Medicine, 2019. https://pmc.ncbi.nlm.nih.gov/articles/PMC6912996/
  22. [22] Van Calster B, et al. “A calibration hierarchy for risk models was defined: from utopia to empirical data.” Journal of Clinical Epidemiology, 2016;74:167–176. NCBI bibliographic record and abstract. https://eutils.ncbi.nlm.nih.gov/entrez/eutils/efetch.fcgi?db=pubmed&id=26772608&rettype=abstract&retmode=text
  23. [23] Commodity Futures Trading Commission. Designated Contract Market Products filing 56561, “Will the FDA approve drug before date” (KalshiEX; certified June 11, 2025). The CFTC record links the associated contract filing. https://www.cftc.gov/IndustryOversight/IndustryFilings/TradingOrganizationProducts/56561
  24. [24] ClinicalTrials.gov. POLARIS-AD study record (NCT05531526), primary outcome CDR-SB change from baseline to week 52. https://clinicaltrials.gov/study/NCT05531526
  25. [25] Mir ZA, Qadir A. “Bayesian Approval Probability Model for FDA Drug Announcements.” SSRN working paper, 2026. DOI: 10.2139/ssrn.6301339. Methods and reported findings cited from the publisher-deposited abstract, not independently checked against the full paper. Publisher-deposited abstract in Crossref
  26. [26] Hamill PA, Hutchinson MC, Nguyen QMN, Mulcahy MB. “FDA approval announcements: Attention-grabbing or event-day misspecification?” Economics Letters 170 (2018):171–174. https://cora.ucc.ie/server/api/core/bitstreams/41096b1a-8b43-462b-8b3f-8ce328637264/content
  27. [27] Duffy J, Liberman D. “Prediction Markets’ Risks Warrant Biopharma Compliance Measures.” Bloomberg Law, September 16, 2026. https://news.bloomberglaw.com/legal-exchange-insights-and-commentary/prediction-markets-risks-warrant-biopharma-compliance-measures
  28. [28] Juergens ET, Zolkind DS, Folly N. “Prediction Markets Trading Must Be Part of Compliance Policies.” Bloomberg Law, June 4, 2026. https://news.bloomberglaw.com/legal-exchange-insights-and-commentary/prediction-markets-trading-must-be-part-of-compliance-policies
  29. [29] CoinDesk. “Polymarket taps Chainalysis to bring Wall Street-level oversight to crypto prediction markets.” April 30, 2026. https://www.coindesk.com/business/2026/04/30/polymarket-taps-chainalysis-to-bring-wall-street-level-oversight-to-crypto-prediction-markets
  30. [30] Kalshi News. “Kalshi to Implement Market Integrity Updates.” June 9, 2026. https://news.kalshi.com/p/kalshi-market-integrity-updates-risk-scoring-employment-verification-whistleblower
  31. [31] Kaile D, Oliwenstein D, Wiktor A. “Prediction Markets and Insider Trading: Compliance Considerations.” Pillsbury, January 6, 2026. https://www.pillsburylaw.com/en/news-and-insights/prediction-markets-insider-trading.html
  32. [32] Hogan Lovells Cadwalader. “Predictions realized: Regulators and enforcement authorities dial up the heat on prediction markets...” August 11, 2026. https://www.hlc.com/en/publications/predictions-realized-regulators-and-enforcement-authorities-dial-up-the-heat-on-prediction-markets