Harnessing AI To Expedite R&D

Life Science AI Exchange Round Table: Practical Application

The August 2026 Life Science AI Exchange brought together industry experts from Merck Healthcare KGaA, Novo Nordisk and Sapio Sciences to explore the challenges of validation, trust and regulatory readiness for agentic AI in life sciences. The session examined why traditional approaches to computerised system validation may not translate directly to agentic AI, highlighting the importance of validating AI within its specific context of use, maintaining meaningful human oversight, ensuring data integrity and designing robust audit trails and governance controls. The roundtable also explored practical challenges including reviewer fatigue, explainability at scale and the need for shared verification standards, reinforcing the importance of continuous governance as AI becomes more deeply embedded in regulated workflows.

Learn more about the Life Science AI Exchange here.

Why Do AI Projects Fail in Drug Development and Pharma?

Insights from Multi-Year Pistoia Alliance Studies

This is a peer-reviewed article from the Pistoia Alliance’s LLM and NLP Use Case Database project, published in the Journal of Computer-Aided Molecular Design (2026).

Artificial intelligence is playing an increasingly central role in drug discovery and across the pharmaceutical industry. Yet despite a steady stream of high-profile success stories, many technically successful pilots never translate into sustained business value. Understanding why and what distinguishes the initiatives that scale from those that quietly stall is one of the most important questions facing R&D leaders investing in AI today.

This paper tackles that question with rigor. It brings together three complementary streams of evidence gathered over several years of Pistoia Alliance work:

  • Industry brainstorming workshops with senior professionals, from scientists to executive directors, held across biotech and pharma since 2017.
  • A global executive survey and interviews, conducted with Zühlke Engineering, capturing where demand is highest, where organizational maturity lags, and where collaboration potential is strongest.
  • A statistical analysis of the Pistoia Alliance’s collection of real-world AI, ML, and NLP use cases, one of the largest and most methodically documented collections of its kind, which captures both successes and candid failures across R&D, pharmacovigilance, manufacturing, and beyond.

Read together, these sources tell a remarkably consistent story and point to a conclusion that may be uncomfortable for teams focused primarily on models and algorithms: the factors that most reliably predict whether an AI project succeeds are not the ones most organizations spend the most time on.

What you’ll find in the paper:

  • Which project characteristics actually correlate with business success, and which widely assumed “success factors” are not predictive on their own.
  • Why reproducibility, explainability, and governance are the real limiting factors for adoption, particularly in high-risk domains.
  • How the barriers to AI success mirror those of earlier waves of digital transformation, and what that means for how AI projects should be planned and led.
  • A concise set of practical, evidence-based recommendations for improving the odds that an AI initiative delivers lasting value.

The paper offers a clear and refreshingly hype-free perspective for AI practitioners, R&D leaders, and decision-makers working to turn AI’s promise into real, scalable impact in life sciences.

DOI: 10.1007/s10822-026-00894-3

From Strings to Things: How Semantic Grounding Defends Against AI Hallucinations in Drug Discovery

Large Language Models (LLMs) are rapidly transforming pharmaceutical research, offering unprecedented capabilities across diverse workflows. However, their pervasive adoption introduces a critical challenge: the propensity for fluent AI hallucinations. These outputs often appear grammatically correct and scientifically plausible, yet can be fundamentally inaccurate in ways difficult to detect without a robust, structured ground truth for validation.

At GSK, they are systematically tackling this challenge by embracing semantic grounding. This involves rigorously classifying our data using established ontology classes—transforming unstructured “strings” into well-defined “things”—and replacing ambiguous free-text references to key research entities with persistent Uniform Resource Identifiers (URIs). This presentation will demonstrate why semantic grounding is more than just a data enhancement; it is a fundamental prerequisite for building truly trustworthy and reliable AI systems in pharmaceutical R&D, essential for driving innovation and ensuring data integrity.

Speakers
  • Alice Augustine, GSK
  • Jim Morris, Progress Software

Japan Roundtable Discussion: AI and AI agents in ELN and LIMS workflows

Led by Yusuke Sato of Kyowa Kirin, this Pistoia Alliance roundtable explored the latest trends, practical applications and future potential of AI and AI agents in ELN and LIMS workflows through expert discussion and knowledge sharing.

Pistoia Allianceが主催し、協和キリンの佐藤祐輔氏が進行した本ラウンドテーブルでは、ELNおよびLIMSワークフローにおけるAIとAIエージェントの最新動向、実践的な活用事例、そして今後の可能性について、専門家による意見交換が行われました。

Life Science AI Exchange Seminar Review – June 2026

This seminar explored how life sciences organizations can successfully deploy agentic AI by focusing on trustworthy data, robust governance, rigorous evaluation, and human oversight, demonstrating that sustainable AI value depends as much on data readiness and trust as on advances in AI technology.

US Life Science Informatics Forum – June 2026

The discussion explored how AI, automation, and interoperable data foundations are shaping the “lab of tomorrow,” with a strong emphasis on practical implementation, scientist-led adoption, and collaboration to accelerate research

Life Science AI Exchange: Agentic AI and Technology Best Practices

Agentic AI — systems capable of autonomous planning, tool use, and multi-step reasoning — represents one of the most transformative and potentially disruptive shifts in the AI landscape. For life sciences, the implications are profound: from automated literature review and hypothesis generation to end-to-end orchestration of research workflows.
 
This two-hour seminar brings together technical experts and practitioners to examine the current state of agentic AI technology, explore emerging best practices, and address the governance challenges unique to autonomous systems. How do you validate an agent’s decisions? How do you maintain audit trails? What infrastructure is needed to deploy agents safely at scale?
 
Whether you are evaluating agentic frameworks for the first time or looking to mature your existing deployments, this seminar provides practical guidance grounded in real-world life science experience.
 

Speakers
  • Frederik Steensgaard Gade, Novo Nordisk
  • Josefa Stoisser, Novo Nordisk
  • Nicholas Larus-Stone, Benchling
  • Sean Blake, Sapio Sciences
  • Arindam Sett, Genentech
  • Kelly Mewes, Genentech
  • Edaeni Hamid, Genentech

From Data Access to Data Readiness

The Next Bottleneck in AI-Driven Small Molecule Discovery
AI Is Advancing Rapidly in Drug Discovery — But Is Your Data Ready?

Join leading experts from AI-first biotech, computational chemistry, and pharmaceutical R&D as they discuss how transforming fragmented scientific data into AI-ready intelligence is accelerating modern small molecule discovery workflows.

Why AI-Driven Small Molecule Discovery Needs Better Data

AI models today can generate predictions faster than ever before — but without standardized, context-rich, and integrated datasets, even the most advanced systems struggle to deliver reliable outcomes.

This exclusive panel discussion hosted on the Pistoia Alliance platform will explore why data readiness — not compute power — has become the true bottleneck in scaling AI-driven drug discovery.

Key Takeaways
  • Why data readiness is becoming the biggest bottleneck in AI-driven drug discovery
  • How leading biotech and pharma teams are overcoming fragmented scientific data challenges
  • What it takes to build scalable, AI-ready discovery workflows and datasets
Speakers
  • Saro Passaro, Co-Founder Boltz PBC
  • Daniel Price, VP , Computational Chemistry, Nimbus Therapeutics
  • Dr. Govinda Bhisetti, Former VP & Head of Computational Chemistry, Cellarity
Moderators
  • Vladimir Makarov, Project Manager, Pistoia Alliance
  • Norman Azoulay, Vice President of Product, Platforms and Data, Excelra

AI Impact and Value Poll

In partnership with Thoughtworks we conducted a poll among the 300 delegates at our 2026 Spring Conference to understand the impact AI is having in life sciences R&D. While AI use is becoming widespread across our industry our poll shows AI is not truly embedded across organizations and the value is not reaching into specialized R&D activities. As AI evolves as a capability, collaboration will be essential for closing the gap between AI investment and value. Successful adoption will require getting both the data and people elements right and ensuring we continue to work together alongside regulators to shape the data and standards required to scale AI safely and effectively.

Natural language querying of biological databases with large language models 

Meaningful querying biological databases today requires specialist knowledge of structured languages such as SQL, SPARQL, or Cypher — making data mining slow, labor-intensive, and inaccessible to many researchers. Reliable natural language querying would change that. It would also be a prerequisite for the next generation of AI co-scientist systems: tools that automate scientific hypothesis generation at scale. 

This peer-reviewed paper, produced by the Pistoia Alliance’s Large Language Models in Life Sciences project and published in Drug Discovery Today (May 2026), reports the outcomes of a systematic assessment of current practices in natural language querying with LLMs. 

Highlights 

  • Accurate natural language data mining is a requirement for AI co-scientist systems 
  • Multi-agent LLM systems combined with deterministic queries offer the best accuracy-flexibility balance 
  • LLM agents must be Findable and Reusable (FAIR) and require open API standards 
  • Shared benchmarks for natural language data mining systems are needed across the industry 

Authors 

Vladimir A. Makarov (Pistoia Alliance), Oleg Stroganov (Rancho Biosciences), Laura I. Furlong (MedBioInformatics Solutions), Brian Evarts (Crown Point Technologies), Loes van den Biggelaar (The Hyve), Alexandros Goulas (AbbVie), Etzard Stolte (Roche), Derek Marren (AstraZeneca), and Lars Greiffenberg (AbbVie). 

__________________________________ 

What did the study test? 

The authors evaluated 21 different strategies for translating natural language into structured database queries, using the Open Targets Platform as a real-world target discovery and validation use case. Five LLMs were tested — GPT-4o, Claude 3.5 Sonnet, o1, GPT-4o-mini, and open-mistral-7b — across naïve, template-based, retrieval-augmented, prompt-optimized, and multi-agent approaches. 

What works — and what doesn’t? 

Naïve prompting failed almost universally on complex scientific questions, even when the full database schema was supplied. Template-based strategies achieved 100% accuracy but are rigid: they cannot be transferred to new data sources without substantial human effort, and they do not scale. 

Multi-agent strategies — in which multiple LLM agents challenge one another’s outputs and interact with a human user — achieved 83–98% accuracy on complex questions with the best-performing model (o1).  

This combination of accuracy, flexibility, and adaptability across data sources represents the most practical path forward identified in the study. Entity recognition (for example, correctly resolving “ALS” to “Amyotrophic Lateral Sclerosis”) was the single largest remaining source of error. 

What does the industry need next? 

The paper makes three calls to action for the field: 

  • Combine multi-agent LLM systems with deterministic API calls for simple, high-confidence retrieval tasks. 
  • Develop a shared industry benchmark for natural language data mining that is resilient to LLM background-knowledge contamination — a subtle but significant source of evaluation error. 
  • Establish FAIR-aligned open standards — such as the Model Context Protocol — for the discovery of and engagement with AI agents across commercial and academic systems. 

This research was coordinated by the Pistoia Alliance Large Language Models in Life Sciences project. To learn more about the project, visit https://pistoiaalliance.org/project/benchmarks-for-natural-language-data-mining-with-llms/