Pharma General Ontology (PGO)-Terminology: Building a Semantic Backbone for FAIR and AI-Ready Knowledge Graphs

Presentations

Delivering Data Driven Value

Pharma General Ontology (PGO)-Terminology: Building a Semantic Backbone for FAIR and AI-Ready Knowledge Graphs

This talk was presented at Bio-IT World 2026 (Knowledge Graphs session, Boston) by Giovanni Nisato, PhD — Project Manager, Pistoia Alliance.

As life science innovation becomes increasingly data-driven, semantic interoperability is essential to scaling FAIR data and trustworthy AI. Yet the pharma ecosystem still lacks a shared reference vocabulary. Today, the same concept (compoundmoleculesubstancedrug) is used inconsistently across domains and organizations, deep domain ontologies proliferate in disconnected silos, and there is no cross-pharma lingua franca. Teams spend enormous effort on semantic reconciliation before any real work begins, and interoperability stays aspirational.

The Pharma General Ontology (PGO)-Terminology project is a Pistoia Alliance–led, industry-governed initiative that tackles this challenge head-on. It is building a community-governed reference set of core pharmaceutical concepts. Think if it as a shared “lingua franca” to serve as a semantic backbone for interoperable knowledge graphs, cross-domain data integration, and AI-ready foundations across the pharma ecosystem.

In this talk, Giovanni Nisato introduces PGO-Terminology’s scope, design principles, current stage, and long-term vision.

What you’ll learn:

  • The problem: Why limited semantic interoperability creates costly “semantic silos” that slow the path to data-centricity across R&D, manufacturing, clinical, regulatory, and commercial domains.
  • What is PGO-Terminology? A community-governed, cross-domain reference set of core concepts, published open access. It is not a replacement for MedDRA, SNOMED, or ChEBI, not a top-down mandate, and not (at this stage) a full ontology.
  • How it’s built: An industry governance model with clear definition rules: reuse existing definitions from known sources, keep them human-readable across domains, and give every concept a publicly resolvable URL/IRI.
  • Where the project stands: Phase 1 delivered 17 core Research & Early Development concepts at controlled-vocabulary level. Phase 2 is expanding beyond R&D, moving toward machine-readable output, and mappings to external ontologies.
  • Why it matters for AI: Knowledge graphs are only as good as their underlying semantics. Features of the PGO-Terminology help to reduce hallucination risk in RAG and LLM pipelines.
  • The long-term vision: To deliver a machine-readable core terminology that is discoverable on resources such as EMBL-EBI OLS and BioPortal, to promote data assets tagged with PGO URIs, and to facilitate cross-company, CRO, partner, and regulator data exchange without manual reconciliation.

PGO-Terminology is developed openly on GitHub under permissive licenses (CC BY 4.0 / MIT)

Learn more on the PGO-Terminology project webpage.

Published on: August 5, 2026