
NIH is translating decades of health records into a format machines can read, converting 70 years of biomedical data into a unified language for artificial intelligence.
Standardizing 12 Petabytes of Records
The effort centers on BioData Catalyst, a cloud-based ecosystem developed by the National Heart, Lung, and Blood Institute in collaboration with the National Library of Medicine and the Office of Data Science Strategy.
According to Sweta Ladwa, chief of the Scientific Solutions Delivery Branch at the NHLBI’s Information Technology and Applications Center, the agency holds over 12 petabytes of multimodal data. This collection includes genomics, clinical imaging, sleep studies, and sensor information. It spans long-running projects like the Trans-Omics for Precision Medicine program, which tracks roughly 180,000 individuals.
Access to this volume of information is one thing, but making it usable for AI requires more than just storage. The technical challenge is interoperability. A cardiovascular variable from a 1990s study must mean the same thing as a similar metric in a recent pulmonary fibrosis project. NHLBI has built a linked data modeling language pipeline to address this gap. Ladwa described the system as a “converter box approach,” where data from a source is plugged in and automatically formatted for analysis.
Related: Major Healthcare Deals Rise in 2025
The pipeline maps information across multiple standards, including LOINC (Logical Observation Identifiers Names and Codes), FHIR (Fast Healthcare Interoperability Resources), and HPO (Human Phenotype Ontology). NHLBI pairs this automated mapping with clinical validation, working with pulmonologists to ensure concepts align correctly. The team uses “publicly available metadata” rather than patient data to train these mapping algorithms, ensuring that a specific hypertensive medication is recognized as the same ontological concept regardless of the study it appears in.
Connecting Research to Clinical Care
While NHLBI focuses on interoperability for existing research, the Office of Data Science Strategy is working to map research-grade standards into clinical systems where new data is generated daily.
Susan Gregurick, NIH’s associate director for data science and director of ODSS, described a push to align NIH research standards with the United States Core Data for Interoperability. This standard is used by electronic medical record systems for accreditation. The work began in oncology and is expanding into other areas, including a cardiovascular partnership with the NHLBI.
The practical effect is that cardiovascular phenotypes appearing in a patient encounter—whether inside or outside a formal study—can be captured in a format researchers can use. “The impact for that sort of cross-agency collaboration is really huge,” Gregurick said. She noted that while this capability drives AI, it serves a broader purpose for medical practice.
Related: AI Bridges Talent Gaps Without Hiring More Staff
Underpinning this work is the National Library of Medicine, which provides the infrastructure for biomedical research. Lisa Federer, acting director of the NLM’s Office of Strategic Initiatives, explained that the library offers assets like Medical Subject Headings (MeSH) and Common Data Elements. These tools standardize how research data is described and collected across the institute.
Federer highlighted a shift in who consumes this information. “We’re not just thinking about humans,” she said. “We’re thinking about how machines are consuming data as well.” NLM has noted a rise in the number of bots and AI agents crawling its digital repositories. Curating data for machine consumption differs significantly from curating for humans because AI agents do not read context clues the way a researcher does.
Building a Foundation for the Future
Gregurick noted that NIH invested nearly $400 million last year in AI-related research grants. While this funding is substantial, the less visible investment in pipelines, ontologies, and standards may be more critical.
Without interoperable standards, 70 years of health data remains locked in the formats it was created in. By establishing these common languages, the data becomes a foundation for the next generation of AI-driven medical research. When researchers can connect disparate datasets across the nation and internationally, the power of the information increases, ultimately helping to treat those affected by diseases and disorders.