

Summary
Well known for her heroic efforts in the Crimean war and innovations in nursing practice, Florence Nightingale was also a gifted statistician. She believed that data collected from London hospitals could and should be used to generate evidence in pursuit of improved patient care. In her work “Notes on Hospitals” she expands on this idea, describing how the future of medical informatics hinges on the need for standardized terminology and data structure. Today, the scope of data has expanded from hospitals in one urban city to databases around the world representing millions of patients, but the need for high-quality data to support research has not changed. This thesis focuses on improving the foundational science of data standardization, data quality, and fitness-for-purpose in federated networks to increase trust in evidence and ultimately improve patient lives.
Data Standardization
Chapter 1 examines the practicalities of standardizing diverse data sources to the OMOP Common Data Model (CDM) through the experience of building the European Health Data and Evidence Network (EHDEN). Anecdotally, we knew the challenges and pitfalls of the extract, transform, and load (ETL) process, but this study allowed us to observe and support twenty-five data partners from 11 countries as they standardized their data. The largest factor that contributed to a data partners’ success was team composition. Such groups should be balanced with individuals with deep knowledge of the source data, individuals with expertise in the OMOP CDM and vocabularies, and individuals with the technical skills to implement the data transformation. The most consistent hurdle for teams was vocabulary mapping, where source codes that could not be automatically assigned to standard concepts required manual review. Large language models are now being leveraged to reduce this burden, and the introduction of community-driven mappings by the OHDSI vocabulary team marks an important step toward a more adaptive and collaborative future for vocabulary standardization. ETL teams also wrestled with issues of granularity, equivalency, and the practical question of how much effort to invest in resolving rare or ambiguous codes, though these challenges can often be mitigated through use case–driven mapping strategies that focus effort where it has the greatest analytical relevance. Finally, this work highlighted the need for clear data governance agreements prior to the initiation of data standardization. Membership in a federated network presupposes the participation of the data partner in network research. Governance requirements to allow such participation should be clear and unambiguous. Without clarity, barriers can emerge later during analysis or publication.
Chapter 2 showcases an opportunity to implement standard methodologies to address assumptions about the underlying population of a database that were previously handled at the time of analysis. 80 definitions of observable time across 11 databases were used to generate incidence rates of five different outcomes of interest. The analysis showed that even relatively minor differences in defining when observation begins and ends for a patient can materially alter incidence rate estimates. Importantly, standardization offers a mechanism to implement consistent approaches to observable time across diverse databases while still accommodating variation in data capture methods.
Data Quality
Standardization not only enables methodological innovation and federated research but also creates the conditions for systematic quality assessment. While scientists can design and test studies without accessing patient-level data, data owners must still execute analyses and return results to study leads. For this process to be trustworthy, there must be a way to evaluate whether each database meets minimum quality standards without compromising patient privacy. Chapter 3 describes the development of the Data Quality Dashboard (DQD) to meet this need. The standardized scaffold of the OMOP CDM provides a foundation for implementing quality checks at scale, allowing data partners to generate transparent, reproducible assessments of their databases. Results can then be communicated to coordinating bodies without exposing patient-level information. Requiring databases mapped to the CDM to run the DQD represents an important first step toward consistently reporting data quality, thereby strengthening confidence in the data used to support observational studies in federated networks.
Chapter 4 then utilizes the data standardization exercises taking place among data partners in EHDEN to measure how well the DQD improves the quality of databases in federated networks. The tool was applied to 25 databases mapped to the OMOP CDM spanning diverse data origins, data capture methods, and underlying populations. Requiring partners to complete the DQD ensured that each database met baseline quality expectations before participating in federated studies. In particular, the tool performed well in assessing conformance to CDM structural requirements, providing assurance that standardized study packages could run without error. While study-specific checks remain essential, the DQD established a consistent foundation of conformance, completeness, and plausibility across the network. This application demonstrated the value of the DQD not only for individual data partners but also for federated networks, ensuring the data are ready when the needs arise.
Fitness for Purpose
Chapter 5 gives an overview of the more than 200 data sources from 29 countries participating in the EHDEN network. The breadth of healthcare systems, data capture methods, and populations represented across these sources underscores the importance of drawing on multiple datasets to capture the full range of patient experiences in Europe. At the same time, this diversity highlights the persistent challenge that no single database can be expected to support every (or even most) research questions posed to a network. Determining whether a database is fit-for-purpose therefore is an essential step before initiating studies. Relying on site-level feasibility assessments based on patient-level data is neither sustainable nor scalable, given the time and resource demands. As federated networks expand in both size and visibility, including within regulatory initiatives such as DARWIN EU®, new approaches are needed to assess database fitness-for-purpose in a way that is both efficient and privacy-preserving.
Chapter 6 offers a solution to this challenge by introducing and validating a method for performing early fit-for-purpose assessments in federated networks using only precomputed summary statistics. By eliminating the need to access patient-level data during study planning, this approach substantially reduces the time, resources, and governance burdens typically required at this stage. The method, known as Database Diagnostics, goes beyond a simple yes-or-no feasibility check by providing diagnostic insights into why a particular study may not be viable in a given database. These insights can guide both data curation efforts and refinements in study design, improving the efficiency and interpretability of network research. In doing so, Database Diagnostics fills an important methodological and operational gap, offering a scalable and privacy-preserving way to evaluate database readiness and ultimately strengthen the generation of real-world evidence across federated networks.
❉ ❉ ❉
This thesis drew parallels to the historical work of Florence Nightingale to illustrate that the needs for high-quality standardized data are not new to the field of medical informatics. However, with the unprecedented amount of data and technologies available, we have both the opportunity and obligation to improve the science of data standardization, quality, and fitness-for-purpose to ensure the reliability of evidence. This thesis argues that these elements make up the foundation of federated observational health research. It is through their interrogation, application, and transparency that we as researchers can trust the science built on these foundations to provide actionable insights that improve patient care. It may start with the data, but it ends with the patient.

















