Micro image of an icy area

Fitness for Purpose: Why “Is This Data Good?” Gets Data Quality Wrong

Article filter
Share this article

Key takeaways:

  • Data quality and fitness for purpose are not the same standard. A dataset can be structurally sound, clinically plausible, and well-harmonized and still be entirely wrong for a specific research question. 
  • The instinct to label data as universally “good” or “bad” creates a false sense of confidence that can obscure deeper evidence risks. Context determines whether data are appropriate, and context requires scientific judgment in addition to the quality score. 
  • Operationalizing fitness for purpose means going beyond baseline integrity checks to explicitly evaluate whether the data can credibly support the question being asked, every time, for every study. 

People ask me all the time whether a particular dataset is any good. My answer is always the same: compared to what? And for what purpose? These are not trivial questions. They are the only questions that matter. 

The real-world data (RWD) community has made remarkable progress on data quality over the past decade. Standards have been established. Harmonization frameworks have been adopted. Automated quality metrics are now applied at scale. Walk into almost any research organization today and you will hear confident claims about high-quality data. 

What you will hear less often is a rigorous answer to whether that data is appropriate for the specific question being asked. This tension surfaced repeatedly at the ISPE Annual Meeting in Milan last month, where conversations about generating robust real-world evidence (RWE) consistently circled back to a gap the field has not yet closed: the distance between data that is technically sound and data that is genuinely fit for purpose. That gap is where a significant share of evidence quality problems originate. And it is a problem that does not resolve itself as data become more abundant or analytics become more sophisticated. It requires deliberate scientific judgment, applied before the analysis begins. 

Data Quality and Evidence Quality Are Not the Same Threshold

It is worth being precise about what we mean when we talk about data quality, because the term is doing a lot of work in this field and not always accurately. 

Data quality determines whether an analysis is possible. It addresses whether a dataset is structurally consistent, clinically plausible, interpretable, and harmonized across systems. These are necessary conditions. They are not sufficient ones. Evidence quality determines whether conclusions drawn from that data can withstand scientific, regulatory, and clinical scrutiny. It requires that the research question is aligned to the intended use, that methodology has been stress-tested against assumptions, that findings are independently reproducible, and that limitations are explicitly understood and handled. 

Confusing one for the other is where defensible evidence begins to fail. High-quality data can still produce low-quality evidence when the research question is poorly framed, the methodology is misaligned, or the underlying assumptions are left unexamined. Conversely, even data with known limitations can support robust evidence when those limitations are explicitly understood and appropriately handled. 

High-quality data are an input. High-quality evidence is a scientific judgment. 

The Invisible Gap Problem

One of the most common and consequential data fitness failures does not involve anything obviously wrong with the data. It involves a failure to understand what the data does not capture. 

Consider vaccination studies derived from health system data. Vaccinations administered at retail pharmacies or community clinics may not flow back into patient medical records. Within the dataset, those vaccinations simply appear missing. If that limitation is not accounted for, data on who was and was not vaccinated will be incorrect, not because the data are wrong, but because of a lack of understanding of what the data do and do not capture. 

Absence of evidence is not evidence of absence. But unless you understand how and why data were captured, it is very easy to draw exactly that conclusion. 

This is the invisible gap problem. The data pass every structural and plausibility check. The analysis runs cleanly. The findings look reasonable. And the conclusions are wrong, not because anyone made an error in the conventional sense, but because a critical assumption about data coverage was never tested. 

Three Dimensions Worth Distinguishing

At TriNetX, we evaluate data quality across three distinct but connected dimensions, and the distinctions matter. 

1. Data Characterization

Data characterization provides a clear, descriptive understanding of what a dataset truly represents. This includes population coverage, care settings, geography, time horizon, source systems, and how the data have evolved over time. It makes explicit what the data do and do not contain. It provides context. It does not, on its own, determine whether data are appropriate for a given question. 

2. Baseline Data Integrity

Baseline data integrity is project-agnostic quality assurance focused on usability. Data must be structurally sound and clinically plausible. Values must fall within reasonable bounds. Concepts must be consistently harmonized across systems. Gaps must be visible and intelligible, not hidden. These checks establish whether data can be analyzed reliably. They do not presume that they should be. 

3.Fitness for Purpose

This is where true data quality work happens. It evaluates whether the data can credibly support a specific research question. It requires scientific judgment, not scoring. If the question is vaccine uptake and the data do not capture where vaccines are administered, the data may be excellent and entirely inappropriate. 

These three dimensions are not interchangeable. Moving from the first to the third requires progressively more scientific judgment and progressively more specificity about the question being asked. Organizations that stop at the first or second dimension and call it quality assurance are leaving the most important work undone. 

Quality Is Continuous, Not Confirmed

One of the more persistent misconceptions in this field is that data quality is something you establish once and then rely on. In practice, quality must be continuously monitored, not treated as a one-and-done exercise. 

At TriNetX, hundreds of expert, automated quality metrics are applied to every incoming dataset using established informatics standards, including Kahn’s Harmonized Data Quality Assessment Framework. Code-description pairs are validated. Deviations are flagged early. Feedback loops with data partners drive corrective action. There are thousands of checkpoints. And we have developed machine-learning techniques to detect privacy-driven date shifting, an industry-wide practice that can bias time-sensitive research, so analysts can be alerted and interpret results appropriately. 

The result is data that are not only clean, but deeply understandable, transparent, and in context. That last part matters as much as the first. Data you cannot explain is data you cannot defend. 

The Standard the Field Needs

Generalized claims about high-quality data exacerbate the trust problem in this field. They encourage a false sense of confidence that can obscure deeper evidence risks. There is no such thing as data that are universally good or bad. A dataset that is ideal for one question may be entirely inappropriate for another, and no quality score will tell you which is which. 

What will tell you which is which is disciplined scientific judgment applied before the analysis begins. That means spending more time understanding the data than analyzing it. It means beginning every study with a structured data deep dive. It means ensuring every research team includes true data characterization expertise. And it means treating assumptions about availability, completeness, and fitness as unproven until verified repeatedly with data partners. 

This is a standard that matters across the full spectrum of RWE work. For the pharmacoepidemiology community, fitness for purpose is foundational to credible study design. For health economics and outcomes researchers preparing evidence packages for health technology assessment (HTA) bodies and payers, it is equally critical: a submission that cannot demonstrate data appropriateness for the clinical question at hand will not survive regulatory or payer scrutiny, regardless of how sophisticated the analysis.  

As the ISPOR Europe community convenes in Vienna next month, these questions of evidence credibility and data appropriateness will be at the center of the conversation. The standard the field agrees to hold itself to in those discussions will shape the trustworthiness of the evidence that reaches decision-makers for years to come. 

Just as quality data are now the expectation, scientific discipline for producing RWE must become the industry-wide standard. Because improving human health depends not just on producing more answers, but on producing more answers we can trust. 

This post is adapted from TriNetX’s new eBook, The Evidence Trust Crisis: Why More Data, Faster Analytics, and Easier Access Are Accelerating Output While Eroding Trust. 

About Matvey Palchuk, MD 

Matvey Palchuk, MD is an expert in clinical informatics, semantic interoperability, and data quality with extensive experience in healthcare information management, data acquisition and interoperability, the design of user interfaces for point-of-care applications, information modeling, knowledge management, and quality measure reporting.