From Real-World Data to Trustworthy Evidence: Designing Studies That the Data Can Support
Dear Colleagues,
One of the things I appreciated most about my recent webinar with Harvard epidemiologist Miguel Hernán was how naturally our two perspectives came together. Miguel focused on target trial emulation and the importance of getting the study design right from the beginning. I focused on something I have spent much of my career thinking about, whether the data themselves are fit for the question being asked.
Those two ideas are difficult to separate. You can have an excellent dataset and still design a poor study. You can also have a thoughtful study design that asks more of the data than they can reasonably provide. Producing trustworthy evidence requires us to consider both at the same time.
Although much of our webinar focused on causal inference, I think that lesson is relevant to almost anyone conducting research with real-world data. Before we think about matching, modeling, or any other analytic technique, we need to understand the question, define the study carefully, and determine whether the available data can actually support the conclusions we hope to draw.
Start With the Question, Not the Analysis
One of Miguel’s points that stayed with me was his description of target trial emulation as a framework for thinking rather than an analytic method.
The idea is to first describe the randomized trial you would ideally conduct if you could. Who would be eligible? What strategies would you compare? When would follow-up begin? What outcomes would you measure? What exactly are you trying to estimate? Only after those pieces have been clearly defined do you ask how well the observational data can approximate that design. The discipline behind this approach is broadly useful across RWD studies.
It is very easy, particularly when working with large datasets and accessible analytic tools, to move quickly from an interesting idea to an analysis. We should resist that temptation. Before choosing a statistical method, we should be able to clearly explain the research question, the population, the exposure or intervention, the comparator where relevant, the index date, the outcome, and the study timeline. Those decisions are not just setup for the analysis; they are part of the science.
Ask Whether the Data Are Fit for This Purpose
Another point I have come to appreciate over many years of working with real-world data is that there is no universally “good” or “bad” dataset.
If your cohort depends on something like body weight or BMI, for example, you may need information available in an EHR that would not be captured in claims. If medication exposure is central to the study, you need to think about exactly what information is required. Is documentation of an exposure enough, or do you need dispensing information, days supplied, dose, units dispensed, or changes in exposure over time?
The same applies to outcomes. Some outcomes are consistently medically attended and relatively well captured in structured data. Others may be under-recorded, documented inconsistently, or difficult to identify with a single code. In those cases, we may need to construct a phenotype or proxy definition and be very thoughtful about what that definition actually represents. Also incredibly important to consider whether to use a specific or broad outcome definition, and the implication of that choice within the context of the research question, dataset, and various influences on data capture.
One of the most important habits researchers can develop is, instead of asking, “Is this a good dataset?” ask, “Is this the right dataset for this question?”
Understanding What the Data Can Support
Data quality can sometimes sound like something that happens before the “real” research begins. In my experience, it is very much part of the research itself and often requires as much or more time and effort as implementing the research question. And if I had my way, we would not be talking about “data quality” but rather “data characterization”, because that is really what we are doing when we are assessing whether a dataset is fit for purpose.
Real-world data come from real healthcare systems, which means they also reflect the complexity of those systems. Units of measure may differ across institutions. Coding practices change over time. Data capture and level of documentation can vary across settings and by the purpose of the data capture, e.g., for clinical documentation or financial reimbursement? Even something that appears straightforward at first can look very different once you begin examining the underlying data more closely. That is why data characterization should be part of every step of the study.
Do the cohort characteristics make clinical sense? Are there unexpected changes in coding over time? Are key variables available often enough to support the analysis? Does an apparent trend reflect something happening clinically, or could it reflect a change in the way the information was captured?
These questions sometimes lead us to refine the original study design. I do not view that as a failure of the process, it’s a feature. In many cases, it is exactly what careful RWD research should look like. We define the question, examine the data, learn what the data can and cannot support, and refine the design accordingly. The important distinction is that those changes should be thoughtful, transparent, and methodologically justified rather than made simply because a particular analysis produced an unexpected result.
A Practical Framework for Real-World Research
If I were distilling our webinar conversation into a practical approach for researchers, I would keep six things in mind:
Start with a clear research question.
Know the population, exposure or intervention, comparator, outcome, timeframe, and scientific objective before you begin the analysis. Use target trial emulation to help crystallize your thinking.
Define the study design before selecting the analytic method.
For causal questions, that may mean explicitly thinking through the target trial. For other studies, it means being equally deliberate about eligibility, index dates, follow-up, and outcome definitions.
Ask whether the data are fit for that specific purpose.
Can the data meaningfully capture each of the major components of the design? Where are the gaps?
Spend time understanding the data.
Look at completeness, coding patterns, temporal trends, plausibility, and anything else that could affect interpretation.
Choose methods that fit both the question and the data.
More complicated methods are not automatically better methods. The analysis needs to make sense for the study you designed and the information you have.
Expect some iteration and document it.
Real-world research rarely moves in a perfectly straight line. Learning more about the data may lead to a better definition, a different timeframe, or a more appropriate analysis. Those decisions should be transparent and driven by the science.
This is less about following a rigid checklist than developing a disciplined way of approaching real-world research.
Building Evidence We Can Stand Behind
Real-world data have opened opportunities to study populations, treatments, outcomes, and clinical questions that are difficult or impractical to examine through traditional clinical research alone. But access to large amounts of data does not remove the need for careful study design. If anything, it makes that discipline more important.
What matters is whether we can explain why the study was designed the way it was, why the data were appropriate for the question, what the important limitations are, and how those limitations should shape interpretation. That is ultimately what I took away from my conversation with Miguel. Good RWD research is not about finding the most complicated method or the largest possible cohort. It is about making a series of thoughtful decisions and being able to explain them.
For those interested in digging further into these topics, I encourage you to watch the full From Data to Evidence: A Framework for RWD Study Design & Causal Inference webinar. Miguel and I go into more detail on target trial design, time zero, longitudinal treatment data, and some of the practical challenges that come with working with observational data.
I also encourage you to take advantage of the scientific, educational, and methodological resources available through TriNetX and to engage with our teams as you are developing your research. Those conversations are often most useful early in the process, when there is still time to refine the question, evaluate feasibility, and think carefully about what the available data can support.
As the ways we use real-world data continue to evolve, I think the fundamentals remain fairly simple: ask a clear question, design the study thoughtfully, understand your data, choose your methods deliberately, and be transparent about what the evidence can tell us.
Best,
Jeffrey Brown, PhD
Chief Scientific Officer
TriNetX





