Research design is the overall strategy of a study that specifies:
- Who will be studied
- What will be measured or explored
- When measurements happen (once, twice, or repeatedly over time)
- How groups are formed, compared, or left untouched
- How the resulting data will be analyzed
It is the plan or framework used to conduct a study. Two teams can ask the same question, collect data from the same hospital, and reach opposite conclusions simply because one used a comparison group and the other did not.
How to choose a research design
Here is a series of steps that help you move from a research problem to the right design that can answer it.
|
Step |
What it answers |
Example |
|
Research problem |
What gap, tension, or practical difficulty requires research? |
Post-operative infection rates are rising in our surgical unit |
|
What precisely do we want to know? |
Does a pre-surgical chlorhexidine bath reduce infection at 30 days? |
|
|
What will we actually do to answer it? |
To compare 30-day infection rates between bathed and non-bathed patients |
|
|
What relationship do we predict? |
Chlorhexidine bathing reduces 30-day infection rates |
|
|
Study design |
How will we gather evidence? |
Randomized controlled trial with blinded outcome assessment |
What is study scope?
Study scope defines the boundaries: which settings, which time period, which subgroups, and, just as importantly, what the study will not attempt. Scope creep is one of the most common reasons student projects stall. A clearly bounded scope makes recruitment targets realistic and protects the analysis from becoming a fishing expedition.
Components of a Research Design
Before choosing a family of design, you need to settle several elements that appear in every study.
Population and sampling
- Population: the entire group you want your conclusions to apply to (e.g. all adults with type 2 diabetes in a district)
- Sampling frame: the actual, accessible list from which you draw participants (e.g. the diabetes clinic register). The gap between the population and the sampling frame is a frequent hidden source of error; patients who never attend clinic are invisible to your study.
- Sampling: the method of drawing participants from the frame
|
Sampling approach |
Type |
Typical use |
|
Simple random |
Probability |
Small, well-defined frames |
|
Stratified random |
Probability |
When subgroups must be represented proportionally |
|
Systematic |
Probability |
Consecutive clinic attendees |
|
Cluster |
Probability |
Multi-site or community surveys |
|
Convenience |
Non-probability |
Pilot work, hard-to-reach groups |
|
Purposive |
Non-probability |
Qualitative studies seeking information-rich cases |
|
Snowball |
Non-probability |
Stigmatised or hidden populations |
Data sources
Primary data are collected by the researcher directly for the study at hand through observation, interviews, measurement, or a questionnaire survey. Secondary data already exist: registries, medical records, census figures, published datasets. Many efficient designs, such as a retrospective chart review, are built entirely on secondary data.
Pilot work
A pilot test/pilot study is a small-scale rehearsal run before the main study. It answers questions the protocol cannot:
- Do participants interpret the questions the way you intended?
- How long does one data collection session actually take?
- Is the recruitment rate high enough to hit the target sample size?
- Do the outcome measures produce usable variation, or does everyone score the same?
Types of Research Designs: Quantitative Designs
Quantitative research measures variables numerically and uses statistics to describe patterns or test relationships. It answers how many, how much, how often, and does X cause Y.
Quantitative designs are of two types: experimental research (the investigator intervenes) and observational study designs (the investigator watches without intervening).
Descriptive and exploratory designs
Descriptive research characterizes a situation without testing relationships: prevalence of anemia, distribution of waiting times, demographic profile of service users. It answers “what is happening?” but not “why?”.
Exploratory research is used when a topic is poorly understood and the researcher is not yet in a position to specify hypotheses. It aims to generate ideas, refine questions, and identify variables worth measuring later. Exploratory work is often qualitative, but small quantitative surveys serve the same purpose.
A case report describes a single patient or unit in detail. It generates hypotheses and documents rare events, but it cannot establish frequency or causation. A case series extends this to several patients.
Correlational designs
Correlational research measures two or more variables in the same group and quantifies how strongly they move together, without manipulating anything. It can establish that a relationship exists and estimate its strength and direction, but on its own it cannot establish which variable causes the other, or whether an unmeasured third variable explains both.
Experimental designs
In experimental research, participants are assigned to conditions by the researcher, ideally at random. This is the only design family that supports strong causal claims, because randomization distributes both known and unknown confounders evenly across groups.
Core features:
- Experimental group: receives the intervention being tested (often called intervention group)
- Control group: receives placebo, usual care, or nothing, providing the comparison against which effect is judged
- Random allocation to groups
- Blinding: concealing group allocation from participants (single-blind), from participants and investigators (double-blind), or additionally from data analysts (triple-blind). Blinding reduces performance and detection bias, i.e., people behave and report differently when they know which arm they are in.
Between-subjects vs within-subjects
Experimental research can also be between-subjects or within-subjects:
|
Feature |
||
|
Structure |
Different participants in each condition |
Same participants experience all conditions |
|
Comparison |
Group A vs Group B |
Each person acts as their own control |
|
Sample size needed |
Larger |
Smaller |
|
Individual variation |
A source of noise |
Controlled out |
|
Main threats |
Group non-equivalence |
Order, carryover, and practice effects |
|
Typical safeguard |
Randomization |
Counterbalancing or washout periods |
A pretest-posttest design measures the outcome before and after an intervention. In its simplest single-group form it is weak, because any change could be due to maturation, seasonal effects, or regression to the mean rather than the intervention. Adding a control group measured at the same two time points converts it into a far stronger design.
Quasi-experimental designs
A quasi-experimental design includes an intervention but lacks random allocation. Groups are formed by convenience, geography, policy, or self-selection, for example, comparing two schools where one adopted a new curriculum. These designs are common when randomization is impossible, unethical, or politically unacceptable.
- Strength: feasible and often high in real-world relevance
- Weakness: groups may differ systematically at baseline, so the comparison is vulnerable to confounding
- Common variants: non-equivalent control group, interrupted time series, regression discontinuity
Types of observational designs
An analytical study compares groups to examine associations between exposure and outcome, in contrast to purely descriptive work. The three most popular analytical observational designs are cross-sectional, case-control, and cohort.
|
Design |
Direction |
What it measures |
Strengths |
Limitations |
|
Snapshot at one point in time |
Prevalence; associations |
Fast, cheap, good for planning services |
Cannot establish temporal order |
|
|
Backwards from outcome to exposure |
Odds ratio |
Efficient for rare outcomes; small samples |
Recall and selection bias; cannot give incidence |
|
|
Forwards from exposure to outcome |
Incidence; relative risk |
Clear temporal sequence; multiple outcomes |
Expensive; long follow-up; loss to follow-up |
A longitudinal study follows the same participants across repeated measurements. This is what allows researchers to distinguish change within individuals from differences between individuals (see our useful comparison of longitudinal vs cross-sectional research).
Prospective vs retrospective
The distinction of prospective vs retrospective concerns when data collection occurs relative to the outcome:
|
Prospective |
Retrospective |
|
|
Outcome status at study start |
Has not yet occurred |
Already occurred |
|
Data quality |
Researcher controls measurement |
Constrained by what was recorded |
|
Cost and time |
High |
Low |
|
Bias profile |
Attrition; observer effects |
Recall bias; missing and inconsistent records |
|
Example |
Multi-year cohort of new mothers |
Retrospective study of readmissions from discharge summaries |
A retrospective chart review (extracting variables from existing medical records) is the most common form of retrospective research in clinical settings. It is fast and inexpensive, but the researcher is limited to what clinicians happened to document, and undocumented does not mean absent.
Survey research
Survey research collects standardised information from a sample in order to describe a population. A questionnaire survey may be administered on paper, online, by telephone, or face to face. Design decisions that determine survey quality:
- Item wording: avoid double-barrelled, leading, and ambiguous questions
- Response format: Likert scales, categorical options, open text
- Validation: prefer instruments with published reliability and validity evidence over items invented on the spot
- Response rate: low rates raise the risk that respondents differ systematically from non-respondents
- Mode effects: people answer sensitive questions more honestly in self-administered formats
Qualitative Designs
Qualitative research explores meaning, experience, process, and context. It uses words, images, and observation rather than counts, and works with small purposively selected samples analyzed in depth. The goal is not statistical representativeness but conceptual richness and transferability.
|
Approach |
Central question |
Typical data |
|
Phenomenology |
What is the lived experience of X? |
In-depth interviews |
|
Grounded theory |
What theory explains this process? |
Interviews, constant comparison |
|
Ethnography |
What is the culture of this group? |
Prolonged fieldwork, participant observation |
|
Case study |
What is happening in this bounded system? |
Multiple sources: interviews, documents, observation |
|
Narrative inquiry |
How do people story their experience? |
Life histories, diaries |
|
Content/thematic analysis |
What themes recur across accounts? |
Interviews, focus groups, open text |
Distinguishing features:
- Sampling is purposive and continues until data saturation, when new participants stop yielding new insight
- Data collection and analysis are iterative rather than sequential
- The researcher is an instrument; reflexivity about their influence is a methodological requirement
- Rigor is judged by credibility, dependability, confirmability, and transferability rather than by internal and external validity in the statistical sense
Qualitative work is indispensable when you need to know why an intervention failed, how patients understand a diagnosis, or what barriers staff perceive. These are questions a questionnaire survey with fixed options cannot answer.
Mixed Methods Research
Mixed methods research deliberately combines quantitative and qualitative components within a single study, integrating them to produce insight neither could deliver alone. Simply running a survey and a few interviews and reporting them in separate chapters is not mixed methods. It is important to integrate the two strands.
|
Design |
Sequence |
Purpose |
Example |
|
Convergent parallel |
Simultaneous |
Corroborate or contrast findings |
Survey scores alongside interviews on the same topic |
|
Explanatory sequential |
Quantitative → Qualitative |
Explain unexpected numerical results |
Interviews with patients who dropped out of a trial |
|
Exploratory sequential |
Qualitative → Quantitative |
Build a measure grounded in local reality |
Focus groups used to develop and then validate a scale |
|
Embedded |
One nested in the other |
Add process insight to a trial |
Qualitative process evaluation inside a clinical trial |
Qualitative vs Quantitative vs Mixed Methods Research
|
Dimension |
Quantitative |
Qualitative |
Mixed |
|
Purpose |
Measure, test, predict |
Understand, interpret, explore |
Explain and confirm |
|
Question type |
How many? Does X cause Y? |
Why? How? What is it like? |
Both |
|
Sample |
Large, probability-based where possible |
Small, purposive |
Both, often nested |
|
Data |
Numbers, scores, counts |
Text, images, field notes |
Both |
|
Analysis |
Statistical |
Thematic, interpretive |
Integrated |
|
Role of hypothesis |
Central and pre-specified |
Usually absent |
Varies by strand |
|
Key quality criteria |
Validity, reliability, generalizability |
Credibility, transferability |
Both plus integration quality |
|
Flexibility during study |
Low, protocol is fixed |
High, design evolves |
Moderate |
How to Choose a Research Design: Practical Tips
- State the problem, then the question. If you cannot write the research question in one sentence, the design will not save you.
- Classify the question type. Prevalence → cross-sectional. Cause → experimental or quasi-experimental design. Meaning → qualitative. Rare outcome → case-control. Incidence over time → cohort study.
- Check feasibility. Time, budget, ethics approval, access to the sampling frame, and available expertise all constrain the realistic options.
- Choose the strongest design the constraints allow. Randomise if you ethically can; if you cannot, plan how you will handle confounding analytically.
- Specify measures and outcomes in advance. Pre-specification prevents outcome switching.
- Run a pilot wherever possible.
- Write the analysis plan before collecting data.
Quick reference
|
If your question is… |
Consider |
|
How common is this condition? |
Cross-sectional study |
|
What happened to these patients in the past? |
Retrospective chart review |
|
Does this intervention work? |
Randomized clinical trial |
|
What exposures preceded this rare disease? |
Case-control study |
|
Who develops the outcome over time? |
Cohort study / longitudinal study |
|
Do these two variables move together? |
Correlational research |
|
Why did the program fail? |
Qualitative case study |
|
What do users think, at scale? |
Questionnaire survey |
|
Does the intervention work, and why? |
Mixed methods research |
|
Is this phenomenon even understood yet? |
Exploratory research, case report |
Quality of Your Study Design
A design is only as good as its defences against systematic error. Bias is any systematic deviation of results from the truth, and it cannot be fixed by increasing the sample size. A larger biased study is simply more precisely wrong.
Internal validity
Internal validity is the degree to which observed effects can be attributed to the intervention rather than to something else. Common threats:
- Selection bias: groups differ before the study begins
- Confounding: a third variable relates to both exposure and outcome
- Measurement bias: outcomes assessed differently across groups, which blinding helps prevent
- Attrition bias: participants who drop out differ systematically from those who remain. Attrition bias is particularly dangerous in a longitudinal study, because those lost to follow-up are often the sickest, poorest, or least satisfied, i.e., the very people whose outcomes matter most. Report dropout rates by arm and use intention-to-treat analysis.
- Recall bias: especially acute in a case-control study, where people with a disease search their memory harder for exposures.
External validity
External validity concerns whether the findings hold outside the study conditions. Generalizability depends on:
- How representative the sample is of the target population
- How typical the setting is of routine practice
- Whether inclusion criteria were so restrictive that real-world patients were excluded
- Whether the intervention can be delivered elsewhere with the same fidelity
There is a persistent trade-off: tightly controlled designs maximize internal validity but often at the expense of external validity, while pragmatic real-world designs do the reverse.


Comment