Pseudo-cluster randomisation and the contamination problem
Individual randomisation, assigning each participant to an arm independently, is the design most researchers default to, and for good reason: it maximises statistical power for a given sample size. But it has a specific weakness in studies where participants at the same site or under the same care provider can be assigned to different arms: contamination. A care provider delivering both a new intervention and standard care to different patients can end up letting elements of one bleed into the other, consciously or not, diluting the measured difference between arms.
Cluster randomisation, assigning whole sites or providers to a single arm rather than randomising within them, solves the contamination problem directly. It introduces a different one: referral bias. If a provider who prefers a particular treatment approach knows which arm their site has been assigned to, referral patterns into the study can shift in ways that undermine the comparison just as much as contamination would.
Why neither approach solves the whole problem
This is a genuine trade-off, not a solved problem with one obviously correct answer. Individual randomisation risks contamination; cluster randomisation risks referral bias; both risks are real and can meaningfully distort a trial's results if left unaddressed. Work applying pseudo-cluster randomisation in practice looked at exactly this tension in a psychiatric care setting, where both risks were plausible: providers delivering care directly, and referral decisions that could plausibly be shaped by provider preference.
What pseudo-cluster randomisation actually does
The method sits deliberately between the two extremes. Rather than assigning a whole site to one arm, or randomising every individual independently within a site, pseudo-cluster randomisation assigns most participants at a site to one arm, with a smaller, randomised minority assigned to the other. The exact ratio is a design choice, but the logic is the same: enough clustering to reduce contamination meaningfully, enough individual-level randomisation retained to reduce the referral-bias risk that a fully clustered design would carry.
It's a genuinely clever compromise, and like most compromises, it's not free. The statistical analysis has to explicitly account for the resulting clustering structure, and the design is more complex to implement and explain to sites than either pure approach.
The three designs, side by side
Laid out together, the trade-off each design makes is easier to see than in prose alone:
| Design | Contamination risk | Referral bias risk | Statistical power | Complexity to implement and analyse |
|---|---|---|---|---|
| Individual randomisation | High, if a provider serves both arms | Low, since assignment is independent of referral | Highest, for a given sample size | Lowest |
| Cluster randomisation | Low, arms are physically separated | High, if providers can steer referrals knowing their site's arm | Lower, effective sample size shrinks with clustering | Moderate, requires cluster-aware analysis |
| Pseudo-cluster randomisation | Reduced, most participants at a site share an arm | Reduced, a randomised minority still crosses over | Between the other two | Highest, needs an analysis that accounts for the specific mixed structure |
No row in this table is free of trade-offs, which is really the point. Pseudo-cluster randomisation doesn't eliminate either risk, it dilutes both to a level a study team judges acceptable, in exchange for a design and analysis that's genuinely harder to execute correctly than either pure approach.
A concrete illustration of why the choice matters
Picture a trial testing a new counselling approach delivered by the same clinicians who also deliver standard care. Under individual randomisation, a clinician sees some patients assigned to the new approach and others to standard care in the same week. Even a careful, well-intentioned clinician can start blending elements of the new approach into standard-care sessions simply because it's fresh in mind, quietly narrowing the measured difference between arms regardless of whether the new approach actually works better.
Move to full cluster randomisation, assigning that clinician's entire caseload to one arm, and the contamination risk disappears. But now consider what happens if referrals into the study come partly through that same clinician's own recommendations. A clinician who knows their site has been assigned to the new approach, and who believes in it, may refer marginally more suitable candidates toward the study than one assigned to standard care would, shifting who ends up in each arm before randomisation ever gets a chance to balance anything.
Pseudo-cluster randomisation is a direct response to exactly this pairing: most of that clinician's patients follow the site's predominant arm, keeping contamination low, while a randomised minority still crosses over, keeping the clinician's foreknowledge from cleanly predicting every referral's eventual arm.
When this design is actually worth the added complexity
Pseudo-cluster randomisation isn't a good default for every study. It earns its complexity specifically when both of these are true:
- Contamination is a real, plausible risk. If a single care provider or site could reasonably be expected to deliver aspects of both arms to different participants, individual randomisation without a clustering mechanism is likely to understate the true effect of the intervention.
- Referral bias is also a real, plausible risk. If providers or sites have visible influence over which patients are referred into the study, and could plausibly steer referrals based on their own treatment preferences, a fully cluster-randomised design introduces a comparable threat in the other direction.
If only one of these risks applies, a simpler individual or cluster design, with the appropriate safeguards for the risk that does apply, is usually the better choice. Pseudo-cluster randomisation is a tool for the specific situation where both risks are live at once, not a generally superior default.
The broader lesson for study design
The value of a case like this isn't the specific ratio the study settled on. It's the reminder that randomisation strategy is a genuine design decision with real trade-offs, not a default setting to be applied uniformly regardless of context. A study that reaches automatically for individual randomisation without asking whether contamination is plausible, or automatically for cluster randomisation without asking whether referral bias is plausible, is skipping a step that pseudo-cluster randomisation exists precisely to force: naming both risks explicitly, and designing the randomisation strategy around whichever combination of them is actually present in the specific research setting.