Navigating data sharing agreements in multi-site research
Multi-site studies, and increasingly studies planning to make their data available for secondary research, depend on data sharing agreements that spell out who can access what, for what purpose, and under what protections. Recent work looking at the current state of data sharing in research makes a point worth taking seriously: the practical difficulty of navigating these agreements has become a genuine barrier to research, not just an administrative step on the way to it.
Why this has become harder, not easier
It's tempting to assume data sharing gets simpler over time as norms mature and institutions gain experience. In practice, several forces are pulling in the opposite direction:
- Increasingly divergent jurisdictional requirements. Different countries, and in some cases different regions within a country, have adopted meaningfully different rules about health data transfer, meaning a single multi-country study may need distinct agreements tailored to each jurisdiction's specific requirements.
- Growing sensitivity to genomic and other highly identifiable data types. Data that's harder to fully de-identify draws more scrutiny and more conservative agreement terms, which is appropriate but adds real complexity for studies working with this kind of data.
- A proliferation of institution-specific templates and requirements. Without a widely adopted standard agreement, each new institutional partner can bring its own preferred template, multiplying the negotiation effort for a study spanning several sites or institutions.
What a well-structured data sharing agreement actually needs to cover
Regardless of jurisdiction, a handful of elements consistently show up in agreements that hold up well in practice:
- Precise scope of permitted use, specific enough to be meaningful, broad enough not to require renegotiation for every minor variation in planned analysis.
- Clear data security and access control requirements, including who at a receiving institution is actually permitted to access the data, not just which institution holds the agreement.
- Explicit provisions for what happens at the end of the agreed period, destruction, return, or continued retention under what conditions, decided in advance rather than left ambiguous.
- A defined process for amendments, since research questions and collaborations evolve, and a rigid agreement that can't accommodate a reasonable scope change becomes an obstacle rather than a protection.
- Clarity on publication and attribution rights, agreed before data starts flowing, to avoid disputes once results are ready to share.
What tends to go wrong when one of these elements is missing
It helps to be concrete about the failure mode each element actually prevents, rather than treating the list as a compliance checklist to tick off:
| Missing element | What typically goes wrong | Who usually discovers it |
|---|---|---|
| Precise scope of use | A genuinely useful secondary analysis turns out not to be covered, and has to wait for a renegotiation | The statistician planning the analysis, months after data collection ended |
| Access control detail | Data reaches more people at a receiving institution than anyone intended, discovered during an audit | A data protection officer, usually after the fact |
| End-of-period provisions | Nobody agreed what happens to the data when the agreement lapses, so it sits in limbo, neither destroyed nor usable | Whoever tries to reuse it for a follow-up study |
| Amendment process | A reasonable, minor scope change requires renegotiating the entire agreement from scratch | The research team, mid-project, when timelines matter most |
| Publication and attribution rights | A dispute over authorship or data ownership surfaces just as results are ready to submit | Whichever collaborator feels their contribution wasn't reflected |
None of these failures are exotic. They're the ordinary, foreseeable consequence of an agreement that covered the legal minimum required to get data moving, without anyone asking what the study would actually need from that agreement a year or two down the line.
A short worked scenario
Consider a three-country observational study collecting real-world data from participating clinics. The original data sharing agreement, drafted quickly to meet a funding deadline, specifies the primary analysis in detail but says nothing about secondary use. Eighteen months later, a collaborator proposes a genuinely valuable secondary analysis combining the dataset with a separate registry.
Without a scope clause anticipating this, the team faces a choice: renegotiate the original agreement (slow, and dependent on every original signatory being willing to reopen it), or treat the secondary analysis as a new data sharing arrangement entirely (slower still, and duplicating work that a better-scoped original agreement would have avoided). Neither option is available quickly, and the analysis, however valuable, ends up delayed by months for reasons that have nothing to do with the science.
This is the practical cost the research described here is pointing at: not that data sharing agreements are unpleasant to negotiate, but that a poorly scoped one becomes a real constraint on what a study's own data can be used for later.
Practical steps for reducing the friction
- Start data sharing agreement negotiations early, in parallel with other study setup activities, rather than treating it as a final step once everything else is settled.
- Use existing, well-tested template agreements where they exist, rather than drafting from scratch for every new collaboration, adapting rather than reinventing.
- Involve institutional data protection or legal expertise from the outset, rather than after a draft has already been negotiated informally between researchers.
- Document data flows explicitly as part of study design, so the agreement reflects what will genuinely happen to the data, rather than a generic description that doesn't match actual practice.
Why this is worth getting right, not just getting done
A rushed or poorly scoped data sharing agreement doesn't just create legal risk. It can quietly constrain what a study is actually able to do with its own data later, an agreement that didn't anticipate a specific secondary analysis, or that expires before a planned follow-up study, can leave genuinely valuable data unusable for research it was collected to support. Treating the data sharing agreement as a document worth real attention at the design stage, rather than paperwork to clear before the interesting work starts, protects the value of the data long after the agreement itself is signed.