How to use this playbook
It collects the documents that together make a dataset responsibly shareable. Most have public templates you can adapt, which means you can arrive at your lawyer with drafts rather than a blank page, and that changes both the cost and the timeline.
Sample content only
Review and customize every section, especially governance, IRB, consent, and legal terms. This is not legal or regulatory advice.
The front door
A public page explaining how someone requests data from you and what happens next. Most groups do not have one, which means every request arrives as an email to whoever is most visible, and gets answered differently each time.
What it needs to say: what the dataset contains at a high level, who may request it, what the process and rough timeline are, what agreements are required, and who to contact. One page. It does more for your credibility than a longer document would.
Data use agreement
The terms a requester signs. At minimum it should address permitted uses, prohibited uses, whether re-identification attempts are forbidden, what happens to the data at project end, publication and acknowledgement expectations, and who owns derived work.
The clause groups forget
Ownership of derived outputs. A requester who builds a model on your data and publishes it has created something. Decide in advance what you expect, because renegotiating afterward is a relationship problem rather than a legal one.
Data access committee
Three volunteers and a written process is a functioning committee. What makes it work is not the seniority of the members but the existence of a rubric.
The SOP needs
- Who sits on it, and how a member is replaced
- How often it meets, and what happens to a request between meetings
- What a complete request contains
- Who may be recused, and when
- How a decision is communicated and whether it can be appealed
The rubric needs
- Scientific merit, is the question answerable with this data
- Community benefit: will the answer help the people who contributed
- Feasibility, does the requester have the capacity
- Governance fit, does the request fall within consent
- Burden, what does approving this cost you in staff time
Score each, write down the threshold, and apply it consistently. That is what makes a refusal defensible.
IRB and ethics documentation
Your protocol, your consent, and the analysis of whether that consent permits what you now want to do. Publish enough that a requester can tell whether their intended use is in scope before they write to you.
Answer these three publicly: is the registry itself under IRB oversight or determined not to be human subjects research; does participant consent permit sharing with third parties; and does it permit re-contact.
Data dictionary and derived fields
Could an outsider understand your dataset from the documentation alone? If not, it is not shareable yet, however much of it you have.
A usable entry has the field name, a plain-language definition, the type and permitted values, whether it is required, when it started being collected, and any known quality caveat. Derived fields additionally need the calculation and its inputs.
Undocumented derived fields end reuse
Undocumented derived variables are the most common reason a dataset cannot be reused. A downstream analyst who cannot reproduce a computed field has to either trust it blindly or discard it, and most discard it.
Readiness assessment
An honest inventory before you publish anything. This is the document that stops you announcing a data sharing portal you cannot staff.
| Ask | If the answer is no |
|---|---|
| Is there a published data dictionary? | Stop. This comes first. |
| Does consent permit this sharing? | Stop. Talk to your IRB. |
| Is there a named person who answers requests? | Stop. An unanswered request is worse than no portal. |
| Is there a written access process? | Write it. A week's work. |
| Do you know your dataset's quality profile? | Assess it before someone else does. |
| Can you support a requester for the length of their project? | Narrow what you offer rather than over-promising. |
Sharing with industry sponsors
Most of the framework is the same. Four things change.
- Consent language. Some consent permits academic sharing but not commercial. Check before the conversation, not during.
- Value. If your data materially advances a commercial program, that has value. Deciding what you want in return (funding, a seat, access to results, a commitment on eventual access, is a governance decision to make before you are asked.
- Publication. Negotiate review windows and the right to publish regardless of findings.
- Community transparency. Decide in advance what you will tell your community about commercial sharing, and tell them.
Depositing with an external repository
A repository gives you durable hosting, a discovery layer, citability, and access governance machinery you would otherwise build. It asks for a documented dataset, consent that permits deposit, and a named responsible person.
Ask any candidate repository what they would require, even if you are years away. Their answer is a free specification for your data dictionary.
FAIR without a budget
Findable, accessible, interoperable, reusable. The common misunderstanding is that FAIR means open. It does not, a controlled-access dataset can be fully FAIR.
The cheapest meaningful step for most groups is publishing discoverable metadata while the raw data stays controlled. A researcher can then find that your dataset exists, understand what it contains, and know how to request it. That is most of the benefit, at almost none of the cost or risk.
Where this sits in the maturity model
This playbook is how a group moves from stage 2 (internal governance) to stage 3 (documented and standardized) and then to stage 4 (shared with a partner) on the governance track, the north star track that caps every other one.