How do I write a data management plan?

Answer six questions: what data you will produce, how it will be documented, where it will be stored and backed up, how it will be protected, what will be shared and where, and who is responsible. Most funder templates are a rearrangement of those six.

Updated

The plan is usually written the week a grant is due, by someone who has not yet lost any data, which is why so many read as compliance rather than as a working document. The version worth writing is the one you would actually follow, and that turns out to be shorter and more specific than the compliance version.

Start with what the data physically is. Interview recordings and transcripts, survey exports, instrument output, code, images, field notes. For each, the format, the approximate volume, and whether it contains anything identifying. That inventory takes twenty minutes and it determines every other answer in the plan.

Write the inventory at the level of detail a reviewer can check. The NIH guidance asks for each dataset's modality, level of aggregation, level of processing, approximate format, and how it will be produced, and it asks for this before the project begins (Gonzales et al., 2022, pp. 2-3). That maps onto a thesis as separate lines for interview audio, transcripts, consent records, coded qualitative files, survey exports, scripts, and derived tables, with each one marked raw, cleaned, anonymised, or analysis-ready. The plan is a description of the whole data life cycle, from collection through organisation, quality control, documentation, preservation, and sharing, and it is peer-reviewed alongside the proposal (Michener, 2015, p. 1).

Documentation is the section most often left thin and the one that matters most three years later. What will let someone else, including future you, understand the files: a readme, a codebook, variable definitions, the version of the instrument used. A dataset nobody can interpret is not preserved in any useful sense.

Name the standards, not just the intention to document. The relevant ones are file formats, data dictionaries, data identifiers, unique identifiers, variable definitions, and the supporting documentation itself (Gonzales et al., 2022, p. 4). Choose standardised and preferably open formats where the analysis allows it, so a CSV export rather than a proprietary Excel workbook, and state the format of the data dictionary as well as its content (Gonzales et al., 2022, p. 5).

Storage needs three copies in two places, with one off site, and it needs to be somewhere your institution supports rather than a personal cloud account. Say who checks the backups and how often, because a backup nobody has ever restored from is a hope rather than a plan.

The standard recommendation is at least three copies in at least two geographically distributed locations, with regular duplication and tests that show the files can actually be retrieved (Michener, 2015, p. 5). Pair that with the error side of the same section. The listed quality assurance and quality control measures include researcher training, instrument calibration, verification tests, double-blind data entry, and statistical or visual checks such as scatterplots and maps for finding anomalies, and which of them apply depends on the study and the sponsor (Michener, 2015, p. 5). A twelve-interview project should not promise instrument calibration, and a survey project should say which coding and entry checks it will run.

Then sharing. What will be made available, where, under what licence, and when. If your consent forms do not permit sharing, that is the answer and you should say so plainly, along with what you will share instead, such as metadata or a synthetic dataset.

What do funders actually check?

Whether the plan is specific to your project rather than generic, whether the storage and backup arrangements exist, whether the sharing plan is compatible with your ethics approval and consent, and whether someone is named as responsible. Vagueness is what gets flagged, not ambition.

The consent question is the one that catches people. A plan promising open data alongside a consent form that promised participants their data would not be shared is an internal contradiction, and reviewers who read both notice it.

Name the repository rather than saying data will be deposited in an appropriate repository. Naming it shows you checked that one exists and accepts your data type.

How long does a data management plan need to be?

Usually one to three pages, and the template decides. Length is not the quality signal. A two-page plan that names the repository, the file formats, the retention period, and the person responsible is stronger than six pages of general commitment to good practice.

There is no universal template, because professional societies, funders, and institutions each set their own requirements (Hassenstein and Jung, 2025, p. 7). Some are very specific. The NIH policy, effective January 2023, requires a one to two page Data Management and Sharing Plan for every funded project (Gonzales et al., 2022, pp. 1-2). Read the requirement your funder publishes before you write, and if none applies, borrow a published one rather than inventing headings.

Write it so a new team member could follow it. That is the test that distinguishes a plan from a statement of intent, and it is also what makes it worth keeping once the grant is awarded.

Revisit it when the project changes. A plan written for a study that became something else is worse than useless, and updating it takes ten minutes at a point where you already know what changed.