Open Lab

Data retention and the right to erasure in practice

How long study data may be kept, when a participant can have it deleted, and where the research exceptions in the GDPR apply. With the retention and deletion options in Open Lab, and what to write in an ethics application.

Open Lab Team
5 min read

Ethics and data-protection forms for online studies commonly ask two questions: how long will the data be kept, and what happens if a participant asks for their data to be deleted. Both have legal answers in the EU General Data Protection Regulation (GDPR), and both have practical answers that depend on how the study was set up. This post covers the relevant provisions, how they apply to a typical browser-based study, and the retention and deletion options available in Open Lab.

The post describes the regulation and the platform. It is not legal advice, and institutions differ in how they interpret the research provisions, so the institution's data protection officer (DPO) remains the person to confirm the final wording of an application.

Who is responsible for what

For the data collected in a study, the researcher or their institution is the controller: the party that decides why and how the data are processed. Open Lab is a processor, which means it stores and processes the data on the researcher's behalf and under their instructions. The terms of that arrangement are set out in Open Lab's data processing agreement (DPA), available at research.open-lab.online/dpa.

This division matters for erasure requests. A request to delete study responses is addressed to the controller, so it is the researcher who decides whether and how to act on it. Open Lab provides the tools to carry the decision out, and participants who have an Open Lab account can also delete their own data directly, as described below.

What the GDPR says about retention

The relevant principle is storage limitation. Article 5(1)(e) requires that personal data be "kept in a form which permits identification of data subjects for no longer than is necessary for the purposes for which the personal data are processed." The same article allows longer storage when the data are processed solely for scientific research, provided that the safeguards required by Article 89(1) are in place. Article 89(1) names data minimisation as the aim of those safeguards and mentions pseudonymisation as one possible measure.

Two points follow from the wording. First, the principle concerns data in a form that permits identification. Recital 26 states that data protection principles do not apply to anonymous information, but that pseudonymised data, which could be attributed to a person with additional information, are still personal data. A dataset in which names have been replaced by codes, with the key held elsewhere, is therefore still covered. Second, the research exception extends the permitted storage period but does not remove the need for one. An application still has to state a period and a reason for it.

Research funders add a requirement in the opposite direction. In Germany, Guideline 17 of the DFG Code of Conduct on good research practice states that research data are generally archived for ten years, counted from the date the results are made publicly available, and that shorter periods may be appropriate in justified cases.

These two requirements are usually reconciled by treating the identifying parts of a dataset separately from the rest. The data needed to reproduce a published analysis are archived for the required period in a form that no longer identifies anyone. Information that identifies participants, such as contact details, panel IDs used for payment, or the key linking codes to people, is kept only as long as the study needs it and then deleted.

Identifying information in an online study

In a browser-based study, identifying information can arrive without the researcher asking for it. Recruitment panels pass participant IDs in the study link, for example the PROLIFIC_PID parameter from Prolific or the survey_code from Sona. These IDs do not contain a name, but the panel can link them to a person, so under Recital 26 they are pseudonymous rather than anonymous. Free-text answers, email addresses entered for a prize draw, and technical information about the device can also contribute to identifiability.

In practice, this means a retention plan usually has two parts. The first part names the identifying fields and states when they will be removed, for example after payment has been made and data collection has closed. The second part states how long the remaining, de-identified data will be kept and where they will be archived. A dataset exported for archiving should have the identifying columns removed before it is deposited.

Setting a retention period for a study in Open Lab

Each Open Lab study has a Data retention tab under Settings. It controls how long participant responses are kept on Open Lab before they are deleted automatically. There are three options:

  • Use plan default. Responses are kept for 12 months on the free plan and 36 months on paid plans.
  • Custom period. Responses are kept for a number of months chosen by the researcher. On the free plan the maximum is 12 months.
  • Keep indefinitely. Available on paid plans only.

The period counts from the moment each response is submitted, so responses collected on different days are deleted on different days. It applies both to completed sessions and to the partial data saved from sessions that were not completed. A change to the setting applies to responses collected after the change. Responses that were already submitted keep the deletion date they were assigned at the time.

When a response reaches its deletion date, it goes through several steps. The study author receives an email 30 days in advance. On the deletion date, the data are removed from the researcher's view. Thirty days after that, the data and the stored file behind them are permanently deleted. Open Lab's database backups are kept on a rolling seven-day window, so a deleted record may remain in a backup for up to a week before that backup is overwritten.

Retention on Open Lab covers the platform's copy of the data only. Archiving for the ten-year period described above is the researcher's responsibility and normally takes place in an institutional or subject repository. Data needed for archiving should therefore be exported, and de-identified where appropriate, before the retention period on Open Lab ends.

When the right to erasure applies to research data

Article 17 of the GDPR gives a data subject the right to have their personal data erased in several situations. Two are common in research. The first is that the participant withdraws the consent on which the processing was based. Article 7(3) allows consent to be withdrawn at any time and requires withdrawal to be as easy as giving consent, while also stating that withdrawal does not affect the lawfulness of processing before it. The second is that the data are no longer necessary for the purpose for which they were collected.

Article 17(3)(d) sets out an exception for scientific research. The right to erasure does not apply "in so far as the right referred to in paragraph 1 is likely to render impossible or seriously impair the achievement of the objectives of that processing", and only where the Article 89(1) safeguards are in place. A researcher who relies on it would need to explain why deleting one participant's data would seriously impair the study. During data collection, removing one participant's responses may have little effect on the study's objectives, so the exception may be harder to justify at that stage than after the data have been analysed or published.

One way to handle this in a consent form is to state a point up to which participants can have their data deleted, such as the date on which the identifying information will be removed. Before that point, a participant can be found in the dataset and their data can be deleted. After that point, the data can no longer be attributed to the participant.

Finding one participant's data

A deletion request can only be carried out if the participant's data can be located. Article 11 of the GDPR covers the case where they cannot. If the controller does not need to identify participants for the purposes of the study, it is not required to collect extra information only to comply with the regulation. When the controller can show that it is unable to identify the person, the rights of access and erasure do not apply, unless the participant provides additional information that makes identification possible.

Studies in which participants take part anonymously can fall into this situation. A participant who wants their data removed later has no way to point to their own records unless they were given something to quote. One option is to show each participant a code at the end of the study and to mention it in the debriefing text, together with an explanation of how to request deletion.

In Open Lab, the study's Data view can be filtered by participant code or by the stable ID that Open Lab assigns to each session. Link parameters such as PROLIFIC_PID can be shown as a column in the same view, which makes it possible to find a panel participant's session when they provide their panel ID. Once the session is found, it can be deleted from that view.

How deletion works in Open Lab

Deleting sessions as a researcher. In a study's Data view, individual sessions or a selection of sessions can be deleted. Deletion removes both the database record and the stored data file. The study author and collaborators with the Owner, Editor or Data Analyst role can delete data. Viewers and Participant Managers cannot. Deletion does not require the decryption key, so it works for all three encryption options, including end-to-end encrypted studies whose content Open Lab cannot read.

Deleting a study. Deleting a whole study, which only the study author or a project owner can do, also deletes the study's responses, the stored data files, the participant records connected to it, its invitations (which contain the invited participants' email addresses), the messages exchanged with participants about the study, and the notifications participants received about it.

Withdrawal by participants with an account. Participants who took part through an Open Lab participant account can withdraw from a study from the study card in My studies. When they withdraw, they are asked whether the data already collected in that study should also be deleted.

Account deletion by participants. Participants can delete their whole account from the Privacy section of their dashboard. As part of this, they choose whether the data collected in their studies is deleted too or kept for the researchers. If they choose to keep it, the data are no longer connected to their account or email address. Other information recorded with the session, such as a panel ID passed in the study link, remains in the dataset, so the kept data may still be pseudonymous rather than anonymous, and the researcher remains the controller for it. The participant receives an email confirming what was deleted and what was kept. Participants who took part without signing in are not linked to an account, so their data are not affected by account deletion and requests from them are handled by locating the session as described above.

Copies outside the platform

Deletion on Open Lab removes the data stored on Open Lab. It does not reach copies that have already been exported, whether on a researcher's computer, a shared drive, a collaborator's laptop or an analysis server. Open Lab's participant interface states this explicitly and advises participants to contact the researcher before deleting their account, because a deleted account can no longer send messages through the platform.

For this reason, a retention plan should list where exported copies are stored and who holds them, so that a deletion request can be carried out in each place. The same list is useful at the end of the retention period, when all working copies of identifying data should be deleted and only the archived, de-identified dataset kept.

What to write in the ethics application

A data-protection section that covers retention and erasure usually needs to state the following:

  • The controller and the processor. The institution or researcher as controller, Open Lab as processor under its DPA, with hosting in Frankfurt, Germany.
  • Identifying fields. Which identifiers are collected, including those passed by recruitment panels, and the date or event after which they will be removed.
  • Retention on the platform. The retention period set in the study's Data retention tab and what happens when it ends.
  • Archiving. Where the de-identified dataset will be archived, for how long, and in what form.
  • How participants request deletion. Who to contact, what information to provide (for example the participant code shown at the end of the study), and the time frame for a reply. Article 12(3) requires a response without undue delay and within one month of receipt, extendable by two further months for complex or numerous requests.
  • The cut-off for deletion. The point after which data can no longer be attributed to a participant, and, if relevant, whether the research exception in Article 17(3)(d) will be relied on after that point and why.
  • Exported copies. Where copies of the data will be kept and how they will be deleted at the end of the retention period or on request.

Which parties can read the stored data depends on the study's encryption option, which is covered in an earlier post on the three encryption options in Open Lab.

Share
All posts →