Open Lab

A researcher's guide to GDPR-compliant online data collection

An overview of the GDPR provisions that apply to a browser-based study: who is responsible for the data, which legal basis to use, what participants must be told, what an online study records without being asked, where the data are stored, and what to put in an ethics application.

Open Lab Team
5 min read

Running a study in the browser changes the data-protection picture in a few specific ways. Data pass through a platform rather than staying on a lab computer, recruitment panels add identifiers to the study link, and the browser reports technical information about the participant's device whether or not the researcher asked for it. The EU General Data Protection Regulation (GDPR) applies to all of this as soon as the data relate to an identifiable person.

This post gives an overview of the provisions that matter for a typical online study and shows how they apply on Open Lab. Two topics have their own posts: the three encryption options for online studies, and data retention and the right to erasure in practice. The post describes the regulation and the platform. It is not legal advice, and institutions differ in how they apply the research provisions, so the institution's data protection officer (DPO) remains the person to confirm the final wording of an application.

Who is responsible for the data

The GDPR distinguishes between a controller, which decides why and how personal data are processed, and a processor, which processes data on the controller's behalf. In a study run on Open Lab, the researcher or their institution is the controller and Open Lab is the processor.

Article 28 requires the relationship between the two to be governed by a contract that sets out, among other things, the subject matter, duration, type of data and the processor's obligations. Open Lab's data processing agreement (DPA) is published at research.open-lab.online/dpa. It is incorporated into the terms of service, so it applies once a researcher accepts the terms at signup or uses the platform to collect study data. Institutions that need a countersigned DPA, or have their own clauses, can request one at info@open-lab.online.

Many institutions also keep a record of processing activities under Article 30, and some require each study to be entered there. The DPO's office can say whether an online study needs its own entry.

Every processing operation needs one of the legal bases in Article 6(1). Two are common in academic research:

  • Consent (Article 6(1)(a)): the participant "has given consent to the processing of his or her personal data for one or more specific purposes".
  • Public interest (Article 6(1)(e)): processing "is necessary for the performance of a task carried out in the public interest or in the exercise of official authority vested in the controller". Public universities often rely on this basis for research, with the research task defined in national or state law.

Research ethics also requires informed consent, and the two are easy to confuse. The European Data Protection Board (EDPB) addressed this in a 2021 document on health research: consent to take part in a study, which ethics standards require, is a different concept from consent as a legal basis for processing personal data. A study can therefore collect ethical informed consent from every participant while relying on Article 6(1)(e) as its legal basis under the GDPR. Where consent is the legal basis, withdrawing it ends the basis for further processing of that participant's data (Article 7(3)). Where public interest is the legal basis, withdrawal is handled under the right to object and the research exceptions instead.

Which basis applies is usually decided by the institution rather than by the individual researcher. The information given to participants should name the basis the institution has chosen.

Special categories of data

Article 9(1) prohibits processing of data "revealing racial or ethnic origin, political opinions, religious or philosophical beliefs, or trade union membership", genetic and biometric data, "data concerning health" and data concerning sex life or sexual orientation, unless an exception applies. Many psychological studies touch these categories: a depression screening questionnaire produces health data, and a political attitudes survey can reveal political opinions.

Two exceptions are relevant to research. Article 9(2)(a) permits processing with the participant's explicit consent. Article 9(2)(j) permits processing for scientific research purposes on the basis of Union or national law, with safeguards in line with Article 89(1). In Germany, for example, Section 27 of the Federal Data Protection Act (BDSG) provides such a basis. An application that involves special categories should identify which exception it uses.

What participants must be told

Article 13 lists the information to give participants when data are collected from them. In summary, it covers:

  • the identity and contact details of the controller, and of the DPO where there is one;
  • the purposes of the processing and its legal basis;
  • the recipients or categories of recipients of the data, including processors;
  • any transfer to a country outside the EU or EEA, and the safeguard used;
  • how long the data will be stored, or the criteria for deciding;
  • the participant's rights (access, rectification, erasure, restriction, portability, objection), the right to withdraw consent where consent is the basis, and the right to complain to a supervisory authority;
  • whether providing the data is required, and what happens if it is not provided.

In an online study this information is usually combined with the ethics consent text on the first screen. On Open Lab, a study can include a consent component in the study builder. Its editor has separate fields for the nature, purpose and duration of the study, procedures, risks and benefits, confidentiality and data handling, the point of contact, the withdrawal process, and the ethics approval number and institution. A participant cannot continue past the consent screen without ticking the agreement box.

Article 7(1) requires a controller relying on consent to be able to demonstrate that consent was given. When a participant agrees, Open Lab records the time of agreement, which researchers can see on that participant's page in the study's Participants view. Open Lab also stores a fingerprint (a SHA-256 hash) of the consent text that was shown, so the version a participant agreed to can be identified later, even if the consent text is edited during the study.

Data minimisation and what an online study records

Article 5(1)(c) requires personal data to be "adequate, relevant and limited to what is necessary in relation to the purposes for which they are processed". In an online study, some data arrive without being requested, so meeting this principle starts with knowing what the platform records. The following information is stored with each Open Lab session, in addition to the participant's responses.

Technical information about the device. The browser reports its user agent (browser and operating system), language and time zone, screen size and pixel density, a rough indication of the device's processor and memory, and basic network information such as connection type. It is recorded by default and is used to describe the sample and to screen data quality. It can be switched off for a study in Settings → Data retention, which makes some quality signals, such as duplicate-device detection, less informative.

A pseudonymised IP address. Open Lab does not store participants' IP addresses. It stores a salted hash: the address is combined with a secret value and passed through a one-way function, so the stored value cannot be reversed to the address but is the same for repeat visits from the same address. It is used to detect several sessions from one connection within a study.

Link parameters. Recruitment panels add identifiers to the study link, such as PROLIFIC_PID from Prolific or survey_code from Sona. Open Lab saves all parameters in the study link with the session, so that the researcher can match sessions to panel records for payment and screening.

Paradata, if enabled. Paradata are records of how a participant interacted with the page rather than what they answered, such as counts of clicks, key presses, copy and paste events, and switches away from the browser tab. On Open Lab, paradata collection is switched off by default and can be enabled per study in the Bot Detection tab of the study's Quality dashboard.

Two points follow for a data-protection section. First, technical information and panel IDs should be listed among the data collected, even though the study's own tasks do not ask for them. Second, these items affect whether the dataset is anonymous. Recital 26 states that pseudonymised data, which could be attributed to a person using additional information, remain personal data. A panel ID does not contain a name, but the panel can link it to a person, so a dataset that contains panel IDs is pseudonymous rather than anonymous. Free-text answers, email addresses entered for a prize draw, and detailed demographic combinations can have the same effect.

Minimisation also applies to what the study itself asks. Age in years can often be replaced by an age band, and a free-text field for occupation by a short list of categories. Questions whose answers are not part of the analysis plan are candidates for removal.

Purpose limitation and later use of the data

Article 5(1)(b) requires personal data to be "collected for specified, explicit and legitimate purposes and not further processed in a manner that is incompatible with those purposes". The same provision states that further processing for scientific research, in accordance with Article 89(1), is not considered incompatible with the original purposes.

Recital 33 acknowledges that "it is often not possible to fully identify the purpose of personal data processing for scientific research purposes at the time of data collection", and allows consent to be given for areas of research rather than one specific project. This form is often called broad consent. The EDPB published draft Guidelines 1/2026 on processing personal data for scientific research in April 2026, with a public consultation that closed on 25 June 2026. The draft confirms that broad consent can be used where research purposes are not fully known at collection, provided that researchers respect ethical standards and add safeguards to compensate for the less specific purpose.

For an online study, the practical question is usually reuse and sharing. If the de-identified data are to be deposited in an open repository, reused in a later project or shared with other teams, the information given to participants should say so at the time of collection.

Data protection impact assessments

Article 35(1) requires a data protection impact assessment (DPIA) before processing that is "likely to result in a high risk to the rights and freedoms of natural persons". Article 35(3) names three cases where one is always required. The case most relevant to research is "processing on a large scale of special categories of data". Each national supervisory authority also publishes a list of processing operations that require a DPIA (Article 35(4)), and these lists differ between countries.

A short anonymous questionnaire on a non-sensitive topic is unlikely to require a DPIA. A large study that collects health data, or one that combines sensitive data with identifiers, may require one. The decision rests with the controller and is usually made together with the DPO. Where a DPIA is needed, the processor is required to assist with it, and Open Lab's DPA includes this obligation.

Where the data are stored and who else handles them

Open Lab stores study data on DigitalOcean servers in Frankfurt, Germany. Open Lab's sub-processors, the companies that handle some data on its behalf, are listed with their purpose, location and transfer safeguard at research.open-lab.online/subprocessors. For participant data, the relevant ones are the hosting provider, the email provider used for messages to participants with accounts, and the payout provider, which receives a participant's name and email address only when a researcher pays participants through Open Lab. Transfers to sub-processors in the United States rely on the EU–US Data Privacy Framework and/or the European Commission's Standard Contractual Clauses.

Any transfer outside the EU or EEA has to be mentioned in the information for participants (Article 13(1)(f)). Two other transfers often arise in an online study and are separate from the platform:

  • Recruitment panels. A panel such as Prolific holds its own records about its participants under its own privacy terms. Prolific is based in the United Kingdom. The European Commission renewed its adequacy decisions for the UK in December 2025, so personal data can continue to flow from the EEA to the UK without additional safeguards.
  • Analysis tools and storage. Data exported for analysis and stored on a cloud drive, or processed by an external service, are covered by that service's terms and location. These copies should be included in the data-protection section as well.

Security of processing

Article 32 requires technical and organisational measures appropriate to the risk, and names encryption and pseudonymisation as examples. On Open Lab, the main choice for a study is its encryption option, set in Settings → Encryption. With the end-to-end options, responses are encrypted before they reach the server and only key holders can read them. With server-side encryption, data are encrypted at rest but Open Lab manages the keys. A new study uses end-to-end encryption with the account key when the researcher has set up an encryption vault in Settings → Security and confirmed that the recovery code has been saved. Without that confirmation, the study uses server-side encryption. The setting is therefore worth checking before data collection starts. The differences between the options, and what to write about each, are covered in the encryption post.

Access within the research team is a second measure. Each Open Lab project role grants a defined level of access, and only some roles can download or delete data, as described in the post on project roles. A data-protection section can state which team members hold which role.

Retention and deletion

Article 5(1)(e) limits storage of identifiable data to the period that is necessary, with longer storage permitted for scientific research under the Article 89(1) safeguards. Each Open Lab study has a Data retention tab under Settings that sets how long responses are kept on the platform before automatic deletion: 12 months on the free plan and 36 months on paid plans by default, with a custom period available. Archiving for the period required by a funder, such as the ten years set out in the DFG's Code of Conduct, normally takes place in an institutional or subject repository after identifying information has been removed. The post on data retention and the right to erasure covers this, along with how to handle a participant's request to delete their data.

A checklist for the ethics application or DPIA

A data-protection section for an online study usually needs to cover the following points:

  • Controller and processor. The institution or researcher as controller; Open Lab as processor under its DPA, with hosting in Frankfurt, Germany.
  • Legal basis. The Article 6(1) basis chosen by the institution, and, for special categories, the Article 9(2) exception. Ethical informed consent is described separately if the legal basis is not consent.
  • Data collected. The study's own measures, plus technical device information (unless switched off for the study), the pseudonymised IP address, link parameters from the recruitment panel, and paradata if enabled.
  • Identifiability. Whether the dataset is anonymous or pseudonymous, and which fields make it so.
  • Information to participants. Where the Article 13 information is given, for example in the study's consent component, and how consent is recorded.
  • Further use. Whether the de-identified data will be shared or reused, and how participants were told.
  • Recipients and transfers. The platform's sub-processors, the recruitment panel, and any analysis or storage services, with their locations and transfer safeguards.
  • Security. The study's encryption option and who in the team has access to the data.
  • Retention and deletion. The retention period on the platform, when identifiers will be removed, where the de-identified data will be archived, and how participants can request deletion.
  • DPIA. Whether one is required, with reference to the national supervisory authority's list.
Share
All posts →