In this blog post, you can find the presentation I gave on June 17, 2026, at the Computational Communication Science Lab at the University of Vienna. The presentation outlines the project conducted with the European Commission’s European Centre for Algorithmic Transparency (ECAT) and draws five lessons from this experience for researchers preparing an application under Article 40(12) of the Digital Services Act (DSA).
Disclaimer: The views expressed in this presentation are my own and do not necessarily reflect the views of the European Commission, DG CONNECT, the Joint Research Centre, ECAT, or any participating platform. The presentation discusses general experiences and lessons learned from the project. It does not disclose confidential information and should not be interpreted as representing official findings, positions, or assessments of any institution involved in the project.

Good afternoon, everyone, and thank you all for inviting me. I’m pleased to be back here to contribute to the research community at University of Vienna. I’m here to share my recent research experience participating in a project coordinated with European Commission teams that explored how the DSA’s data-access provisions operate in practice. The DSA is one of the most significant European laws for social science researchers in recent years, as it establishes the conditions under which researchers can access data from major digital platforms. In this presentation, which should last about half an hour, I’d like to share some lessons I learned during this year-long collaboration. Specifically, I have selected five lessons that may be useful as you navigate data-access requests under the DSA.

The research project I am going to talk about is titled “Systemic Risks of Very Large Online Platforms with Publicly Available Data”, and last Tuesday, we presented the final report—one of the three main deliverables—at a meeting held at the headquarters of the Directorate-General for Communications Networks, Content and Technology (DG CONNECT) in Brussels.

The goal of the project was to generate evidence that could inform the implementation of the DSA by exploring how the data-access mechanism established under Article 40(12) operates in practice.
To do so, the European Commission needed researchers to work in practice under the DSA framework and submit actual data-access requests. This made it possible to observe how the data-access mechanism operated in practice, and to what extent the data provided met researchers’ needs for conducting meaningful research. The project therefore offered a practical opportunity to assess the effectiveness of the data-access mechanism from the researchers’ perspective.

This slide shows the main actors involved in the project and how they interacted throughout the process.
At the top is the European Commission, which initiated and funded the project. Within the Commission, two bodies played a central role. The first was the Joint Research Centre, or JRC, which provides scientific and technical support to EU policymaking. The second was DG CONNECT, the Directorate-General responsible for digital policy.
The operational coordination of the project was carried out through the European Centre for Algorithmic Transparency, or ECAT, which brought together researchers and acted as the bridge between the research teams and the Commission.
Researchers were then tasked with designing research projects aimed at assessing systemic risks, submitting actual data-access requests under Article 40(12) DSA to a range of Very Large Online Platforms, and conducting their research using the data provided. The idea was not simply to study the law in theory, but to test how the data-access mechanism works in practice.
The project covered different types of platforms. These included social media platforms, AI chatbot services, online marketplaces, and adult-content platforms. Our role in the project focused on online marketplaces.

The research process that we followed throughout the project consisted of several stages.
The first step, as mentioned, was to design a research project addressing a specific systemic risk covered by the DSA.
The second step was to translate our research needs into formal data-access requests under Article 40(12), once the research questions had been defined. This required identifying which data would be necessary to answer the research questions and explaining why access to those data was needed.
The third phase consisted of interacting with the platforms. This often involved a substantial amount of communication, clarification requests, and negotiations regarding the scope of the data to be provided. Over the course of the project, and as of today, we have exchanged several emails with the five platforms we worked with.
Fourth, after receiving the data, we assessed its quality and evaluated the extent to which it matched both our original requests and our actual research needs. This was a crucial step because obtaining data is not, by itself, sufficient; the data must also be usable and relevant to the intended analysis.
Finally, we conducted the research itself. In many cases, however, the data received differed from what had originally been requested. As a result, the research design often had to be adapted to the information that was actually available. This was a very important point for the Commission, because if the DSA is to work in practice, researchers must receive the data they need to answer systemic-risk research questions.
Throughout this entire process, the work was conducted under the oversight of DG CONNECT and ECAT, which coordinated the project and monitored its progress. The outcomes and documentation gathered during the process may inform the understanding of how the access mechanism works in practice.

Before moving on, I think it is useful to provide a bit of background on the DSA. You have probably all heard of it, and some of you may even have already tried to access data under its provisions, but a brief overview will not hurt.

The Digital Services Act, or DSA, establishes a common regulatory framework for online platforms operating in the European Union. While the Regulation applies to a wide range of digital services, it imposes additional obligations on Very Large Online Platforms and Very Large Online Search Engines—commonly referred to as VLOPs and VLOSEs. These are platforms and search engines with more than 45 million active recipients in the EU, roughly 10% of the Union’s population.
The DSA rationale is straightforward: because of their scale, very large platforms can generate societal effects that extend far beyond individual users. For this reason, the DSA requires them to identify, assess, and mitigate what the Regulation calls “systemic risks”.
These systemic risks fall into four broad categories. The first concerns the dissemination of illegal content. The second relates to risks affecting fundamental rights, such as privacy, non-discrimination, or freedom of expression. The third covers risks to civic discourse, electoral processes, and public security. Finally, the fourth category includes risks related to health, safety, the protection of minors, gender-based violence, and the physical or mental well-being of individuals.
To enable independent scrutiny of these issues, the DSA also includes data-access provisions. Researchers can access public data under Article 40(12), while Article 40(4) provides a mechanism for access to certain non-public data under stricter conditions. Our project focused on Article 40(12), which is the provision that allows vetted researchers to request access to publicly available platform data for the study of systemic risks.

For our purposes, the key provision is Article 40(12) DSA. The article establishes that researchers can obtain access to publicly accessible data held by Very Large Online Platforms when the research contributes to the detection, identification, or understanding of systemic risks.
The concept of public data is important here. The DSA 40(12) does not primarily concern private user information or confidential business data. Rather, it focuses on data that are already visible through the platform’s interface and available to users, but that may not be easily accessible for systematic and large-scale research.
The Regulation also specifies that access should be provided without undue delay and, where technically possible, may include real-time data.
In short, Article 40(12) seeks to transform publicly visible platform information into a resource that can be used for independent scientific scrutiny of the societal risks generated by large digital platforms.
Access is limited to eligible researchers who satisfy a number of requirements and who use the data exclusively for research on systemic risks.

In fact, not every researcher can simply request data from a platform. Article 40 establishes a number of conditions that applicants must satisfy.
First, researchers must be independent from commercial interests. The idea is that the mechanism is intended to support independent scientific scrutiny rather than commercial exploitation of platform data.
Second, researchers must disclose the funding sources of their research. This allows platforms and regulators to assess potential conflicts of interest and promotes transparency regarding who is supporting the project.
Third, applicants must demonstrate that they are capable of handling the data securely and confidentially. This includes describing the technical and organizational measures they have put in place to protect personal data and ensure compliance with data-protection requirements.
Finally, researchers must justify their request. They need to demonstrate that access to the requested data is necessary and proportionate to their research objectives, and that the expected results will contribute to the understanding of systemic risks covered by the DSA.
In other words, access is not granted simply because a researcher is interested in a topic. The request must be independent, transparent, secure, and clearly linked to the study of systemic risks. The researcher must demonstrate that these conditions are met in order to obtain access to the requested data.

In this presentation, I will focus exclusively on Article 40(12) DSA, which I had the opportunity to experience directly and observe first-hand through my involvement in this project. I will leave aside Article 40(4) for now, as it concerns a different access mechanism and was not the focus of my work.
Just to briefly mention some of the main differences between the two provisions: Article 40(4) does not involve direct interaction between researchers and platforms; researchers must be affiliated with a research institution; and the data that can be accessed include not only publicly visible information but also certain non-public data.
There are also differences in the scope of the research. While Article 40(12) focuses on research aimed at understanding systemic risks, Article 40(4) additionally covers research assessing the adequacy, effectiveness, and impacts of the platforms’ risk-mitigation measures.

I will now discuss some of the key lessons that emerged from our experience with Article 40(12) DSA. The first concerns the design of data-access requests.

This first lesson emerged very early in the project, during the design of the data-access requests.
One thing we quickly learned is that it is not enough to have an interesting research question. Under Article 40(12), every request must be clearly connected to a systemic risk recognized by the DSA. In practice, this means that you need to explain not only what you want to study, but also why answering that question contributes to the detection, identification, or understanding of a systemic risk.
This is important because your requests are typically reviewed by the platforms’ legal and compliance teams. If the connection between your research question and the systemic risk is weak or ambiguous, platforms have more room to challenge the request or argue that the requested data fall outside the scope of Article 40(12).
For this reason, our approach was to explicitly anchor every research question to the relevant DSA provisions, recitals, and, where available, European Commission guidance. The example shown on the right comes from one of our requests. Before presenting the research question itself, we explained why the issue could constitute a systemic risk under the DSA and identified the specific legal provisions supporting that interpretation.
The practical lesson is therefore quite simple: do not assume that the systemic-risk relevance of your research question is self-evident. Make it explicit, support it with official sources, and clearly demonstrate how the requested data are necessary to investigate that risk. This can substantially strengthen your request and reduce disputes about its scope.

A useful way to strengthen your request is to link it to European Commission documentation and other official sources. These documents can help demonstrate that the issue you are studying is relevant from a systemic-risk perspective and falls within the scope of the DSA.
For example, reports and guidance published by the European Commission can provide useful evidence connecting specific research questions to the categories of systemic risks. Referring to these sources can make the rationale for your request clearer and more difficult to challenge.

The second lesson emerged from a relevant discussions we had both with experts from the European Commission and with the platforms themselves: how much data is actually necessary for research.
A frequent argument you may also hear is that researchers should request only the smallest possible amount of data, often in the form of platform-generated samples. This is presented as a straightforward application of the GDPR’s data-minimization principle. However, this interpretation is not entirely correct.
It is true that the GDPR requires data minimization, but data minimization does not mean requesting the smallest dataset possible. It means requesting the minimum amount of data necessary to answer the research question reliably. Depending on the research question and methodology, that minimum may well be the full dataset.
In many cases, determining what is relevant, representative, or anomalous is itself part of the research process. If researchers do not have visibility over how data are selected, sampled, or filtered, they may lose the ability to assess whether the resulting dataset is adequate for answering the research question.
The example shown on the right comes from our requests. We argued that relying on platform-generated samples would not allow us to verify how the sample had been constructed or whether it was representative of the underlying population. In other words, researchers should not rely solely on the platforms’ own selection of data. Independent verification is important, especially where data-selection choices affect validity. As a result, we requested the full dataset within clearly defined categories and timeframes, arguing that this was the minimum amount of data necessary to produce scientifically robust results.
The broader lesson is that researchers should carefully align the scope of their requests with the actual requirements of their research design. Good science requires oversight of every stage of the process, including data selection and sampling. These steps should not be entirely delegated to the very platforms whose systems are being studied.

The DSA grants access to public data, but it says very little about the format in which those data should be provided. As a result, platforms retain considerable discretion over what information to include, how variables are defined, and how datasets are structured.
This can become a serious challenge for research. A variable that seems obvious to a researcher may be interpreted differently by a platform. Important metadata may be omitted. Data may be delivered in formats that make analysis difficult or even impossible. In practice, the quality of the research can depend heavily on these implementation choices.
For this reason, one of the most useful decisions we made was to attach a detailed data dictionary to our requests. Rather than simply asking for “product data” or “review data,” we specified exactly which tables we wanted, which variables should be included, how those variables should be defined, and what formats should be used.
The example shown on the right is an excerpt from one of the data dictionaries we submitted. For each table, we listed the variables required, their definitions, their data types, and any additional metadata needed to support the analysis.
The broader lesson is simple: do not leave data definitions to the platform. Researchers should define data requirements themselves, based on scientific and methodological considerations. The more specific you are about the data you need, the more likely you are to receive data that can actually support the research you intend to conduct.
From this perspective, an important question for future discussion is how standardized forms of data access, such as APIs, can accommodate the diverse data needs of systemic-risk research under the DSA.

We can also derive a number of lessons from the challenges and obstacles we encountered while seeking access to data.

The fourth lesson concerns situations in which platforms may request information that appears to go beyond the requirements explicitly established under Article 40(12).
As researchers, it is natural to assume that requests coming from legal or compliance teams reflect obligations that must be satisfied before access can be granted. However, it is important to distinguish between information that may be useful to a platform and information that is required under the legal framework governing the request.
For example, platforms may seek additional details about a research project, including performance indicators, benchmarks, assumptions, quality-control procedures, or other aspects of the methodology. While such information may be relevant from a scientific or operational perspective, researchers should consider whether these requests are necessary for demonstrating compliance with the conditions set out in Article 40(12).
The DSA requires researchers to demonstrate that they satisfy the eligibility criteria established by the Regulation and that the requested data are necessary for research addressing systemic risks. At the same time, the Regulation does not appear to assign platforms a general role in evaluating or approving the scientific merits of a research design.
For this reason, it may be useful for researchers to distinguish between information that is necessary to establish eligibility under the DSA and information that goes beyond the legal requirements of the access request. Doing so can help ensure that discussions remain focused on the conditions established by the Regulation.
The broader lesson is the importance of understanding the legal framework within which data-access requests are made. Platforms may request additional information for a variety of legitimate reasons, including internal compliance, operational, or risk-management considerations. Before investing substantial effort in responding, however, researchers may wish to consider whether the requested information is necessary under the DSA and how it relates to the legal conditions governing access. A solid understanding of the Regulation can help researchers navigate the process more efficiently and engage with platforms in a clear and informed manner.

This final lesson concerns the practical challenges that can arise when accessing data, particularly when technical issues affect the collection or retrieval process.
When such issues occur, identifying their source is not always straightforward. Researchers may receive different explanations for unexpected behaviour, access limitations, or data-collection errors. In these situations, it is important not to rely solely on initial assumptions but to approach the problem systematically and empirically.
A useful practice is to carefully document the research workflow, maintain detailed logs, replicate tests where possible, and compare observed outcomes with the expected behaviour of the systems involved. This makes it easier to identify the source of technical problems and to communicate clearly about them.
The broader lesson is that researchers should have confidence in their methods and rely on evidence when evaluating technical issues. Independent verification remains a fundamental part of the research process. Careful documentation, testing, and validation can help clarify misunderstandings, resolve technical challenges more efficiently, and support productive engagement with platforms and other stakeholders.
More generally, researchers should approach technical interactions in the same way they approach scientific questions: by gathering evidence, testing alternative explanations, and drawing conclusions based on the available information. This approach can help avoid unnecessary delays and contribute to a more effective data-access process.

Let’s move to the final takeaway: What makes a strong DSA request? This final lesson brings together the points I have discussed so far.

DSA data-access request is not just a research proposal. It is also a legal document. While researchers naturally focus on the scientific aspects of the project, the reality is that these requests are reviewed by legal and compliance teams. As a result, they are assessed not only from a scientific perspective but also from a legal one.
For this reason, researchers operating under Article 40(12) need to think not only like scientists but also, to some extent, like lawyers. A strong request requires both scientific rigor and legal precision.
From a scientific perspective, the request should clearly define the research questions, explain how they relate to systemic risks under the DSA, and present a detailed research design. The scope of the data, the relevant timeframes, the methodological approach, and the necessity of each data category should be clearly defined and explicitly linked to the relevant provisions of the Digital Services Act (DSA). The objective is to demonstrate that the requested data are genuinely required to answer the research questions.
From a legal perspective, researchers should become familiar with the relevant DSA provisions and explicitly ground their requests in them. The connection between the research and the systemic risks identified in Article 34 should be clearly documented. The request should demonstrate compliance with the requirements established by Article 40 and anticipate the most common objections that platforms may raise.
Ultimately, the strongest requests are those in which the scientific and legal arguments reinforce one another. The goal is not to create an adversarial relationship with platforms, but to produce a request that is so carefully constructed that there is little room to argue that the research falls outside the scope of Article 40(12) or that the requested data are unnecessary.
This is therefore my final takeaway from working under this provision: build the project and the request like a scientist, but defend them like a lawyer.