Workshop held December 13, 2024
Proceedings
Funding: Gordon and Betty Moore Foundation
Published: July 2, 2025 (Last Updated: July 2, 2025)
On Friday, December 13, 2024, the California Council on Science and Technology (CCST) convened stakeholders to discuss health data sharing and access for policymaking during public health emergencies. CCST initiated work on public health emergencies during the COVID-19 pandemic, convening experts to discuss recommendations for improving public health response during future occurrences. This convening was a follow-up on previous recommendations and information obtained from conversations with public health experts, data privacy researchers, and policymakers in the California State Legislature.
Four themes emerged from the convening’s conversations and breakout sessions:
Authors: Kleeman, Michael; Peisert, Sean; Readhead, Heather
Suggested Citation: Kleeman, Michael, Sean Peisert, and Heather Readhead. 2025. Health Data Access and Sharing: CCST Workshop Proceedings. Sacramento, CA: California Council on Science and Technology. DOI: https://doi.org/10.63429/CAFO7362
On Friday, December 13, 2024, the California Council on Science and Technology (CCST) convened stakeholders to discuss health data sharing and access for policymaking during public health emergencies. CCST initiated work on public health emergencies during the COVID-19 pandemic, convening experts to discuss recommendations for improving public health response during future occurrences. This convening was a follow-up on previous recommendations and information obtained from conversations with public health experts, data privacy researchers, and policymakers in the California State Legislature.
A public health use case was included in pre-convening materials to gather attendees’ input on public health data needs for policy and decision-making. Discussions centered on the utility of sharing infectious disease reporting data and various considerations, including protecting data privacy and establishing a robust data governance structure, among others.
There was consensus among attendees that such a process would require considerations for the following:
In 2021, amidst the COVID-19 pandemic, CCST assembled an expert COVID-19 Steering Committee comprised of medical doctors, researchers, and public health and policy experts to identify the key areas and questions to study (see committee members in Appendix E). The Steering Committee developed an integrated set of lessons learned and recommendations from the public health response to the pandemic. They also identified two key issue areas—modernizing the public health system and health data sharing—and expanded each into CCST workshops, convening cross-sector experts for candid discussions, to develop policy and planning solutions for future pandemics. Two workshops were held with corresponding proceedings documents:
The latter convened over 50 public health and data privacy experts to explore California’s challenges to public health delivery during the pandemic.
In November 2023, recommendations from the workshop discussions were published as 11 policy briefs:
In 2023, findings from the initial set of workshops were presented to and reviewed by state legislators, health officials, and other public policy experts. Feedback from these interviews and meetings highlighted data privacy as the primary concern regarding data sharing among healthcare, public health, social services, and other community partners during a public health emergency. Legislative office staff sought to understand the impact and limitations of the new California Health and Human Services Data Exchange Framework, established via AB 133 (2021), and how it would facilitate the data sharing necessary for responding to public health emergencies. For additional interventions, they inquired about the type of support required, including new legislation, regulations, or funding.
The logic underlying the secure and privacy-preserving sharing of public health infectious disease data is that it will enhance the ability to respond more quickly and effectively to health emergencies. The COVID-19 pandemic demonstrated the value of timely access to data analysis and insights from a broad range of professionals, which enabled different and often novel ways of evaluating the data on COVID-19 and enhanced the understanding of the evolving nature of the pandemic. With the hypothesis that an expanded capacity for timely data sharing will yield better insights and lead to faster and more effective responses, a data analysis sharing framework was proposed (Appendix C). This framework, or environment, has three key attributes:
In any epidemic situation, a delay in understanding the spread of the disease can lead to additional infections, which in turn can lead to a pandemic. Thus, access to timely, complete, and accurate information about the illness and its causative agent is critical to protecting public health. The proposed framework was conceived to provide faster and broader-scope data analyses, helping to limit the spread of infection and facilitate collaboration across jurisdictional lines, thereby providing more complete and timely information and analyses. Ultimately, the goal is to enhance the quality and usability of data both before and during a health or other relevant emergency.
On December 13, 2024, CCST convened a virtual workshop focused on the value of infectious disease reporting and data sharing, including the issues raised in policy brief recommendations 6 through 11, with an emphasis on preserving privacy and security (see Appendix A for the agenda). We invited representatives from local, state, and national public health agencies, data scientists, privacy experts, healthcare providers, and academic researchers. The workshop asked participants to comment, provide feedback, or provide extra details on a use case to ensure appropriate and accurate reflection of the data needs during a public health emergency (see Appendix B). The workshop also provided an overview of two different types of privacy-preserving data governance and trusted execution environments, soliciting questions and feedback from participants regarding their use in supporting data analysis and ensuring access to timely information during public health emergencies. Lastly, participants were asked to develop recommendations and next steps for pursuing the goal of more accurate, usable, and timely data for California during public health emergencies.
To ground the conversation and provide a shared context, a use case was presented for a novel communicable virus detected in a community (see State and Local Health Departments Facing a Novel Virus with Pandemic Potential). In the scenario, an environment or framework in line with the one described above was described as being available to qualified users from any field—not only healthcare (such as sociology, economics, policy, etc.)—and the results of the requested analyses could also be seen by others to foster the sharing of insights and stimulate further investigation. In addition, a standard set of analyses could be quickly shared by all county public health agencies and health providers, allowing for better-informed responses.
Four themes emerged from the convening’s conversations and breakout sessions:
It is important to note that the basic concept of an environment or framework where public health infectious disease data could be accessed in a privacy-preserved, secure, and cross-jurisdictional manner was received positively by all in attendance and that the discussions focused on how to proceed with creating and sustainably maintaining such a capability in the right way.
Three general themes emerged in the context of stakeholders and benefits:
Several participants emphasized that resource disparities within the healthcare system limit the ability to respond effectively to health-related crises. Larger health systems, such as the University of California, Sutter Health, or Kaiser Permanente, have the personnel and systems in place to handle changing conditions and increased information demands. Likewise, larger counties, such as Los Angeles, have large staffs that can better deal with a surge of data and coordinate a response compared to smaller counties, many of which have small health staffs and no dedicated epidemiologists.
In that context, participants saw the capabilities of a data-sharing environment as a way to provide benefits to smaller healthcare practices and counties by quickly and automatically making the same quality of reports and analyses available to their larger counterparts. The ability to share clinical insights, such as evolving clinical syndromes associated with new viral variants, across all counties automatically would allow smaller jurisdictions to see these data in their context, potentially saving lives. This concept of ‘equal access’ would extend to affordability and usability. It is expected that the proposed system would reduce the workload within a public health department during ‘normal’ times, enabling personnel to keep up with data analytics tasks during peak times, and creating an economic benefit. However, it was also considered essential that the cost and local technical expertise not restrict smaller organizations from benefiting from the framework, and that equitable costs and economic considerations are addressed.
The ability to reduce the time from identification of an infected individual to inclusion of information on that case in the data reporting was seen as essential. Essential to reducing time through the use of data sharing environments is the automated reporting of reportable conditions from healthcare providers to public health agencies—something a few larger providers and counties had begun to do during the COVID-19 pandemic.
Rapid access to accurate and meaningfully analyzed data, along with the tools to help organizations quickly translate that into effective action, was seen as a key value. To achieve this goal and provide quality analyses and tools to both small and large jurisdictions, a core set of analyses with appropriate graphic and geographic outputs was recommended. The automatic generation and updating of these analyses were seen as a way to avoid overload and confusion, especially if users of these data had been receiving similar reports during non-peak periods and were comfortable with their interpretation and use.
A standard set of queries and analyses, which could be automatically run at routine intervals, was considered necessary, as was the ability to create additional custom queries and integrate other data (such as environmental or vaccination data) with the infectious disease reports. Attendees also discussed potential opportunities to use predictive analytics or machine learning to help identify emerging patterns in the data and recommend further analyses. Integral to the ability to generate stakeholder benefits and trust was the recommendation that diverse stakeholders, including policymakers, community organizations, and private industry, be engaged in defining relevant question sets.
Lastly, the concept of clear and consistent communications with the public was emphasized as a benefit and one that would help build and maintain trust. Participants noted that publicly available, consistent, reliable, and accessible information during health emergencies would improve transparency and trust in public health services and providers. We expand on the topic of trust below.
When the concept of a data-sharing environment was first proposed in 2023, several questions arose regarding privacy and security. Could such a system actually allow access to data for analysis and to inform policy while protecting the privacy of both patients and providers? How would this system mitigate either malicious or inadvertent exposure of sensitive data? In short, how could such a system be trusted with some of the most personally sensitive data?
Convening attendees explored different dimensions of these critical issues. One dimension involved determining the minimum necessary data required to answer the questions. The other related to technical means to protect privacy, such as avoiding large, centralized databases or leveraging tiers of users with corresponding levels of access privileges. Lastly, the formal (clear regulations and policies) and informal (guidelines and training for users) means of ensuring compliance were covered.
‘Minimum necessary data’ refers to the concept that there must be a legitimate, public-serving reason for including data in the system. This suggests that the scope of allowable analyses would largely be determined in advance. Studies that require additional data can be conducted only after review for compliance with the system’s boundaries. This also aligns with the concept of different levels of user access privileges, in which some users have a more limited set of questions (and thus data) that they can explore than others.
Regarding the technological means of protecting privacy and personally identifiable information, there was consensus among those who had used technologies such as differential privacy or confidential computing that this was achievable (see Appendix D for information about these technologies). It was noted that a project at the National Institutes of Health (NIH), the National COVID Cohort Collaborative (N3C), employs a similar approach, where the underlying data is not made available to researchers, but they can submit specific queries against the data. This has proven successful in protecting privacy while enabling a wide range of research on COVID-19-infected individuals. However, the N3C Consortium utilizes a large central data repository that runs on a platform operated by a single private sector entity.
In contrast, the recommended framework would maintain each county’s data in a separate, encrypted repository that can only be accessed by the query engine, which can span multiple repositories for authorized analyses. Avoiding centralization of databases also minimizes the risks of data breaches and unauthorized access. This dual layer of protection, which involves separating the data and employing end-to-end encryption, enables cross-jurisdictional analyses and reports without the need for a large central repository.
Another important output of the workshop discussions was the idea that basic and repeated analyses should be the core output of the system and that these can be designed and tested for security and privacy well in advance of use. Analysis guardrails should strike a balance between data access and security, ensuring that sensitive information is protected while allowing for meaningful analysis that helps protect public health. This can enable expanded or ad hoc queries through the use of a secure ‘clean room environment’ that automatically constrains queries to be compliant with privacy policies. Outputs for basic and common analyses would be reviewed by the appropriate stakeholder(s) to ensure there were no violations of privacy rules or laws. A finding that a user attempted to circumvent or violate the privacy guidelines would result in that user being barred from future access.
The primary purpose of this system is to support timely and effective responses to health emergencies, rather than solely for academic research. Existing access by public health officials and healthcare system care providers with authority for broader access would not be impacted. These personnel, as well as researchers authorized by public health officials, can currently access detailed records to facilitate a timely response to public health events. This system would enable an additional level of analysis, allowing personnel to access this data in a manner consistent with applicable privacy laws and regulations.
While the technology was not considered a significant challenge, the design, administration, and governance of such a data-sharing environment were essential elements. Suggestions were made about how major technology firms might support the development of this capability. However, they would likely not be trusted with the administration and operation of the system, as state or academic control was preferred.
Participants voiced that leveraging the capabilities of the private sector would be beneficial in developing the Framework’s capabilities. Partnerships with tech companies, such as Google and Microsoft, could enhance data tools and analytics capabilities. However, concerns were raised about the involvement of the private technology sector in operating and administering systems with highly sensitive personal data.
Participants emphasized the need for clear regulatory and legal guidelines. Currently, the different means of sharing data among providers and between providers and the government are not harmonized, and not all users are connected in a two-way manner. This restricts data sharing and collaboration, essential elements of any proposed data sharing framework that must get data inputs from providers and share results with providers and public health collaborators.
Finally, participants explored how a data-sharing environment might evolve to include a broader community of users who would either access the data for analysis or leverage the outputs for interventions that would improve public health. In this context, recommendations were made to consider additional tools, such as application programming interfaces (APIs) connected to the Framework environment, tailored for smaller providers. The APIs would ideally bridge resource gaps and support equitable data usage and digital tools that leverage outputs, such as environmental engineering models, to support interventions, especially in under-resourced areas like schools or rural counties.
Although not initially a specific topic area, an ongoing discussion of trust emerged, underscoring its importance for any information system, especially one handling sensitive personal information. This discussion examined trust from multiple perspectives, including:
Ultimately, participants recognized the potential of a robust and transparent data sharing framework to play a critical role in helping restore trust in the public health system by providing timely, accurate, and actionable insights that benefit the public. As one participant said:
“If they [the public] feel like you’re on top of things and you know what you’re doing and you’re doing your best and trying to help improve health outcomes for everybody at the population level, that will help build trust.”