Project Name:

Secure Multiparty Computation: A Case Study

Contractor: Stealth Software Technologies, Inc.

Lessons Learned

  • Why SMPC?
    • A long-standing hypothesis of SMPC pioneers is that this technology will make data sharing much safer and therefore, easier to convince stakeholders. Several recent examples just in the past 5 years have shown that this is the case, but it has to come from a combination of technological and administrative alignment and not simply SMPC alone. Even within privacy-enhancing technologies, SMPC is only one tool in that toolbox, but it uniquely offers a much stronger privacy and functionality guarantee that other tools cannot achieve (of course, at the cost of performance). In some cases, this may be overkill, and “secure enough” is secure enough, but we have found that this additional strength may be what ultimately moves the needle.
    • However, there is still a general feel of “too new” or not having enough previous use-cases to point to. One hospital we spoke with preferred a Privacy-Preserving Record Linkage solution using secure hashing as opposed to SMPC simply because they saw it being used in a previous city-wide study. Like many technologies, there is a network effect of just having more organizations using the technology and having heard of it. These points only further highlight the importance of this project.
    • For our specific use case, we have noted that even with the technology being offered to protect privacy, it appears that there needs to be an internal business need to motivate the private sector partner to engage their legal team to collaborate on this work.
  • Technology Insights
    • In order for the SMPC technology to scale for use in multiple concurrent studies, there requires a standardized, auditable account management capability. To support future studies, we should implement (or adopt) an account management system that allows each user to authenticate and manage the studies they are involved in (e.g., view status, track necessary actions). For this demonstration project, we are assessing whether existing secure data platforms, within a future National Secure Data Service, can integrate with an open-source identity/account stack (e.g., Django-based authentication or comparable frameworks) or whether a separate, agency-approved solution is required.
  • Data Governance and Access
    • For commercial data governed by restrictive data-use and redisclosure controls, the data governance or legal stakeholders should be engaged at project initiation and treat approval timelines as a critical-path dependency. In the early outreach, it is important to use precise language that clearly distinguishes (a) internal project use vs. redisclosure to additional agencies/contractors, (b) access to schema/metadata vs access to full datasets (including synthetic), and (c) who will have access, where data will be stored, and under what controls. For this demonstration project, data access approvals have remained unresolved after ~3 months, future project plans should allocate 4-6 months for commercial data acquisition that involves multi-party access and redisclosure constraints, especially when approvals span multiple organizations. In this case study, we initially expected data acquisition from J.P. Morgan Chase (JPMC) to be straightforward because the datasets are synthetic and publicly listed, and we had a supportive point of contact. In practice, however, their terms of use restricted downstream sharing, which conflicted with our multi-agency project structure; subsequent clarification requests were declined by the data governance team, creating a schedule impact. To mitigate schedule risk (particularly given year-end holiday slowdowns), the team is continuing engagement with JPMC’s data governance stakeholders to identify a compliant access/sharing path while, in parallel, searching for alternative data sources as a contingency should JPMC ultimately be unable to support the project’s multi-party access requirements.
    • For the other datasets in the project, data acquisition has progressed more efficiently because we had a clear path to the decision-makers and the U.S. Department of Veterans Affairs (VA) saw clear value in participating—i.e., securely analyzing their dataset alongside others to benefit their users, veterans. A key accelerant was the team’s preparation of a concise, one-page partner brief using VA’s own terminology and publicly available references, clearly mapping the project objectives to VA’s stated priorities and articulating the mutual benefits of participation. Additionally, having Western Institute for Veterans Research represented on the project team (and funded under the project) improved prioritization, responsiveness, and timeline adherence. For future efforts, we should explicitly align participation benefits for each data contributor, package the ask in the partner’s language with a brief, evidence-backed value proposition, and ensure early engagement with the correct owners of data access and approval pathways.
    • For future reference, the projects should treat data rights and governance alignment as a formal prerequisite to cross-organization linkage work and plan accordingly. This does not necessarily require completing all approvals before submitting a project proposal; rather, proposals should scope and resource these activities as early-phase tasks with clear entry/exit criteria, and allocate a realistic schedule margin for multi-party review cycles. Specifically, 1) confirm the data use and redisclosure terms early and document them as an entry criterion for cross-agency analysis; 2) obtain full stakeholder buy-in and documented commitment, and 3) ensure empowered points of contact are identified at each partner organization and that incentives and accountability (deliverables/timelines, and when applicable, funded participation) are aligned to support on-schedule execution.
  • From the DUA and data acquisition journey during this quarter:
    • Based on continued engagement with external data partners during the DUA process, the team identified several lessons regarding partner coordination and data acquisition planning. In January, after multiple rounds of continued communication to clarify the scope of the request, the team successfully obtained access to a synthetic dataset from a large U.S. banking institution This outcome was facilitated by the support of an internal stakeholder of the banking institution who helped escalate the issues and advocate for the project within the partner organization.
    • For the MDClone dataset, the team has continued to make progress through multiple meetings and ongoing communication with the VA to navigate internal processes and clarify technical and governance requirements. Engagement has been accelerated through the involvement of key stakeholders, WIVR on the team, who helped ensure the discussions reached the appropriate decision-makers. At the same time, the overall timeline has also been affected by internal organizational restructuring within the VA, independent of the project’s technical scope.
    • These interactions highlighted several recurring lessons: First, identifying and engaging key internal stakeholders early in the process is critical for facilitating introductions, escalating issues, and maintaining momentum. Second, the framing of the initial data request can significantly influence partner engagement; in practice, partners often respond more constructively when invited to outline feasible options within their own legal and governance frameworks (e.g., a designated point-of-contact model).
  • Output Privacy Concerns
    • Data that require SMPC for data sharing or joint computation are either sensitive or are provided in a low-trust setting. The National Secure Data Service Demonstration (NSDS-D) projects have thus far been mostly interested in using these data for statistical purposes, which include releasing summaries of the data, including counts, totals, means, order statistics, and regression coefficients to the public. Potential data providers, such as credit bureaus and financial service providers, who are unwilling to share data without a tool like SMPC, are also unlikely to share unaltered results from SMPC.
    • In other words, if input privacy (e.g., SMPC or PPRL) is a precondition of sharing data for statistical purposes, then output privacy (e.g., noise infusion or differential privacy) may also be a precondition of sharing data for statistical purposes. This project has not needed to directly address this challenge because our proof-of-concept data are fully synthetic datasets. Even so, JPMC has strict requirements about the sharing of its synthetic data. From experience and conversations with colleagues, most implementations of SMPC for statistical purposes will likely need to be paired with other privacy procedures or technology, such as differential privacy.
  • Regulatory and Release Protocols
    • Obtaining permission to release information about our SMPC prototype from a prior project (Catalyst) with the Department of War would have to go through the DARPA Public Release process, documented here: <https://www.darpa.mil/news/public-affairs/public-release>.
    • As we move forward in seeking potential data partners, we learned that clear guidance on information disclosure should be established at the onset of a project. Well-defined instructions on what project details can be shared during initial outreach would help ensure consistent communication, reduce uncertainty around disclosure boundaries, and support more efficient engagement with prospective partners.
  • Institutional and Legal Timelines (Task-3 Adjustment)
    • Consultations with legal counsel (e.g., the VA) revealed that developing and securing institutional buy-in for novel DUAs tailored for Privacy Enhancing Technologies (PETs) within 10–11 months is not feasible. In addition, since the prior ADC project (with agreement number ADC-PPRL1-23-N03) largely covered the development of the Data Sharing Agreement (DSA), after consulting with NSF/NCSES, we decided to re-align Task 3 to focus on securing DUAs from the data owners and documenting this data acquisition journey for future reference.
  • Operational Blockers for Data Acquisition
    • Data procurement was hindered by internal VA reorganization, which impacted the Office of Health Innovation and Learning (OHIL) and delayed the delivery of MDClone-generated synthetic data within the ARCHES environment. Such external factors can hinder data procurement and delay the delivery of synthetic data environments, even when the technical scope remains unchanged.
  • Software Licensing and Risk Mitigation
    • During the development of the SMPC implementation for the targeted functionalities, we chose to rely on an external GitHub library originally released under the MIT License. While MIT is permissive, easy to understand, and imposes only minimal obligations, it does not include an explicit patent grant. To better mitigate potential patent infringement risk, we preferred to use the Apache License 2.0, which remains similarly permissive but provides additional protection through express patent license and patent retaliation provisions. These features make Apache 2.0 generally more suitable for commercial use and for collaboration across larger teams. Because the library’s main maintainer is part of our team, we were able to convince them to add an Apache 2.0 license, thereby reducing legal risk while preserving the flexibility of use.
  • Our experience during the data acquisition phase demonstrated that assessing data fitness for use often requires deeper inspection (e.g., executing exploratory queries directly on the dataset) rather than relying on high-level descriptions. For instance, the initial synthetic finance dataset ultimately lacked the specific longitudinal variables required for the study, prompting the team to seek alternative data sources. Because privacy-preserving exploratory querying was not within the scope of the current contract, we had to manually run exploratory Python scripts, provided by the study Principal Investigator (PI), on the Data Partner (DP)’s behalf to identify these missing variables. This experience yielded two key lessons: First, determining a dataset’s properties prior to direct access remains an unsolved challenge in data science. To align with the broader NSDS mission of creating standardized, reusable data infrastructure, future data-sharing frameworks should adopt protocols based on the Federal Committee on Statistical Methodology (FCSM) Framework for Data Quality. Second, navigating this challenge led our team to come up with a novel future application: utilizing VaultDB specifically as a privacy-preserving “pre-filtering” tool for the exploratory phase. Maturing VaultDB so that it can automatically handle any general SQL queries without manual edits falls outside our current scope of work, but is a high-value area for future research and a separate program. In an envisioned future, a study PI would be able to write exploratory SQL queries, governed by privacy-preserving release mechanisms, to investigate a dataset’s fitness for use. VaultDB converts these queries into MPC operations, allowing the PI to receive the query results and verify a dataset’s fitness for use without the raw data ever leaving the DP’s machine.
  • Our project relied on synthetic datasets to meet tight project timelines and satisfy the exploration of deploying an MPC tool. The project did not explore the lengthy compliance processes (e.g., Authority To Operate) required for sensitive data due to time constraints. During the data acquisition journey, however, we observed that external organizations occasionally confuse “synthetic data” with arbitrary “fake data”, which can stall early negotiations. Future projects should establish clear terminology at the onset of a project to ensure alignment. For instance, emphasizing that high-quality synthetic data are mathematically generated to preserve real-world statistical properties without exposing sensitive information (e.g., personally identifiable information (PII)), making it a secure, valid tool for prototyping or piloting new technologies, data integration, early data access, and more. Contrary to fake data that may not have any of the properties of the true data and therefore may not shed light on the processes needed for prototyping.
  • We encountered a recurring challenge where the study Principal Investigator (PI) needs structural or statistical information about a dataset to write accurate preprocessing scripts and establish the study operations over the data but cannot legally access this information without a fully executed Data Use Agreement (DUA). Because Data Partners (DPs) rarely have the internal engineering resources or motivation to run exploratory scripts on the PI’s behalf, a member organization of our team had to go through the standard, lengthy legal provisioning processes to get access to the synthetic datasets and collaborate closely with the study PI on the DPs’ behalf. However, in real-world deployment, our Secure Data Statistics Platform (SDSP) architecture can alleviate this friction. By leveraging VaultDB’s SQL-to-MPC conversion alongside SDSP’s user-friendly Web Portal (or CLI tool), the platform allows non-expert DPs to execute the PI’s queries locally with just a couple of clicks. This capability will allow DPs to securely share cryptographically protected data shares directly, without exposing their raw data or requiring third-party engineering intervention, thereby minimizing legal friction and enabling the use of real datasets. This architecture has been shared with the NSDS team for future implementation in an NSDS environment.
  • As noted in the previous quarter, institutional data provisioning is highly susceptible to external delays. Navigating these roadblocks reinforced the immense value of identifying and empowering committed “champions” within the partner organizations who can advocate for the project and push it through internal bureaucratic bottlenecks.
  • The project utilized mature MPC software (EMP-toolkit1) and VaultDB2, to seamlessly translate standard relational SQL database queries into secure MPC protocols. We knew from the start that such a framework does not and would not cover all possible data science activities but had enough of a foundation to serve most of the query types used in our project. However, our use-case study requires linear regression, which cannot be effectively expressed or optimized using standard SQL statements. Because our team includes the primary maintainer of the EMP-toolkit, we successfully custom-built the MPC implementation for linear regression in a short period of time. Nonetheless, this highlighted that executing complex workflows still requires cryptographic expertise to manually write and tune MPC code.While the long-term goal is to fully automate translation of general-purpose languages (e.g., Python/R) that data scientists normally use into optimized MPC circuits, our team found this goal impractical in the near term. A more practical and feasible near-term approach for NSDS infrastructure would be to define a concrete set of common statistical operations and natively support them as pre-packaged software modules.
  • Looking forward, the application of advanced Artificial Intelligence (AI) and Large Language Models (LLMs) presents a transformative opportunity. During this project, we also utilized AI tools to draft custom MPC engine code, which drastically expedited the development process. However, the AI-generated code still required debugging and revision by our internal cryptographic experts. As AI models mature, future AI tools could be leveraged to automate MPC code generation, data harmonization, and schema preprocessing, drastically lowering the barrier to entry for privacy-preserving computation.
  • Real-world data linkage requires extensive preprocessing to resolve variable inconsistencies (e.g., typos or varying coding standards) and therefore, exact matching on key variables without close collaboration cannot be assumed. While advanced cryptographic primitives like fuzzy Private Set Intersection (fuzzy PSI) exist to match “close enough” records, they are currently orders of magnitude more computationally expensive than standard, exact-matching PSI. Consequently, rigorous data cleaning and standardization must be performed locally on the DP’s side before feeding the data into the current MPC engine until future cryptographic research improves the performance of fuzzy PSI and allows SDSP to efficiently integrate these techniques.

Disclaimer: America’s DataHub Consortium (ADC), a public-private partnership, implements research opportunities that support the strategic objectives of the National Center for Science and Engineering Statistics (NCSES) within the U.S. National Science Foundation (NSF). These results document research funded through ADC and is being shared to inform interested parties of ongoing activities and to encourage further discussion. Any opinions, findings, conclusions, or recommendations expressed above do not necessarily reflect the views of NCSES or NSF. Please send questions to [email protected].