Project Name:
Linking State and Federal Data to Strengthen Federal Program Evaluation
Contractor: The Urban Institute
Lessons Learned
During the project’s first quarter, Urban submitted and revised detailed analytic plans for the two major project tasks: Task 1, conducting a qualitative study on barriers and opportunities for federal–state data linkages, and Task 2, executing privacy-preserving record linkage (PPRL) within the National Secure Data Service (NSDS).
A key lesson learned is that successful PPRL projects require extensive coordination across legal, procedural, and technical governance domains. Understanding the privacy and security capabilities of existing systems proved challenging, particularly when engaging stakeholders with varied technical expertise. To address this, we tailored outreach materials to distinct audiences to ensure consistent interpretation of project goals, privacy requirements, and technical constraints. This approach improved our ability to assess system readiness and identify logistical barriers early in the PPRL process.
Another major challenge was timely access to raw data, which is often delayed due to the high coordination burden inherent in PPRL projects. To mitigate this bottleneck, we developed an end-to-end PPRL simulation using synthetic data informed by publicly available sources. This workaround enabled parallel progress on code development and governance planning while formal data access processes were underway.
Lessons from this quarter highlight opportunities for NSDS to strengthen its role beyond providing secure computing infrastructure. Most federal–state linkages rely on shared unique identifiers or common PII fields; however, for this project, address data provided by states is likely to be central due to both the nature of HUD data and the incompleteness of state records. While PPRL systems can be designed to handle address data, many existing open-source solutions (such as those offered by MRAIA) do not currently support address pre-processing under robust security certifications.
These findings suggest potential value in NSDS investing in PPRL software and infrastructure that supports secure address pre-processing. The broader utility of such an investment depends on whether this challenge is common across federal–state linkages or largely specific to HUD and CCDF data. Combining insights from Task 1 (qualitative interviews) and Task 2 (technical PPRL testing) will help NSDS determine where standardized tools would have the greatest impact.
More broadly, NSDS could play a proactive role in improving data interoperability by facilitating the collection and dissemination of metadata and linkage-quality assessments. This would make potential linkages more discoverable while helping researchers anticipate match quality issues and mitigate risks associated with low-quality or incomplete identifiers.
During the project’s second quarter, Urban began its work on the two major project tasks: Task 1, performing a qualitative study on barriers and opportunities for federal-state data linkages, and Task 2, testing privacy-preserving record linkage (PPRL) inside the National Secure Data Service (NSDS). Early work has reinforced the importance of aligning legal, administrative, and technical requirements when establishing federal-state data linkages.
Federal-state data linkage presents legal and administrative challenges.
Our work on Task 1 indicates that merging state and federal data is extremely challenging. Some of the primary challenges are having to share identifiers across systems, privacy laws that vary from state to state, and the time it takes both at the federal and state level to get data sharing approved and develop and sign a data sharing agreement. Because there are not well-established legal or procedural precedents for data sharing between federal and state agencies, these processes create more administrative burden than in the federal-to-federal case.
State participation depends on minimizing burden and demonstrating value.
Recruiting state agencies to participate in Task 2 has been challenging, in part because of limited staff capacity. To address this challenge, the team has focused on reducing burden for states, including requesting data that is already reported and offering technical assistance. The team has also focused on emphasizing the benefits of participation for the state, including but not limited to state-specific evaluation goals, data modernization, and infrastructural preparation for state-to-state linkages.
Balancing technical robustness with usability remains a challenge.
Task 2 has highlighted the need to balance strong privacy and security protections with practical implementation requirements. PPRL software solutions must work effectively within the NSDS environment while also being assessable to state partners with varying levels of technical capacity. There is a needed balance between robust software that addresses all privacy and security concerns and options that are flexible enough for both state partners who might have limited technical capacity and ensuring that software requirements align with the NSDS environment has presented challenges.
Simulation with realistic, non-sensitive data is critical.
The team began its data linkage simulation exercise, the goal of which is to emulate the expected data challenges using public or non-sensitive datasets. Using realistic, non-sensitive data and simulated identifiers is crucial to the success of the simulation, as it directly affects encryption and matching performance.
PPRL can reduce the need to share direct identifiers, but implementation burden remains a consideration.
PPRL offers an important and valuable solution to one of the primary challenges in matching state and federal data by ensuring that human-readable identifiers do not need to be shared across agencies. However, staff burden and long review processes remain significant barriers to implementation.
Future NSDS should minimize burden and emphasize value.
States are more likely to participate when the burden is minimal and the benefits are clear. One way to minimize burden is to focus on collecting data that is already reported. Many state agencies must report data to the federal government; however, this data is often de-identified and therefore cannot be matched at the federal level. NSDS can minimize burden by prioritizing the use of already-reported administrative data, which is often standardized from state to state due to federal reporting requirements. A limitation is that these data sets may not include all variables of interest, so it is important to effectively communicate the benefits of participation.
Clear benefits can encourage state participation.
States are most motivated by two types of benefits. One benefit is the ability to answer policy-relevant questions which could be supported through access to the matched data or analytic services. The other benefit is access to PPRL technology that is being used, which States can utilize to match data sets across state agencies. Together, these incentives can encourage participation while supporting broader state data modernizing their own systems.
During the project’s third quarter, Urban Institute completed Task 1, completing a qualitative study on barriers and opportunities for federal-state data linkages, and continued its work on Task 2, implementing privacy-preserving record linkage (PPRL) inside the National Secure Data Service (NSDS).
The project team identified several lessons that can help inform the future federal-state data linkage efforts:
- Consistent and proactive communication is essential for advancing a data sharing agreement (DSA).
Regular meetings and check-ins can ensure that processes are not delayed and that issues that arise are handled quickly without losing momentum. Clear communication with all partners about expected timeframe and potential constraints is also critical for maintaining momentum and developing realistic project schedules.
- Data sharing agreements benefit from clear visualizations and plain-language explanations of data flows.
Providing a clear description or visual map of the data sharing process including what data are being transferred, who owns the data, and who will receive them, can help partners understand the data sharing process and can expedite the data sharing process. In addition, ensuring the PPRL process is described in plain language can also help stakeholders understand the proposed approach and facilitate review of the DSA.
- Data ownership is often complex and requires explicit clarification.
Many states rely on external sources, for instance external data sources, making it important to clearly identify data ownership and the roles and responsibilities of relevant stakeholders. Establishing this information early can help ensure that the appropriate parties are involved in the data-sharing process.
- Data sharing agreements often require longer timeline than initially expected.
DSAs may require extended review and approval periods involving multiple stakeholders. Future projects should account for these timelines when developing project schedules and allow sufficient time for review, revisions, and approvals.
- Given the complexity of PPRL, limited state capacity, and the stringent legal requirements of state partners, our team was tasked with developing preprocessing tools that could be used by state partners before data were transferred to NSDS. This approach was designed to reduce the need for direct access to personally identifiable information (PII) by third parties while supporting consistent preparation of data for linkage.
- Developing tools for data that the project team cannot directly inspect requires careful planning and testing. The tools must be sufficiently flexible to accommodate differences in data structures and formats across state partners.
- Comprehensive validation and clear error reporting are also important for identifying issues efficiently and reducing time-consuming back-and-forth during implementation.
- Clear, plain-language instructions are similarly important. State partners need to understand the steps required to prepare and submit their data, including the sequence of activities, the purpose of each step, and the technical requirements for implementation.
- In some cases, additional variables or transformations may also be needed to support analysis while protecting sensitive information.
- Consistent communication among all partners involved in any DSA remains essential for maintaining momentum and ensuring timely resolution of issues. DSAs require extended timelines due to multi-party legal review and approval processes, and future NSDS work should explicitly account for these constraints in planning and scheduling.
In addition, NSDS could maintain a repository of “template code” and guidance for common data preprocessing and encrypting activities. While these resources would need to be adapted for individual data-use cases, they could provide a useful starting point for state partners, reduce duplication of effort, and help streamline future data linkage projects.
Disclaimer: America’s DataHub Consortium (ADC), a public-private partnership, implements research opportunities that support the strategic objectives of the National Center for Science and Engineering Statistics (NCSES) within the U.S. National Science Foundation (NSF). These results document research funded through ADC and is being shared to inform interested parties of ongoing activities and to encourage further discussion. Any opinions, findings, conclusions, or recommendations expressed above do not necessarily reflect the views of NCSES or NSF. Please send questions to [email protected].




