Project Name:

Fraud in Public Healthcare Programs: Developing Fraud Detection Models with Linked Data

Contractor: CORMAC Corporation

Lessons Learned

  1. Access to federally protected data is complex and time-consuming.

a) CMS data environments allow authorized users to access CMS datasets under datasharing agreements and to perform analyses across those datasets. However, gaining access is a multi-step process that can take several weeks, beginning with the approval of a data use agreement. Additional requirements include completing business and technical training (e.g., database and BI tools), configuring VPN access on local computers, and navigating agency-specific procedures. Timely progress depends on support from internal CMS resources and clear, open communication channels to resolve issues efficiently.

b) Early decisions about the data access and analysis environment are critical. Any environment used to access federal data must comply with federal cybersecurity and data security requirements. The team learned that access to a secure agency environment could have been expedited if an interagency agreement had been established before the project launch.

  1. A narrow analytical focus is essential for short-duration projects.

Given the project’s limited timeframe, it was critical to work closely with the Department of Justice’s Fraud team to focus on one or two related types of fraud that are both high-risk and feasible within existing data and system constraints. The following factors guided the selection of the project’s focus:

  • Data availability, accessibility, and usefulness
  • Likelihood of fraud based on known or similar cases
  • Care setting
  • Billing or business practices
  • Provider type
  • Geographic location
  1. Careful definition of key variables is foundational to AI/ML-based fraud detection.

During the first quarter, Cormac focused on identifying the data elements and contextual information most likely to indicate fraudulent activity, as well as the analytical processes and tools needed to flag suspicious patterns. The team also considered the required outputs and how results should be communicated clearly, including in plain language, to support effective interpretation and use.

The project had the following lessons learned for this quarter that could inform a future NSDS: 

  • Structured sprint execution strengthens delivery clarity and accountability. With a short performance period of work, we decided to adopt a defined sprint framework to enable a clear sequencing of technical work, resulting in improved visibility across complex, interdependent tasks.
  • Foundational data engineering demands rigorous upfront validation. As we began exploring the data, we needed to build a master list of providers, episodes, and claims structures, which required iterative validation steps to ensure consistency across datasets, underscoring a future NSDS would need to allocate sufficient time for data preparation.  
  • Metric development is most effective with iterative testing. In addition to setting up rigorous validations at the start, we needed to break metrics into discrete, testable units, which allowed us to improve accuracy and transparency while enabling early detection of data quality issues.  
  • Data access timing remains a critical dependency.  Because we continued to experience delays in securing access to the compute environment, we resequenced the work, highlighting the importance of contingency planning in data-dependent efforts.  
  • Continuous stakeholder engagement ensures technical alignment. For the success of any project in an NSDS, ongoing collaborations and discussions with subject-matter experts and clients (i.e., DOJ and NSF partners) has helped keep metric development and analytic approaches relevant and aligned with project objectives.
  • Data Integration is a Critical and Resource-Intensive Phase Linking claims, enrollment, beneficiary data and clinical surveys is typically the most time-consuming and complex stage, requiring significant effort in data cleaning, standardization, and entity resolution.
  • Attribution Requires Advanced Modeling Due to Data Limitations Data used to identify referring and certifying providers is often incomplete or inconsistent. As a result, teams should be prepared to address these gaps by developing robust attribution models to accurately infer provider relationships. 

Disclaimer: America’s DataHub Consortium (ADC), a public-private partnership, implements research opportunities that support the strategic objectives of the National Center for Science and Engineering Statistics (NCSES) within the U.S. National Science Foundation (NSF). These results document research funded through ADC and is being shared to inform interested parties of ongoing activities and to encourage further discussion. Any opinions, findings, conclusions, or recommendations expressed above do not necessarily reflect the views of NCSES or NSF. Please send questions to [email protected].