Project Name:

Building Capacity for State, Local, and Territorial Governments to Use Administrative Data for Evidence-Building

Contractor: BrightQuery, Inc.

Lessons Learned

  • AI Backend Setup: The process of setting up the AI backend and website was a learning experience that will help streamline future developments.
  • State Collaboration: Early engagement with state partners is critical to ensure project milestones are met on schedule.
  • Data Sharing Agreements: We leveraged existing Data Sharing agreements wherever possible. Some states like CA have data linked up end to end and ready to be used while others have an MoU in place between different participating agencies where data needs to be linked and extracted on specific uses like in the case of CT. Both arrangements have their own pros and cons and are useful in different states of maturity of insight generation. The final long-term objective should be to generate insights at the state level which can then be aggregated at the federal level for a concerted policy making and evidence building.
  • AI Backend Setup: The process of setting up the AI backend and website was a learning experience that will help streamline future developments.
  • Platform Development: The development of the UI and data exploration tools provides valuable experience for future iterations of the platform.
  • Data Sharing Concerns: Data sharing concerns between agencies are legitimate and based on legal and privacy considerations. These cannot be solved simply through effort; new techniques will be required in combination with political and administrative considerations.
  • Data Silos: Data siloed within individual states will always have a “visibility horizon” at the state’s borders.
  • State Size: Larger states have an inherent advantage in collecting and linking data. Smaller states have a higher proportion of their population leaving and entering the state.
  • Workforce Churn: People drop out and re-enter the workforce in different states due to multiple reasons like the workforce/skills being seasonal, people taking a break from working and re-entering the workforce, people taking retirement etc. Eg: moose hunting or fishing which is possible in summer only, some people in financial services did not seek jobs for a few quarters after the 2008 stock market crash. In these cases, linking at a statistical level is the only alternative.
  1. Data Sharing & Governance:
    • Existing Agreements Are Crucial:
      Leveraging pre-existing Data Sharing Agreements (DSAs) expedited project execution in
      some states (e.g., California). These agreements allowed for smoother integration and
      insight generation.
    • Different States, Different Models:
      States like CA have fully linked datasets, whereas others like CT rely on MoUs for data
      use on a case-by-case basis. Each model has its pros and cons depending on the state’s
      readiness and legal context.
    • Privacy & Legal Concerns Persist:
      Legal frameworks and privacy considerations remain a major barrier. Technical solutions
      alone are insufficient; policy-level collaboration is essential for broader data integration.
  2. Data Infrastructure & Standardization
    • Value of SDMX & .Stat Suite Adoption:
      Using international standards (like SDMX) enabled consistent, interoperable data
      handling. These tools reduced manual errors and supported real-time data updates.
    • Scalability Through Open Standards:
      Open-source tools like .Stat Suite proved both cost-effective and scalable, making them
      well-suited for long-term use across varied government agencies.
  3. Platform & Technology Development
    • Rapid Progress Through Prototyping:
      The development of the AI-powered platform and visualization tools significantly
      accelerated data usability. These early-stage efforts lay a strong foundation for future
      expansion.
    • Natural Language Search Was a Win:
      A natural language interface helped make data accessible to non-technical users,
      promoting more inclusive engagement with the platform.
    • Website Deployment Provided Real-World Testing:
      Deploying a live version of the platform with California data allowed for valuable
      feedback and iteration, helping refine both backend and frontend components.
  4. Analytical Frameworks & Usability
    • Standardized Templates Are Transformative:
      Canned analysis templates promote repeatability and ease of use, especially for states
      with limited analytical capacity. These templates balance rigor with usability.
    • Need for Tailored Visual Outputs:
      Customizable visualizations and recommended methods make analysis outputs more
      actionable for policymakers.
  5. Strategic & Policy Insights
    • Interstate Comparisons Are Limited by Data Silos:
      Siloed data systems and the “visibility horizon” at state borders challenge cross-state
      analyses, especially for smaller states with high population mobility.
    • Workforce Churn Affects Longitudinal Studies:
      Seasonal labor patterns, economic shifts, and population movement complicate
      workforce tracking. Statistical linking remains the most viable long-term solution.
    • Larger States Have an Advantage:
      Bigger states naturally collect more comprehensive data, aiding longitudinal tracking
      and more complex analysis
  1. Validation Is Crucial: Early and thorough validation of Data Commons outputs and canned reports ensured quality, reduced downstream errors, and increased stakeholder confidence.
  2. Template Portability: Modular design of canned reports allowed for flexibility and cross-state adaptability but requires continuous refinement to account for state-level data variations.
  3. Search Optimization: Combining lexical and semantic search delivered significantly improved user experiences compared to either method alone.
  4. State-Specific Needs: Tailoring platform deployment for each state’s data maturity level will remain key to successful expansion.

From the Agency Perspective

  1. Dynamic Over Static Integration: Agencies recognize that dynamic, AI-driven synthesis avoids the need for costly and time-consuming hard-coded integrations, offering faster scalability across programs and datasets.
  1. Improved Transparency and Trust: Semantic and lexical search together deliver higher-quality, more explainable results, which agencies see as vital for building trust with policymakers and the public.
  2. Interoperability Value: Using the Data Commons framework ensures outputs are consistent with federal metadata and FAIR principles, reducing downstream technical debt.

From the State Perspective

  1. Practical Demonstrations Drive Adoption: Florida and Connecticut emphasized that live, state-specific demonstrations (wage + education data) are critical to understanding the value of the platform.
  1. Scalability Matters: States are drawn to solutions that can be extended easily to other policy domains, such as workforce readiness, re-employment programs, or healthcare outcomes.
  1. Ease of Use Is Paramount: Chatbot-based querying significantly reduces reliance on technical staff, enabling policymakers and analysts to self-serve data insights.
  1. Evidence Pathways: Dynamic linkages create more complete narratives for decision-making, e.g., tracing how education outcomes connect to wage data in real time.
  1. Administrative Data Integration Across Jurisdictions

Cross-domain integration significantly improves evidence-building capabilities. Combining state datasets with related federal datasets provides a much richer analytical environment than either data source alone. The integration of Connecticut’s wage and education datasets with federal demographic and economic indicators demonstrated the value of a unified data environment for policy analysis.

Data alignment requires careful schema normalization and metadata harmonization.

State and federal datasets often differ in structure, naming conventions, and metadata quality. Successful integration required explicit field mapping, schema alignment, and normalization processes to ensure that chatbot queries returned consistent and interpretable results.

  1. Natural Language Interfaces Expand Accessibility of Administrative Data

Conversational AI interfaces reduce the technical barrier to data exploration.

The chatbot interface implemented in the ADEB24 platform allows users to ask natural language questions and retrieve answers derived from multiple datasets simultaneously. This approach makes the platform accessible to policy analysts and agency staff who may not have specialized technical skills in data querying or statistical software.

Transparency and source traceability remain essential for trust.

While conversational interfaces simplify access, users must still be able to verify the origin of results. Providing direct references to underlying datasets and maintaining traceable links to source data proved essential for maintaining confidence in the system’s outputs.

  1. Stakeholder Testing Is Critical for Government AI Deployments

Early engagement with stakeholders improves system usability.

Stakeholder testing using Connecticut datasets helped validate the platform’s functionality and ensured that the system addressed real policy analysis needs. Feedback during testing helped refine query workflows, sample questions, and user guidance materials.

Providing sample questions accelerates adoption.

Sample queries provided during testing helped stakeholders understand how to interact with the system and demonstrated the types of questions the platform could answer. This proved particularly useful in helping users transition from traditional data browsing to conversational interaction with datasets.

  1. Data Quality and Documentation Remain Foundational

High-quality documentation improves AI-driven retrieval.

Accurate data dictionaries, clear variable descriptions, and consistent metadata significantly improve the ability of AI systems to interpret and retrieve relevant data.

Structured public data portals simplify ingestion pipelines.

Datasets provided through well-organized open data portals allowed the BrightQuery ingestion pipeline to operate efficiently and reduced the amount of manual preprocessing required.

  1. Integrated Platforms Enable Cross-Domain Policy Insights

Linking workforce, education, and demographic data enables deeper analysis.

The ADEB24 platform demonstrated that combining datasets across domains allows analysts to explore relationships that would otherwise remain difficult to identify using isolated datasets.

Unified platforms reduce fragmentation across public datasets.

State and federal datasets are often distributed across many separate portals. Providing a unified interface significantly improves the efficiency of evidence-building workflows.

  1. Conversational Data Systems Require Both Structured Data and Retrieval Pipelines

Retrieval pipelines must be carefully validated.

The project implemented Q&A validation tests to ensure that the conversational interface retrieved correct data and maintained consistency with underlying datasets.

Combining browsing and conversational workflows improves usability.

Users benefited from the ability to both directly browse datasets and query them using natural language. This dual approach allowed users to verify results while still taking advantage of AI-assisted discovery.

  • A standardized data modelling and integration approach simplified analytical development across multiple use cases. Consistent use of SOC, NAICS, and other crosswalks enabled seamless integration of datasets from multiple public and proxy sources. This reusable framework reduced duplication of effort and supported consistent analysis across all research questions.
  • Knowledge graph-based data linkage improved the effectiveness of search, discovery, and workforce analysis. Organizing relationships between occupations, wages, transitions, credentials, and benchmark datasets provided a scalable foundation for answering complex research questions. The approach also simplified future extension of the analytical framework with additional datasets.
  • Using proxy datasets allowed project progress despite restricted access to state-level data. Public and proxy datasets enabled validation of the analytical methodology while access requests for restricted datasets were pending. Designing the solution to accommodate future replacement of proxy data minimized project risk and maintained delivery timelines.
  • Continuous quality assurance and stakeholder validation improved solution reliability. Comprehensive testing against all Statement of Work questions, combined with stakeholder demonstrations, helped identify issues early and confirm that analytical results met project expectations. This reduced rework during the final stages of the project.
  • Comprehensive documentation throughout the project simplified knowledge transfer and project closure. Preparing milestone reports, technical documentation, and supporting artifacts in parallel with implementation ensured that deliverables accurately reflected the completed work. This approach also facilitated final reporting and future maintenance activities.

Disclaimer: America’s DataHub Consortium (ADC), a public-private partnership, implements research opportunities that support the strategic objectives of the National Center for Science and Engineering Statistics (NCSES) within the U.S. National Science Foundation (NSF). These results document research funded through ADC and is being shared to inform interested parties of ongoing activities and to encourage further discussion. Any opinions, findings, conclusions, or recommendations expressed above do not necessarily reflect the views of NCSES or NSF. Please send questions to [email protected].