Project Name:
Measuring Large Language Model Understanding of Federal Statistical Data
Contractor: NORC at the University of Chicago
Lessons Learned
- Balance breadth of prompt coverage. Balancing prompts across scenarios (Data Discovery, Access, and Use) and personas (Casual/Intermediate/Experienced) can prevent over-indexing on advanced analysis prompts and ensure that the evaluation reflects a range of realistic user questions (e.g., about data update frequency).
- Design prompts to support one-to-many data asset relationships. Real users’ questions often span domains (e.g., education and income). Allowing prompts that can depend on multiple data assets reflects realistic workflows and can exposes cross-dataset metadata dependencies.
- Distinguish domain expertise from technical expertise. Generic personas (“Experienced, Intermediate, and Casual”) provide a helpful scaffold but distinguishing experienced users into domain experts and data scientists reflects real-world diversity in user needs aligned with the appropriate response context.
- Early Engagement Streamlined the PRA Clearance Process. Early planning and proactive engagement with our government partners enabled a quick and efficient Paperwork Reduction Act (PRA) clearance process. By sharing our user engagement plan with government clients early, they were able to coordinate with the appropriate officials regarding the planned outreach. A brief consultation with the Commerce PRA Clearance Officer confirmed that PRA clearance would be required for the effort. The team then promptly submitted the necessary documentation, which initiated an expedited five-day review. At the conclusion of that review, the team received OMB approval to move forward with outreach activities. This experience reinforced the value of early coordination, clear communication, and advanced preparation when navigating PRA requirements.
- Need for Understanding. The need to understand the IT infrastructure the contractor is building for. Since this is R&D, we are exploring what to build and what to deliver. This has led to a lot of questions about the long-term maintainability of the application that’s being delivered at the end. There are cost implications for how the software is developed that should be considered when developing the software to account for these AI models constantly changing.
- Accurate and complete LLM benchmark responses are critical to evaluating the AI-readiness of any data asset. Subject matter expert review of prompt-response pairs is therefore essential to ensure accuracy, completeness, and alignment with intended use cases. Wherever possible, expert validation of both prompts and responses should be incorporated into the evaluation process, as it significantly improves prompt quality and the reliability of the resulting AI-readiness assessment.
- Evaluation frameworks for LLM responses to federal data prompts should recognize that multiple valid analytical paths may lead to a correct answer and therefore emphasize conceptual correctness and authoritative sourcing rather than enforcing a single canonical response.
- User experience level and task context (discovery, access, use) significantly affect what constitutes a “good” AI response, requiring evaluation frameworks to account for audience‑appropriate depth and framing.
- Metadata, footnotes, release notes, and revision documentation contain critical contextual information that AI systems frequently miss unless explicitly prompted, making them essential focal points for AI-readiness evaluation.
- The importance of expert validation in ensuring accurate benchmark responses supports the NSDS vision of a shared service that federal agencies can trust and rely upon. When the AI-evaluation tool scales to additional subject-matter domains and agencies, the NSDS infrastructure should promote expert review that balances quality assurance with operational efficiency. This lesson suggests that a future NSDS could establish a cross-agency expert network that can support review of prompt-response pairs.
- The finding that multiple valid analytical paths can lead to correct answers has implications for how a future NSDS positions the AI-readiness evaluation as a shared service. The NSDS should provide flexible evaluation frameworks that recognize domain-specific analytical conventions, while maintaining certain standards for conceptual correctness and authoritative sourcing.
- The vision for how a shared service will be developed and deployed within the NSDS is important for the design of tools. For example, if the tool will be centrally hosted, factors such as identity, access, and cost management should be taken into account at the design stage, rather than after delivery.
- Recognizing that user experience level and task context significantly affect response quality underscores the need for a future NSDS to serve multiple user communities with varying experience levels and needs. The AI-readiness evaluation tool our team is developing will help agencies understand how their data assets perform across different use cases and audience segments. Additionally, the finding that AI systems frequently miss critical contextual information in metadata, footnotes, and documentation validates the NSDS focus on helping federal agencies gain insights on the AI-readiness of their statistical data and to optimize their statistical data assets for improved AI response quality.
• Several challenges were present at the outset of SME testing and feedback on the AI-readiness evaluation tool. SMEs had limited familiarity with the application, and the tool required navigating unfamiliar inputs, configurations, and reporting features, which created the possibility of varying feedback quality. Additionally, the evaluation design included prompt quality, dataset context, and response accuracy across knowledge-only and context-augmented modes for multiple LLMs (Claude, Gemini, GPT‑5.4) imposed a high burden, increasing the risk of fatigue and incomplete assessments. To mitigate these issues, the NORC team implemented targeted solutions to streamline engagement and improve data quality. A dedicated workspace with 25 pre-configured prompt–response pairs, linked datasets, and fully executed experiments reduced setup complexity and enabled SMEs to focus on evaluation. Detailed, screenshot-based instructions standardized navigation, and an embedded Microsoft Forms instrument structured feedback collection across key dimensions. These approaches improved consistency and produced more comparable data, enabling us to refine and strengthen our scoring methodologies.
• Overall performance metrics provide an important summary of AI capabilities, but targeted examination of challenging or unusual cases often yields deeper insights. Testing prompts involving ambiguity, data limitations, methodological nuances, or multiple valid answers can help identify strengths and limitations of tools and of LLMs that are not apparent from aggregate scores alone.
• Lessons learned highlight the need to reduce user burden and improve usability to meet NSDS goals around accessibility and impact. Challenges with SME onboarding, limited exposure to the tool, and evaluation fatigue reinforce the importance of intuitive design, detailed instructions, and streamlined evaluation processes to enable consistent participation. The use of pre-configured workspaces, structured evaluation criteria, and embedded feedback mechanisms demonstrates a scalable approach aligned with NSDS priorities such as standardization, data quality, and cross-system comparability.
• Lessons from the prior quarter show that aggregate performance metrics alone are insufficient to evaluate AI systems. Incorporating targeted tests using ambiguous, complex, or edge-case scenarios provides a more accurate assessment of system reliability, data quality, and real-world applicability.
Disclaimer: America’s DataHub Consortium (ADC), a public-private partnership, implements research opportunities that support the strategic objectives of the National Center for Science and Engineering Statistics (NCSES) within the U.S. National Science Foundation (NSF). These results document research funded through ADC and is being shared to inform interested parties of ongoing activities and to encourage further discussion. Any opinions, findings, conclusions, or recommendations expressed above do not necessarily reflect the views of NCSES or NSF. Please send questions to [email protected].




