Topics covered

Subject ▸ validity

When F1 Is Not Enough: Auditing Codebooks for LLM-Assisted Annotation

Large language models (LLMs) are increasingly used for text annotation in social science, but standard performance metrics do not explain why errors occur. Low F1 may reflect limitations of the annotation system, defects in the written codebook, or inconsistency in reference labels. We propose a pre-deployment codebook audit for LLM-assisted annotation. The audit uses class-specific metrics and confusion matrices to identify problematic classes and boundaries, structured document-level review to attribute error sources, and targeted codebook revision followed by held-out evaluation.

Read More…

Can We Trust LLM-Generated Data? The CRAFT Framework for Measurement and Inference in Political Science

Large language models are increasingly used for data generation in political and social science, yet the discipline lacks a shared standard for validating their output. Existing frameworks address pieces of the workflow, mostly covering a single stage. We propose C-R-A-F-T, a five-step framework that connects construct definition through inferential adjustment within a single, model-agnostic specification: C-onstruct roles and tasks; R-eport dual-track metrics; A-ssess stability across prompts, and models; F-ield human audit and adjudication; T-ranslate to inference incorporating uncertainty.

Read More…