Large language models (LLMs) are increasingly used for text annotation in social science, but standard performance metrics do not explain why errors occur. Low F1 may reflect limitations of the annotation system, defects in the written codebook, or inconsistency in reference labels. We propose a pre-deployment codebook audit for LLM-assisted annotation. The audit uses class-specific metrics and confusion matrices to identify problematic classes and boundaries, structured document-level review to attribute error sources, and targeted codebook revision followed by held-out evaluation. We apply the audit to two published multi-class codebooks from distinct domains. In both applications, the audit identifies under-specified or conflicting coding rules, and targeted revisions reduce the diagnosed errors on held-out data. Analyses with two LLM annotators show that some problematic boundaries and codebook revisions transfer across models, whereas other revision effects are annotator-specific. The results show that LLM annotation errors can help identify weaknesses in the written measurement instrument, not only in the model. Codebook auditing complements prompting, retrieval, and fine-tuning by clarifying whether persistent errors call for model improvement, codebook revision, or renewed scrutiny of reference labels.
Presented at PolMeth 2026, Michigan State University, July 2026, and — as “When F1 Lies: Taking Constructs Seriously in AI-Assisted Measurement” — at AI for Social Science Research Methods, Yale University, May 2026.