Interactive analysis of LLM performance on charity document data extraction
Aggregated stats per provider (active models only, F1 > 0). Fields and cost are averages per model.
Each cell shows how accurately a model extracted that field across all 11 documents. Click a cell for details.
Average accuracy across all functional models
Each cell shows accuracy for a specific document/model pair.
How this extraction benchmark evolved — from raw data preparation to a multi-provider, multi-model scored leaderboard.
From extraction_stats.csv — actual observed time and cost per model (where available).