Interactive analysis of LLM performance on charity document data extraction
Aggregated stats per provider (active models only, F1 > 0). Fields and cost are averages per model.
Each cell shows how accurately a model extracted that field across all 11 documents. Click a cell for details.
Average accuracy across all functional models
Each cell shows accuracy for a specific document/model pair.
Side-by-side comparison across all providers (active models only).
How models are distributed across pricing tiers per provider.
Higher is better — F1 score achieved per dollar spent (active models with cost data).
Text-only vs multimodal model distribution and performance per provider.
All configured models with capabilities, pricing, and performance. Filter by provider.
How this extraction benchmark evolved — from raw data preparation to a multi-provider, multi-model scored leaderboard.
From extraction_stats.csv — actual observed time and cost per model (where available).