Interactive analysis of LLM performance on charity document data extraction
Same models, order, and F1 values as the leaderboard (scroll to see all).
Aggregated stats per provider (active models only, F1 > 0). Fields and cost are averages per model.
Each cell shows how accurately a model extracted that field across all 11 documents. Click a cell for details.
Exact-match accuracy among functional models (F1 > 0).
Average exact-match accuracy across functional models — table and chart share the same rows.
Each cell shows accuracy for a specific document/model pair.
Average exact-match accuracy across functional models (error rows excluded from the average). Table and chart share the same rows.
Exact-match correct / wrong / missing shares — same order and values as the chart.
Same models and percentages as the table (scroll to see all).
Active models only — same aggregates as Rankings → Provider Summary (Avg/Best F1).
How models are distributed across pricing tiers per provider.
Higher is better — F1 score achieved per dollar spent (active models with cost data).
Text-only vs multimodal model distribution and performance per provider.
All configured models with capabilities, pricing, and performance. Filter by provider.
How this extraction benchmark evolved — from raw data preparation to a multi-provider, multi-model scored leaderboard.
From extraction_stats.csv — actual observed time and cost per model (where available).