Model Extraction Playground

Interactive analysis of LLM performance on charity document data extraction

Model Leaderboard

F1 Score Overview

Provider Summary

Aggregated stats per provider (active models only, F1 > 0). Fields and cost are averages per model.

Scoring Methodology

Field-Level Accuracy Heatmap

Each cell shows how accurately a model extracted that field across all 11 documents. Click a cell for details.

Best Model per Field

Field Difficulty Ranking

Average accuracy across all functional models

Document × Model Accuracy Heatmap

Each cell shows accuracy for a specific document/model pair.

Document Difficulty (avg accuracy across functional models)

Error Category Breakdown per Model

Correct
Wrong value
Missing

Common Error Patterns

Per-Document Field Comparison

Decision Helper

Key Insights

Improvement Suggestions

Project Evolution Timeline

How this extraction benchmark evolved — from raw data preparation to a multi-provider, multi-model scored leaderboard.

Numbers at a Glance

Cost & Speed Summary

From extraction_stats.csv — actual observed time and cost per model (where available).