Model Extraction Playground

Interactive analysis of LLM performance on charity document data extraction

Model Leaderboard

F1 Score Overview

Provider Summary

Aggregated stats per provider (active models only, F1 > 0). Fields and cost are averages per model.

Scoring Methodology

Field-Level Accuracy Heatmap

Each cell shows how accurately a model extracted that field across all 11 documents. Click a cell for details.

Best Model per Field

Field Difficulty Ranking

Average accuracy across all functional models

Document × Model Accuracy Heatmap

Each cell shows accuracy for a specific document/model pair.

Document Difficulty (avg accuracy across functional models)

Error Category Breakdown per Model

Correct
Wrong value
Missing

Common Error Patterns

Per-Document Field Comparison

Decision Helper

Key Insights

Improvement Suggestions

Provider Comparison

Side-by-side comparison across all providers (active models only).

Tier Distribution by Provider

How models are distributed across pricing tiers per provider.

Cost Efficiency (F1 per $)

Higher is better — F1 score achieved per dollar spent (active models with cost data).

Modality Breakdown

Text-only vs multimodal model distribution and performance per provider.

Model Catalog

All configured models with capabilities, pricing, and performance. Filter by provider.

Project Evolution Timeline

How this extraction benchmark evolved — from raw data preparation to a multi-provider, multi-model scored leaderboard.

Numbers at a Glance

Cost & Speed Summary

From extraction_stats.csv — actual observed time and cost per model (where available).