benchmark playgroup

UK Charity Document Extraction — Multi-Model Benchmark

Why
A playgroup needed a fair comparison of models on the same real charity PDFs, with the same fields and the same score.
Problem
UK charity annual reports are PDFs. Eight financial fields have to come out correctly, and models differ in accuracy, failed runs, time, and cost.
Built
A scored extractor across OpenRouter, the Doubleword Batch API, and V7 Go, plus a public playground of every run.
Result
122 scored runs. The best models reach about 0.95 F1. Provider averages and failures are on the leaderboard below.
repo: playgroup_202602_docextract dataset: Kleister Charity (11 PDFs) fields: 8
Open latest playground All 25 snapshots View on GitHub QUICKSTART

Results snapshot: 2026-09-18 · 122 scored runs · playground 2026-09-18T0710Z (25 archived versions from git history)

Who is this for?

Key findings

Provider aggregates match python score.py over every data/*_dev_extracted__*.tsv file on 2026-09-18: 122 scored runs (46 OpenRouter, 43 Doubleword, 33 V7 Go). Registries: 39 OpenRouter keys, 36 Doubleword models, 32 V7 keys.

Provider Models Active Failed Avg F1 Best F1 Best model Avg fields Avg time Avg cost
Doubleword 43 41 2 0.827 0.947 dw-deepseek-v4-pro-0813 65.1/85 (77%) 735s $0.067
OpenRouter 46 33 13 0.788 0.946 gemini-3-pro 59.9/85 (71%) 1030s $0.158
V7 Go 33 33 0 0.833 0.852 gpt4-1 62.8/85 (74%) 124s $0.000

Doubleword’s average F1 among active models (0.827) edges OpenRouter (0.788), and Doubleword retains the highest single-model F1 (dw-deepseek-v4-pro-0813, 0.947). gemini-3-pro ranks 2nd at 0.946; new entry dw-deepseek-v4.1-flash joins at rank 10 (F1 0.941, $0.032/run — highly cost-efficient). V7 averages 0.833 F1 with no failed runs in this snapshot. V7 cost stays at zero until price_in / price_out are set in config_models_v7.py.

Top 5 models by F1

Rank Model Provider F1 Precision Recall Fields found
1dw-deepseek-v4-pro-0813Doubleword0.9470.9750.92078.2/85
2gemini-3-proOpenRouter0.9460.9750.91878.1/85
3kimi-k3OpenRouter0.9440.9750.91677.8/85
4kimi-k2.6OpenRouter0.9440.9750.91577.8/85
5glm-5.3-flashOpenRouter0.9440.9750.91477.7/85

Takeaways

Interactive playground

The hosted playground is a static snapshot generated by python playground.py in the repo. Eight tabs: Rankings, Field Heatmap, Document Analysis, Error Breakdown, Deep Dive, Recommendations, Provider Analysis, and Project Evolution.

Two scoring views: Rankings and Provider Analysis use semantic F1 from score.py; heatmaps and error tabs use exact-match vs expected TSV. F1 can exceed exact-match when values are close but not identical — that is expected.

Open latest playground Browse all snapshots

Playground version history

Each row is a frozen copy of which-models-extracted-playground.html from git, named with its commit time in UTC. Use this to see how the benchmark grew from 40 OpenRouter-only models (Feb 2026) to 122 scored runs across three providers (Sep 2026).

Snapshot (UTC) Models F1 runs Highlights
2026-09-18T0710Z122122Latest — dw-deepseek-v4.1-flash (F1=0.941, rank 10)Open
2026-09-11T0000Z121121DW/OR catalog prices & cost backfillOpen
2026-09-10T2232Z121121OpenRouter Gemini, Kimi, GLM resultsOpen
2026-09-09T2357Z115115Four new Doubleword model resultsOpen
2026-08-22T0830Z111111dw-deepseek-ocr-2 OCR resultsOpen
2026-08-21T2219Z111111dw-qwen3.8-27b, dw-muse-glimmer-30bOpen
2026-08-21T2145Z109109Tab sync fix across all dashboardsOpen
2026-08-09T0746Z109109Aug 2026 benchmark refreshOpen
2026-07-29T1317Z10710734 Doubleword modelsOpen
2026-04-11T1748Z8585V7 Go provider addedOpen
2026-03-07T2116Z4747Provider Analysis tabOpen
2026-03-07T2104Z4747F1 scoring & provider summaryOpen
2026-02-28T1312Z40—First leaderboard (exact-match tabs)Open

View all 25 snapshots with search and filters →

Dataset & fields

Small export from the Kleister Charity dataset (dev-0): 11 UK charity financial PDFs (up to 200 pages each) with pre-extracted OCR text. Data is UK open data — see Kleister license.

charity_number charity_name report_date income_annually_in_british_pounds spending_annually_in_british_pounds address__postcode address__post_town address__street_line

Workflow (local)

Step-by-step video guide for Doubleword batch runs: extractor.py --all-doubleword.

# Smoke test
python llm_openrouter.py

# Extract (auto-detects OpenRouter / Doubleword / V7 from registries)
python extractor.py gemini-2.0-flash
python extractor.py --all-doubleword
python extractor.py --all-v7

# Score and regenerate playground
python score.py
python playground.py

Repo essentials

github.com/neomatrix369/playgroup_202602_docextract Back to home