benchmark poolside laguna hackathon

py-bug-trace — Laguna model sweep

Evaluation of poolside/laguna-xs.2 against 28 comparison models on the py-bug-trace environment from the Poolside Laguna hackathon. Models predict outputs of subtly broken Python (L1–L2) or fix buggy code (L3). All artifacts originate from the laguna-eval-experiments dataset on Hugging Face.

target: poolside/laguna-xs.2 29 models matrix: 2026-06-04 reports: 2026-07-17
Headline finding: Laguna-XS.2 ranks 12th of 29 overall (79.0%) but 2nd on Level 3 (86.7%) — an inverted difficulty profile where the target is weak on easy tasks and strong on hard code-fix problems. Thirteen of 28 flagged anomalies are level inversions across the field.
Credit & source of truth (Hugging Face)

Pages on this site are hosted mirrors for navigation and readability. When citing results, link to the canonical Hugging Face artifacts below.

Open interactive explorer Reports archive

Target model at a glance

LevelScoreRank (cohort)Percentile
L1 — Python gotchas77.1%3 / 825th
L2 — Concurrency / asyncio73.3%4 / 854th
L3 — Code fix86.7%2 / 879th
Overall79.0%12 / 29

Hosted reports

Each page links back to its canonical Hugging Face artifact.

Interactive explorer Filter models, levels, tasks One-pager 30-second summary Write-up 5-minute narrative Technical report Executive summary Level scorecards L1, L2, L3 breakdown

HF-only artifacts