benchmark poolside laguna hackathon

py-bug-trace — Laguna model sweep

Why
The Poolside Laguna hackathon asked how laguna-xs.2 behaves on Python bugs next to a field of other models.
Problem
One overall score hides the shape of a model. Easy tasks and hard code-fixes are different jobs.
Built
A 29-model sweep on py-bug-trace, with an explorer, a short write-up, and per-level scorecards.
Result
12th of 29 overall (79%), and 2nd on the hardest level (86.7%): weaker on easy tasks, stronger when the job is fixing code.

Evaluation of poolside/laguna-xs.2 against 28 comparison models on the py-bug-trace environment from the Poolside Laguna hackathon. Models predict outputs of subtly broken Python (L1–L2) or fix buggy code (L3). All artifacts originate from the laguna-eval-experiments dataset on Hugging Face.

target: poolside/laguna-xs.2 29 models matrix: 2026-06-04 reports: 2026-07-17
Beyond the rank: thirteen of 28 flagged anomalies are level inversions across the field, so the weak-easy strong-hard pattern is not only Laguna-XS.2.
Credit & source of truth (Hugging Face)

Pages on this site are hosted mirrors for navigation and readability. When citing results, link to the canonical Hugging Face artifacts below.

Open interactive explorer Reports archive

Target model at a glance

LevelScoreRank (cohort)Percentile
L1 — Python gotchas77.1%3 / 825th
L2 — Concurrency / asyncio73.3%4 / 854th
L3 — Code fix86.7%2 / 879th
Overall79.0%12 / 29—

Hosted reports

Each page links back to its canonical Hugging Face artifact.

Interactive explorer Filter models, levels, tasks One-pager 30-second summary Write-up 5-minute narrative Technical report Executive summary Level scorecards L1, L2, L3 breakdown

HF-only artifacts