benchmark poolside laguna hackathon
py-bug-trace — Laguna model sweep
Evaluation of poolside/laguna-xs.2 against 28 comparison models on the py-bug-trace environment from the Poolside Laguna hackathon. Models predict outputs of subtly broken Python (L1–L2) or fix buggy code (L3). All artifacts originate from the laguna-eval-experiments dataset on Hugging Face.
Headline finding: Laguna-XS.2 ranks 12th of 29 overall (79.0%) but
2nd on Level 3 (86.7%) — an inverted difficulty profile where the target is weak on easy tasks and strong on hard code-fix problems. Thirteen of 28 flagged anomalies are level inversions across the field.
Credit & source of truth (Hugging Face)
Pages on this site are hosted mirrors for navigation and readability. When citing results, link to the canonical Hugging Face artifacts below.
- poolside-laguna-hackathon/laguna-eval-experiments — main dataset repo
- environments/py_bug_trace/README.md — environment docs
- reports/README.md — pick-your-path guide
- reports/ — full reports tree
- poolside-laguna-hackathon/datasets — rollout datasets
Target model at a glance
| Level | Score | Rank (cohort) | Percentile |
|---|---|---|---|
| L1 — Python gotchas | 77.1% | 3 / 8 | 25th |
| L2 — Concurrency / asyncio | 73.3% | 4 / 8 | 54th |
| L3 — Code fix | 86.7% | 2 / 8 | 79th |
| Overall | 79.0% | 12 / 29 | — |
Hosted reports
Each page links back to its canonical Hugging Face artifact.
Interactive explorer
Filter models, levels, tasks
One-pager
30-second summary
Write-up
5-minute narrative
Technical report
Executive summary
Level scorecards
L1, L2, L3 breakdown
HF-only artifacts
- matrix/stats.md — stats digest
- matrix/report.md — full matrix report
- stats.json · comparison.csv