llmcoderlab

Models

Every model that has run in the lab, with latest-run scores per mode. Meters show the mean blended score over graded attempts in the most recent run.

glm4-9b
oneshot87
agentic49.4
1
runs
15/18
graded
3
failures

full profile →

llama3.2-3b
oneshot91.4
agentic77.9
1
runs
16/18
graded
2
failures

full profile →

qwen2.5-coder-3b
oneshot87.2
agentic49.8
2
runs
21/28
graded
7
failures

full profile →

qwen3-30b-a3b
oneshot100
agentic100
1
runs
12/18
graded
6
failures

full profile →