llmcoderlab

llama3.2-3b

1 run(s)16/18 graded2 failures

← all models

// score history

RunModeScore GradedFailures
run 2agentic77.99/90
run 2oneshot91.47/92

// every attempt