Evaluation & safety · Glossary term
What is Pass@k?
Across a task set, the fraction of tasks for which at least one of k sampled candidates passes a defined correctness test.
Why does Pass@k matter?
It measures the value of sampling several attempts for tasks such as code generation where an automatic verifier can check each candidate.
Pass@k in practice
Generate candidates independently under a fixed configuration, run the same isolated tests on each, and report k with the sampling and estimator details.
What is the common confusion about Pass@k?
Pass@k is not single-attempt accuracy, and a higher score can reflect a larger attempt budget rather than a better first answer.
Learn Pass@k in the course
No lesson links to this term yet. Search the course catalog for it.
Related terms
- Coding AgentAn agent specialized for software work that can inspect a repository, edit files, run development tools, and use their outputs to advance…
- Regression TestA repeatable check that protects behavior known to work, especially after code, prompt, model, retrieval, or tool changes.
- Eval SetA versioned collection of inputs, expected properties, scoring rules, and metadata used to measure an AI system against a defined…
- Test OracleThe mechanism, specification, reference, invariant, or human judgment used to decide whether observed program behavior is correct.
- Exact Match (EM)A metric that counts an output as correct only when its normalized representation exactly equals an accepted reference answer.
Sources
More terms in Evaluation & safety
This entry comes from glossary/terms.md on GitHub. Browse all 250 glossary terms.