Evaluation & safety · Glossary term

What is Pass@k?

Across a task set, the fraction of tasks for which at least one of k sampled candidates passes a defined correctness test.

Why does Pass@k matter?

It measures the value of sampling several attempts for tasks such as code generation where an automatic verifier can check each candidate.

Pass@k in practice

Generate candidates independently under a fixed configuration, run the same isolated tests on each, and report k with the sampling and estimator details.

What is the common confusion about Pass@k?

Pass@k is not single-attempt accuracy, and a higher score can reflect a larger attempt budget rather than a better first answer.

Learn Pass@k in the course

No lesson links to this term yet. Search the course catalog for it.

  • Coding AgentAn agent specialized for software work that can inspect a repository, edit files, run development tools, and use their outputs to advance…
  • Regression TestA repeatable check that protects behavior known to work, especially after code, prompt, model, retrieval, or tool changes.
  • Eval SetA versioned collection of inputs, expected properties, scoring rules, and metadata used to measure an AI system against a defined…
  • Test OracleThe mechanism, specification, reference, invariant, or human judgment used to decide whether observed program behavior is correct.
  • Exact Match (EM)A metric that counts an output as correct only when its normalized representation exactly equals an accepted reference answer.

Sources

More terms in Evaluation & safety

Open the Evaluation & safety list in the glossary

This entry comes from glossary/terms.md on GitHub. Browse all 250 glossary terms.