Back to archive
#ai#papers#aigen#llm#agents

pass@k

The model produced ten programs, and one passed the tests. Does that mean a single response is likely to succeed? You need to distinguish “it worked on the first attempt” from “we found a success among many attempts”.

pass@k measures the chance that at least one of k attempts at a task passes a specified check. Results are summarized across tasks. The number k is part of the metric's meaning because it determines the attempt budget.

In a pool of four attempts with one success, the pass@1 estimate is 1/4. Of the six possible pairs, three contain a success, so the pass@2 estimate from this pool is 1/2. Do not confuse the number of attempts collected for estimation with the number allowed in the situation being measured.

A high pass@10 does not promise an equally good pass@1. Nor does it tell us whether the user can identify the correct attempt without a tester. Chen and colleagues, §3.1, describe pass@k estimation; the paper on agents uses this metric in a different context.