nDCG@10
A search engine returns ten results. Good documents may be on the list, but if the most useful is only tenth, the reader may never reach it. You need a measure accounting for both relevance and position.
nDCG@10 measures the quality of a ranking's first ten positions against available relevance judgments. A useful result contributes more at the beginning of the list than at the end. The sum of these contributions is divided by the score for the best possible ordering for the same query.
If there is one highly relevant document, moving it from first to tenth worsens the score. Swapping two documents with identical relevance judgments does not change it. In the usual correctly defined case, 1 means ideal ordering. The scikit-learn documentation explains NDCG normalization and limitations.
The metric does not say whether numbers assigned to documents are well calibrated for a specific threshold. The same ranking can therefore produce different “keep/discard” decisions. The article on score dependence on order examines this difference.