Glossary · DISCRY METHODOLOGY

Closed-book baseline

A closed-book baseline runs the same measurement tasks with no documentation provided, establishing what a model already knows about an API from training. Whatever a model can answer closed-book is excluded from measurement, so results reflect what the documentation itself contributes. Without this control, a famous API would score well even if its live docs were empty — the model answers from memory.

Frontier models have read the famous APIs' documentation in training, which makes memory the central confound for any documentation benchmark. The closed-book baseline is the control that removes it: by first recording what a model answers with no docs at all, the measurement can separate knowledge the docs delivered from knowledge the model walked in with.

The baseline is also what makes the resulting score fair in both directions. A famous API earns nothing from being famous, and an obscure API loses nothing for being obscure — what remains after the closed-book exclusion is performance the documentation had to earn.

How Discry measures this

In the Discry methodology, every task first runs closed-book: anything a model can answer without the docs is excluded from the eligible set, and only genuinely doc-dependent tasks count toward the score. In the extreme case where models know an API's docs essentially by heart, there is nothing left to measure, and the profile publishes as insufficient coverage rather than being handed an unearned grade.

What does an AI agent make of your API?

Find out in about a minute — no signup.

Discry your API — free