Glossary · DISCRY METHODOLOGY

Trap task

A trap task asks a model to do something the API does not support; the correct answer is a refusal. It measures whether documentation prevents confident hallucination — given the docs, a model should recognize the capability is absent and say so, rather than inventing endpoints or parameters. Documentation that fails its trap tasks is documentation agents will hallucinate against in production.

The trap task models a failure that is already showing up in support queues: a developer asks a coding agent to integrate an API, the agent scaffolds against endpoints that do not exist, and the resulting ticket blames the API. The hallucination happened while the agent was reading the docs — which makes it a documentation property worth measuring, not only a model flaw.

Passing trap tasks is a statement about the docs, not just the model. Documentation that states its boundaries clearly — what the API does and, implicitly or explicitly, what it does not — gives a model the ground to refuse. Vague or aspirational docs leave a gap the model fills with invention.

How Discry measures this

Trap tasks are part of Discry's published task banks, run alongside factual and request-construction tasks with the fetched docs as the model's only source. Grading is mechanical: a refusal passes, a confident invention fails. Because the tasks probe for capabilities the API genuinely lacks, a failure decodes to something verifiable — this documentation, read by a model, produced a fabricated capability.

What does an AI agent make of your API?

Find out in about a minute — no signup.

Discry your API — free