
A retrieval corpus is the body of documents a system searches at query time to ground its answer, as distinct from the data the underlying model was trained on.
Retrieval happens live. The corpus can include indexed web content, licensed databases, uploaded documents or a specific curated set, depending on the system.
Retrieval corpus vs. training data
Training data is historical and fixed at a point in time. The retrieval corpus is queried in the moment, which is why a document published this week can shape an answer given today by a model trained a year ago.
Why it matters
This is the mechanism by which current, well-structured third-party research outweighs a decade of a vendor's own marketing.
Volume of published material confers very little advantage here. What matters is whether a document answers the question being asked, clearly enough to be extracted, from a source the system treats as credible — which describes analyst research almost exactly.
Related: Grounding · Citation surface · Answer engine · Machine-readable claim
Ready to elevate your Product Marketing & Analyst Relations strategy
Contact us to learn more about how we can help you accelerate your business success.


