← All articlesPRACTICAL KNOWLEDGE

Analysing documents with AI: extraction before interpretation

Dependable document analysis separates evidence location, extraction, interpretation and judgement. If these stages are mixed, plausible claims can no longer be traced to the original. The article provides a controlled method, a realistic CTPM practice example and a concrete transfer artefact.

Realistic enterprise scene illustrating Analysing documents with AI: extraction before interpretation
Short answer

Dependable document analysis separates evidence location, extraction, interpretation and judgement. If these stages are mixed, plausible claims can no longer be traced to the original.

What the concept actually means

Before analysis, document type, version, language, tables, annexes and confidentiality are clarified. Structured questions and an evidence-linked output format follow.

Why it matters in the enterprise

For repeated processing, ad-hoc sampling is insufficient. Representative test documents, known expected values and rules for unreadable or conflicting passages are required.

A controlled method

The CTPM practice framework for controllable AI applications uses seven stages: understand the task, clarify context and data, apply AI deliberately, review professionally, handle deviations, approve accountably and document transfer. It is a transparent working framework, not a certification.

  • Define task and impact
  • Clarify data, context and permissions
  • Review against domain criteria
  • Control deviations, approval and evidence

CTPM practice example

CTPM practice example: Obligations are extracted from 30 policies. Each row contains document, version, section, original quote, normalised obligation and uncertainty. Grouping occurs only afterwards.

Quality and test criteria

The following criteria make quality observable for this use case:

  • Completeness is measured against known evidence.
  • Quotes match exact wording and location.
  • Tables and annexes are reviewed separately.
  • Unreadable passages are not silently completed.

Risks and common misconceptions

Risks include OCR errors, wrong document versions, omitted tables, mixing commentary with normative text and unauthorised processing of confidential content.

Example transfer artefact

Transfer artefact: a document-analysis template with source inventory, extraction schema, review sample and approval rule.

Sources and references

  1. NIST: Artificial Intelligence Risk Management Framework: Generative Artificial Intelligence Profile (2024)
  2. OWASP GenAI Security Project: OWASP Top 10 for LLM Applications 2026 (2026)
  3. OpenAI: Working with evals (2026)