← All articlesPRACTICAL KNOWLEDGE

Defining quality criteria for AI outputs

AI quality is task-specific. An output is not generally “good”; it may be accurate, complete, grounded, relevant, consistent, safe, understandable and conformant to a required format. The article provides a controlled method, a realistic CTPM practice example and a concrete transfer artefact.

Realistic enterprise scene illustrating Defining quality criteria for AI outputs
Short answer

AI quality is task-specific. An output is not generally “good”; it may be accurate, complete, grounded, relevant, consistent, safe, understandable and conformant to a required format.

What the concept actually means

Criteria must be observable. “Professional” is too vague; “contains all five mandatory fields and every number has a source reference” is testable.

Why it matters in the enterprise

Enterprises combine automated checks, reference cases, domain review and operational feedback. Not every dimension should be collapsed into one overall score.

A controlled method

The CTPM practice framework for controllable AI applications uses seven stages: understand the task, clarify context and data, apply AI deliberately, review professionally, handle deviations, approve accountably and document transfer. It is a transparent working framework, not a certification.

  • Define task and impact
  • Clarify data, context and permissions
  • Review against domain criteria
  • Control deviations, approval and evidence

CTPM practice example

CTPM practice example: A summary is evaluated for complete decisions, correct assignment of owners, date format, source grounding and absence of newly invented claims.

Quality and test criteria

The following criteria make quality observable for this use case:

  • Criteria have definitions, methods and thresholds.
  • Critical errors are handled separately rather than averaged away.
  • Reviewers calibrate using shared examples.
  • Metrics are checked against real failure consequences.

Risks and common misconceptions

Popular but unsuitable metrics may optimise behaviour that conflicts with the business goal. An average can conceal rare critical failures.

Example transfer artefact

Transfer artefact: a quality rubric with dimension, weight, critical failure, measurement method, threshold and owner.

Sources and references

  1. NIST: Artificial Intelligence Risk Management Framework (AI RMF 1.0) (2023)
  2. NIST: Artificial Intelligence Risk Management Framework: Generative Artificial Intelligence Profile (2024)
  3. OpenAI: Working with evals (2026)