Generative AI is very good at producing plausible state.

That is not the same as authoritative state.

A model can extract a deadline from a document, classify a transaction, summarize a policy, infer a category or propose a structured record. Those capabilities are useful. But the moment generated output is written directly into authoritative application data, a different question appears:

What gave this output the right to become true inside the system?

That is not primarily a hallucination question. It is a question about authority, evidence, provenance and system state.

Generation and Authority Are Different Operations

A traditional deterministic system often has explicit authority.

A customer submits a form. A payment network returns a settlement result. A database transaction commits. A signed document establishes an obligation. A rule engine evaluates a known policy.

Generative models operate differently. They produce output based on probabilistic inference over inputs and learned representations.

The output may be highly accurate. It may even be more useful than a human's first interpretation. But accuracy alone does not define authority.

A system needs to distinguish:

Source evidence
  ↓
AI interpretation
  ↓
Candidate state
  ↓
Validation
  ↓
Policy decision
  ↓
Canonical state

The candidate state is where AI becomes operationally useful without being granted unchecked authority.

Canonical State Is a Product Concept

"Canonical" means more than "stored in the database."

Canonical state is the version of information the application treats as authoritative for a particular purpose.

A competition platform may ingest a public page and use AI to extract:

  • registration dates;
  • eligible grades;
  • categories;
  • stages;
  • fees;
  • result dates.

The extracted values are useful candidates.

But what if the registration deadline appears in two documents with different dates? What if the model confuses the publication date with the event date? What if a PDF from last year ranks higher in retrieval than the current page? What if the organizer changes a date after extraction?

The system needs an epistemic model, even if the UI never uses that term.

It needs to know what it knows, where that knowledge came from, how confident the process was, and what evidence can override it.

Provenance Should Travel With the Data

Provenance is the connection between a claim and the evidence from which it was derived.

Without provenance, AI-assisted data becomes difficult to inspect and correct.

A useful extracted field might include:

field: registration_deadline
candidate_value: 2027-03-15
source_id: official_rules_pdf_2027
source_location: page 4
extraction_method: model-x
model_version: 2026-09
confidence: 0.93
validation_status: pending

The exact schema will vary. The principle is durable.

The system should be able to answer:

  • Which source produced this value?
  • Which model or process interpreted it?
  • Was the value directly extracted or inferred?
  • Has it been validated?
  • Was it later superseded?
  • Who or what approved it?
  • Can the system reconstruct why the canonical value exists?

NIST's Generative AI Profile extends the AI Risk Management Framework with guidance for risks specific to generative systems. The broader lesson is that trustworthy AI is a system property, not simply a model property. Evaluation, governance, context and controls around the model matter.

Confidence Is Not Permission

Confidence scores can be useful, but they are easy to misuse.

A score of 0.96 does not mean the output should automatically become canonical. Confidence may not be calibrated. It may describe token probability rather than factual correctness. It may not account for source quality. It may not capture whether the model answered the right question.

Permission to promote a candidate should depend on policy.

For a low-risk classification, policy may say:

confidence >= threshold
AND source_type == trusted
→ auto-accept

For a high-risk decision, policy may require deterministic validation or human approval regardless of confidence.

The important design move is separating model output from promotion policy.

That allows the system to change models without silently changing the organization's definition of truth.

Source Quality Matters More Than Fluency

A model can confidently extract from a weak source.

That creates a subtle failure mode: good inference over bad evidence.

Consider a system collecting regulatory obligations. A model retrieves a secondary blog that summarizes a regulation. The extraction is perfect. The summary is outdated.

The model has not hallucinated. The system is still wrong.

This is why provenance needs source policy.

Useful source attributes may include:

  • first-party versus third-party;
  • official versus commentary;
  • current versus archived;
  • jurisdiction;
  • effective date;
  • publication date;
  • revision;
  • trust tier;
  • access scope.

AI systems that govern only the model and ignore the source layer have incomplete controls.

Disagreement Is Information

When two authoritative sources disagree, the system should not always force immediate resolution.

Disagreement can be a meaningful state.

A robust pipeline can represent:

Source A → Candidate X
Source B → Candidate Y
                 ↓
          Conflict detected
                 ↓
        Review or policy rule
                 ↓
          Canonical value

This is more honest than asking a model to choose whichever value seems most plausible and discarding the conflict.

The unresolved disagreement may matter operationally. It may indicate a publication error, a changed policy, stale documentation or a genuine exception.

Preserving disagreement is part of preserving evidence.

Human-in-the-Loop Is Not a Universal Answer

"Add human review" is often the first governance response to AI uncertainty.

Human review can be appropriate, but it is not free and it is not automatically reliable.

Review creates:

  • cost;
  • latency;
  • inconsistency;
  • reviewer fatigue;
  • throughput limits;
  • new authorization requirements.

The better question is which decisions deserve which controls.

A risk-based model may look like:

WorkflowExampleAppropriate control
Low risk, reversiblecontent taggingautomated validation
Medium riskstructured extraction for internal reviewthresholds + spot review
High impactfinancial classification that changes customer outcomedeterministic checks + human approval
Safety/regulatoryaction with legal or safety consequencestrong policy, evidence and authorized review

The system should spend human attention where uncertainty and consequence justify it.

Deterministic Controls Still Matter

Generative AI does not replace ordinary software validation.

If a field must be a date in the future, validate it.

If a category must belong to an allowed vocabulary, enforce it.

If a total must equal the sum of line items, calculate it.

If an action requires permission, authorize it outside the prompt.

If a record transition is forbidden, enforce the state machine.

The model can propose. The application still governs.

This distinction is especially important because prompts are not access-control systems. A prompt that says "do not show confidential information" is not a substitute for permission-aware retrieval and enforcement.

Model Changes Are Production Changes

A team can change an AI model without changing application code and still change system behavior materially.

That means model version, prompt version, retrieval configuration and validation policy belong in production change management.

Questions include:

  • Did extraction accuracy change?
  • Did classification distribution change?
  • Did latency or cost change?
  • Did failure behavior change?
  • Did a new model become more willing to infer missing values?
  • Are previous evaluations still representative?

Versioning makes these questions answerable.

Without versioning, a team may know that "the AI got worse" without being able to reproduce when or why.

Evaluation Needs Production Semantics

Generic model benchmarks rarely answer the product question.

A document extraction system needs evaluation based on its own fields, sources and error costs.

False positives and false negatives may not be symmetric. Missing a registration deadline may be inconvenient. Inventing a deadline may actively mislead users.

Evaluation should therefore connect model behavior to workflow consequence.

A useful evaluation set includes:

  • representative source documents;
  • known difficult cases;
  • conflicting sources;
  • missing fields;
  • historical versions;
  • adversarial or malformed inputs;
  • permission boundaries;
  • expected abstention behavior.

The goal is not simply "accuracy." The goal is confidence that the system behaves acceptably in the states that matter.

AI Is a Component in a Larger Knowledge System

The most useful architecture often treats the model as one stage rather than the center of the product.

Source
  ↓
Acquisition
  ↓
Normalization
  ↓
AI Extraction
  ↓
Candidate
  ↓
Validation
  ↓
Review / Policy
  ↓
Canonical State
  ↓
Downstream use

This structure creates places to observe, test and improve the system.

It also allows components to evolve independently. A new model can improve extraction without changing canonical policy. A new source can enter the pipeline without bypassing validation. A human correction can become feedback for evaluation without rewriting history.

This is also why Why Product Decisions Need Memory matters in AI-assisted work. Context, evidence and provenance need to survive beyond a single model interaction.

Canonical Truth Is a Governance Decision

The practical mistake is not using generative AI in authoritative workflows.

The mistake is failing to design the promotion path from generated candidate to authoritative state.

A strong system makes the boundary visible:

  • AI output is a proposal.
  • Evidence supports or weakens the proposal.
  • Validation checks structure and invariants.
  • Policy decides which controls apply.
  • Humans intervene where consequence and uncertainty justify it.
  • Canonical state records the accepted result and its provenance.
  • Versioning preserves how the decision was made.

The model is powerful precisely because it can operate in spaces where deterministic software is weak: ambiguous language, messy documents, classification and inference.

That power becomes safer and more useful when the system does not confuse plausibility with authority.

AI output can be valuable knowledge.

It should earn the right to become canonical truth.

References

  1. NIST. Artificial Intelligence Risk Management Framework: Generative Artificial Intelligence Profile. 2024. https://www.nist.gov/publications/artificial-intelligence-risk-management-framework-generative-artificial-intelligence
  2. NIST. Artificial Intelligence Risk Management Framework (AI RMF 1.0). 2023. https://www.nist.gov/itl/ai-risk-management-framework
← All Insights