Data foundations for AI

Our enterprise data is not ready for AI use

AI-ready data is modeled and governed data, not raw access plus embeddings.

A model can read a table without understanding which customer identity is authoritative, what “revenue” includes, whether a value is current, or which downstream action may rely on it. Retrieval makes data available; it does not make the meaning explicit or the use defensible.

For one consequential AI use case, readiness means the system can identify the data boundary, interpret it consistently, detect material change, trace where an answer came from, and name who owns correction and retirement. Embeddings and larger context windows cannot repair ambiguous or ungoverned source meaning.

See the Data Foundations for AI pillar →

Five decisions before the model consumes the data

  1. AI-readable meaning: Define the entities, measures, identifiers, grain, time semantics, and permitted uses the model must interpret without relying on tribal knowledge.
  2. Data contracts: Make schema, quality, freshness, and change expectations executable at producer-consumer boundaries, with an explicit response to violations.
  3. Semantic layer: Give business terms and metrics one governed definition that can be resolved to authoritative data and cited in an answer or action.
  4. Data products: Name the owner, consumers, service expectations, provenance, access scope, and lifecycle for the bounded source—not just the pipeline that moves it.
  5. Production evidence: Test representative distributions, identity resolution, lineage, temporal correctness, exception handling, and change behavior against the decision the AI system will make.

Read the data-readiness path

Proof and limits


Assess one use case, one decision, and one data boundary

A bounded data-readiness assessment identifies the contract, semantic, provenance, identity, quality, and lifecycle gaps that block one production decision. It uses existing representative artifacts rather than asking for a broad data-platform transformation.

In your note, include: the single AI use case; the production decision you need to make; the named data product or source boundary; known contract, semantic, provenance, or quality gaps; representative artifacts that are available; and the decision deadline. Do not send secrets or unrestricted production data by ordinary email.