Fixed-scope paid diagnostic · one production decision

Production AI Architecture Review

Make one defensible production decision about one existing AI agent workflow. I reconstruct how the system behaves from available evidence, identify the causal failure paths that matter to that decision, and give the accountable team a sequenced remediation path with acceptance evidence.

This is not a general review of whether an architecture is “good.” It is a bounded diagnostic for a decision such as moving a claims-triage workflow from a supervised pilot to limited production traffic, or deciding whether an existing agent can safely expand its authority, tenancy, data access, or autonomy.

When the review fits

The primary buyer is a CTO, VP or Head of Engineering, or Head of AI accountable for a named agent workflow and a consequential production decision in the next 30–60 days.

  • There is one named system and workflow in prototype, pilot, or production.
  • There is one consequential decision, a deadline, and an accountable owner.
  • The implementation and its operating evidence can be inspected.
  • A technical or operating owner can join the intake and working handoff.

When it does not fit

  • The work is still only an idea, or the question is enterprise-wide AI transformation.
  • The team cannot yet name the workflow, decision, owner, deadline, or available evidence.
  • The desired output is a vendor shortlist, general strategy, certification, or guarantee.
  • The required work is a penetration test, compliance audit, load test, line-by-line code review, or implementation.

If the system is too broad for the delivery window, I narrow the production decision rather than silently expand the engagement.

Seven production seams the review follows

The lenses are prompts for tracing real trajectories, not seven separate audits or a promise of complete coverage.

  1. Purpose, authority, and autonomy. Permitted and prohibited decisions and actions, actor identity, delegated authority, approvals, and reversibility.
  2. Orchestration, tools, and external effects. Termination, budgets, retries, idempotency, checkpoints, cancellation, typed contracts, least privilege, and attributable writes.
  3. Context, state, memory, and data. Provenance, freshness, typing, scope, tenancy, versioning, retention, forgetting, retrieval evidence, and schema change.
  4. Models and control plane. Eligibility, routing, fallback, structured-output validation, prompt and configuration versions, latency budgets, and spend controls.
  5. Evaluation, observability, and release evidence. Production-shaped scenarios, replay, correlated trajectories, drift, release gates, cost attribution, and reconstruction.
  6. Security, privacy, and governance. Identity, authorization, trust boundaries, untrusted content, secrets, egress, retention, policy exceptions, and approvals.
  7. Reliability, recovery, and ownership. Degraded dependencies, blast radius, rate limits, kill paths, exception queues, rollback, incident recovery, change control, and accountable operators.

Every material conclusion is labelled verified, asserted, inferred, or unknown. Severity and confidence remain separate: a high-severity, low-confidence risk becomes an urgent evidence request or bounded test, not a falsely certain finding.

How the review works

  1. Agree the charter. Name the system, workflow, production decision, intended operating conditions, owner, evidence window, and exclusions.
  2. Index the evidence. Record the source, version, environment, date, and access boundary of supplied material.
  3. Reconstruct the system. Compare the as-stated architecture with what the artifacts and runtime evidence demonstrate.
  4. Trace critical trajectories. Follow one representative success path and at least one consequential, exceptional, or failed path.
  5. Test decision-changing hypotheses. Use only approved, bounded, non-destructive probes.
  6. Write atomic findings. Express each as condition → trigger → failure → consequence → control gap, with evidence state, severity, confidence, and affected boundary.
  7. Sequence remediation. Contain unbounded effects, make the system observable and recoverable, then improve capability, autonomy, scale, or economics.
  8. Record and hand off the decision. State the verdict, unknowns, accepted residual risks, expiry or next-review trigger, and accountable owners.

See every stage, input, output, and decision gate →

Six written outputs

  1. Review charter and evidence index. The exact decision, scope, exclusions, artifacts inspected, provenance, and material gaps.
  2. As-evidenced architecture and critical trajectories. Trust boundaries, authority, state and data flow, external effects, operating path, and differences from the stated model.
  3. Prioritized atomic findings. Causal failure paths with evidence state, severity, confidence, existing controls, blast radius, and detection and recovery implications.
  4. Sequenced remediation register. Required capabilities, immediate containment, structural changes, dependencies, owner roles, effort bands, residual risk, and acceptance evidence.
  5. Bounded executive verdict. One of proceed, proceed with conditions, do not proceed or expand yet, or insufficient evidence.
  6. Working handoff and decision log. Evidence, disagreements, first actions, owners, decision dates, unknowns, verdict expiry, and next-review trigger.

Recommendations name a required capability and its acceptance evidence, not a vendor shopping list. “Add guardrails” is not a useful recommendation; a deterministic test for preventing a duplicated external write is.

Evidence needed from the team

I ask for material that already exists, not newly polished documents:

  1. A system summary covering users, intended outcome, lifecycle stage, environment, allowed actions, decision owner, deadline, and current blocker.
  2. An architecture view with components, dependencies, product and model versions, and relevant trust boundaries.
  3. One representative end-to-end trajectory, including context, model decisions, tool calls, writes or external effects, final state, and available latency or cost evidence.
  4. At least one failed, exceptional, degraded, unauthorized, interrupted, or otherwise high-consequence trajectory.
  5. Contracts and controls: tool and API schemas, identity and authorization, prompt and policy versioning, human approvals, and exception handling.
  6. Evaluation and operating evidence: evals and release gates, traces, deployment and rollback material, recovery procedures, runbooks, and ownership.

The delivery clock starts only after the charter is agreed and the evidence pack is ready enough to support the decision. Missing evidence may become a finding. If it prevents a responsible conclusion, the verdict is insufficient evidence or the start date changes.

Do not send passwords, tokens, private keys, unrestricted production data, or unnecessary personal data by ordinary email. We agree a secure transfer mechanism before restricted evidence is shared.

Timing, scope, and commercial boundary

Delivery
Three to four working days after evidence readiness.
Scope unit
One system, one workflow, one production decision.
Commercial posture
A fixed-scope paid engagement, quoted after the fit check. There is no public price.
Scope change
A broader or materially different system, workflow, decision, evidence environment, or requested test requires a narrower boundary, schedule change, or separate scope.
Follow-on work
Remediation implementation, a reference build, or embedded enablement is separately proposed only after the verdict. None is bundled or implied.

What the review does not claim

  • It is not enterprise AI strategy, a transformation roadmap, or a generic best-practice workshop.
  • It is not a complete risk inventory or a guarantee that no failure remains.
  • It is not a penetration test, certification, compliance audit, privacy or legal opinion, load test, benchmark, or line-by-line code review.
  • It is not proof of production reliability, security, performance, regulatory compliance, or client outcome.
  • It does not include implementation, managed operations, staff augmentation, or vendor selection.
  • It is not authorization to deploy or expand. The accountable client owner retains that decision.

No hosted or client work begins on verbal interest alone. Scope, access, confidentiality, and secure evidence transfer must be agreed first.

The decision trail

  1. Question

    Can one named AI workflow proceed, expand, or enter production under its intended operating conditions?

  2. Decision

    Reach one bounded verdict. Proceed, proceed with conditions, do not proceed or expand yet, or insufficient evidence.

  3. Evidence

    Specified. Separate verified artifacts, assertions, inferences, and unknowns; keep severity distinct from confidence.

  4. Implementation

    Specified. Findings become required capabilities, acceptance evidence, accountable owners, and a sequenced remediation register. This describes the method; it does not claim a client outcome.

Next step

Request a Production AI Architecture Review

Send the seven fields below. I will reply with a direct fit/no-fit view and, if it fits, the next intake step. Do not include secrets or unrestricted production data.

  1. What the team is building
  2. Lifecycle stage: prototype, pilot, or production
  3. The one production decision to make
  4. What is currently blocking that decision
  5. The consequence of getting it wrong
  6. Evidence currently available
  7. Decision owner and deadline