Security agents should earn containment authority separately from detection accuracy

Agentic AI SeedlingPlanted Sep 2026

I treat threat detection and containment as two different production contracts. A security agent that correctly identifies a compromised endpoint has demonstrated judgment about evidence; it has not yet demonstrated the right to isolate that endpoint, disable an account, or terminate a process. Joining those claims into one “accuracy” score turns a model evaluation into an authorization grant.

The separation matters because the error costs are asymmetric. A false positive in an analyst queue consumes attention. The same false positive at the containment layer can interrupt payroll, lock out an administrator during an incident, or destroy the volatile state needed for investigation. Detection can be optimized around precision, recall, ranking, and time to triage. Containment needs a different test suite: target identity, action scope, reversibility, stale-state handling, concurrent change, evidence preservation, and recovery time after a mistaken intervention.

I would therefore design a security agent as an authority ladder. The first rung observes and enriches. The next recommends a response with the evidence and uncertainty attached. A higher rung prepares an exact action against an exact resource. Only the final rung executes, after policy checks the principal, target, permitted effect, and current incident state. This is the security-specific form of putting guardrails at the action-decision layer: a good detector placed after the effect is merely a commentator.

The Vidhyarthi security-agent architecture makes the boundary concrete. Alert-triage agents ingest, enrich, prioritize, and summarize. Incident-response agents may isolate systems, disable accounts, remove persistence, rotate credentials, and restore known-good state. Those are not adjacent confidence levels on one scale. They are different capabilities with different blast radii. The response path should consume a signed decision record containing the evidence snapshot, target identifiers, proposed action, policy result, approval state, and expiry. If any authority-bearing input changes, the system should decide again rather than replay an old verdict.

Human approval is useful, but “put a human in the loop” is too vague to be a control. The approval interface should show the affected hosts, identities, network segments, expected service impact, reversibility, and evidence age — the same reason approval UX should communicate blast radius. Low-risk actions can be pre-authorized by policy. High-impact actions should require named approvers or two-person control. The agent should never infer broader scope from a hurried click.

Containment also needs a technical boundary beneath the workflow. A restricted execution environment can enforce target allowlists, egress limits, capability allowlists, resource budgets, immutable action logs, and clean reset. That is why I size the sandbox to the blast radius. The runtime must reject an out-of-scope command even if the model, detector, and approver all make the same mistake. Audit records then describe both the decision and the enforcement result, not just the agent’s narration.

I concede that fully automated containment is justified for narrowly specified, highly reversible actions where target identity is strong, state is fresh, the policy is explicit, and recovery is tested — quarantining a disposable workload under a known compromise signature may fit. That still does not let detection accuracy stand in for containment readiness. It defines the bounded conditions under which authority has been earned.

The production question is not “How accurate is the security agent?” It is “Which evidence may trigger which effect, against which target, under whose authority, with what recovery path?” I want the detector to improve quickly and the containment boundary to change deliberately. Keeping their contracts separate lets both happen without making every model promotion an undeclared expansion of operational power.