Agent capability platforms should expose policy surfaces, not catalog size
An enterprise agent platform should be judged by the policy surfaces it exposes, not the number of agents, connectors, models, or skills in its catalog. Capability abundance is easy to demonstrate. The production question is whether an operator can state who may use a capability, against which data, under what conditions, with what budget, and how that decision will be reconstructed later.
The platform landscape makes the choice clear. Vendors emphasize different layers — shared business context, governance control planes, managed runtimes, or developer experience. Those are meaningful architectural bets. A catalog count is not. Ten thousand actions without coherent identity and policy create ten thousand ways for ambient authority to escape review.
I want every advertised capability to expose an enforceable contract: principal, action, resource, context, data classification, side-effect class, cost envelope, approval rule, and revocation state. Action-centric governance answers whether an agent may invoke a tool. Data-centric governance answers whether the resulting information may be accessed, retained, or transmitted. Neither is sufficient alone. The platform earns trust when both policies meet at the invocation and egress boundaries.
This is why policy-as-code is how enterprises say yes to agents. Declarative rules can be versioned, tested, reviewed, and evaluated before credentials reach the runtime. A visual builder that configures an agent but sends operators elsewhere to understand its effective permissions is not a control plane. It is a creation interface attached to an authorization mystery.
Discovery must also be downstream of eligibility. A registry should index revocation before capability, because a disabled identity or compromised trust path should disappear before ranking begins. Then routing can buy verified capability using execution evidence, current policy, and operating history rather than self-described skill labels.
The operator experience matters because controls hidden in configuration do not survive incidents. A useful policy surface answers concrete questions quickly: which agents can write to this system, which policy granted access, what changed since yesterday, which runs used the old rule, and what will break if access is revoked now? Denials should be visible alongside permits, and simulations should show the blast radius before promotion. Governance that leaves the admin console does not exist for operators.
Composability raises the bar. If an agent is built in one ecosystem, deployed in another, and governed by a third, effective policy cannot live as undocumented glue between consoles. Identity, capability, data labels, and decision evidence need portable interfaces, while each enforcement point reports the rule it actually applied. Interoperability without policy continuity merely distributes the ambiguity.
There is one precise concession: in a bounded internal sandbox with synthetic data, no persistent credentials, and disposable effects, catalog breadth can be the right evaluation criterion because discovery speed is the purpose of the environment. That criterion ends when a capability touches customer data, durable state, money, communications, or shared infrastructure. Production value then shifts from finding an action to governing it.
This also changes procurement. I would ask a platform vendor to demonstrate a denied high-risk call, a policy version rollback, cross-tenant data isolation, immediate revocation, and an audit query that traces one action to its human sponsor. I would not start with how many prebuilt agents exist. Catalogs help teams begin. Explicit policy surfaces determine whether they can continue without accumulating invisible authority. The durable platform moat is not more things an agent can do — it is clearer control over when those things are allowed.