Start from a neutral baseline and add what matters to you. Criteria are labeled by source — the platform baseline is architecture-neutral; buyer-contributed criteria are shown separately.
This evaluation is stored in your browser only. We cannot see it, and it is not tied to any account. Save it to a link or create an account to keep it across devices — you can export it at any time either way.
Signing up adds sharing with your team, sending this as an RFP to vendors, and private document sharing. Nothing above is taken away, and nothing here is sent anywhere until you choose to.
Strong answers give a concrete, customer-referenced full-autonomy percentage rather than a blanket 'AI-powered SOC' claim; be skeptical of vendors who can't separate 'AI recommends' from 'AI executes.'
A credible vendor can describe a real case where the guardrail worked, not just that guardrails exist in theory; ask specifically what happens when the AI is uncertain — does it default to human escalation or to action?
Graceful degradation to human escalation on genuinely novel patterns is a sign of a mature, safety-conscious design; overconfidence on out-of-distribution cases is a real operational risk worth probing directly.
False-negative rate is inherently hard to measure (you only know what you eventually catch some other way); a credible vendor explains their measurement methodology rather than citing an unverifiable low number.
No buyer-contributed criteria yet
Verified buyers can suggest criteria (anonymized before pooling).
Full decision-reasoning audit trail and easy reversibility of autonomous actions are essential safety properties; a black-box action with no explainable reasoning trail is a real operational and compliance risk.
Incremental autonomy rollout (not day-one full automation) is a materially safer and more realistic adoption path; ask for a concrete example of how a real customer graduated their autonomy level over time.
Cross-customer model training on shared telemetry raises real confidentiality questions security teams should ask about explicitly; a vendor should be precise about tenant isolation, not vague.
Insist on speaking directly to a reference customer rather than accepting only a vendor-produced case study; this is an early, fast-moving category where marketing claims often run ahead of production reality.
A documented, tested failover to human-only operation during an outage is essential; a platform with no plan for its own downtime creates a real coverage gap exactly when defenses might be needed most.
Ask for a real, customer-validated cost comparison against equivalent human-analyst capacity, not just a platform price in isolation — the whole value proposition rests on this comparison being real.
An autonomous action that can't be explained afterward is a real liability if it's later challenged — ask for a specific answer on decision-explainability, not just action logging.
Trend-over-time autonomy-progression reporting is a distinct, meaningful metric for this category specifically — confirm this exists as a maintained, exportable report showing the trust-earning trajectory, not just a static current-state autonomy level.
Expanding AI autonomy is a consequential decision that should have real governance — ask for a specific, documented approval process, not an informal or unilateral configuration change.
If the AI platform is effectively authoring detection logic, that logic should be reviewable by the customer's own security engineers — a fully opaque detection layer creates real governance risk.
An AI system with autonomous action authority is itself a potential attack surface — ask for evidence of genuine adversarial testing against the AI's decision-making, not just traditional security testing of the platform's infrastructure.
Ask for an honest per-environment reliability breakdown; uneven autonomous-action reliability across environments is a real gap a vendor should disclose rather than obscure, especially given the higher stakes of autonomous (versus advisory) action.
A silent model update that changes autonomous decision-making behavior is a real operational risk — ask for a specific change-management process around model updates, not just confirmation the model improves over time.