Tool inventory
Do you have a complete inventory of every tool your agent can call (manifest, MCP export, OpenAPI schema), classified read versus write?
Interactive Resource
Most benchmarks ask whether an agent can finish the task. This test asks whether its write-actions are safe to enable.
Enterprise teams do not block agents on intelligence. They block them at the write-access door, with questions about evidence, policy, approval, and traces. Use this scorecard on the workflow where a wrong write would create customer, financial, or operational damage.
Answer for the agent whose write access your enterprise customer has to approve. The score updates immediately.
Do you have a complete inventory of every tool your agent can call (manifest, MCP export, OpenAPI schema), classified read versus write?
Would you notice a new write-capable tool appearing in the agent surface before your users or your customer does?
Are the agent write-actions ranked by business risk: financial, customer-facing, workflow-state-changing, destructive?
Can you name, today, the single riskiest write-action and its worst-case business outcome?
Is the evidence each risky write requires (linked records, thresholds, customer tier, environment) explicit and checked, rather than implied by the prompt?
When required evidence is missing, does the risky write reliably stop or escalate instead of proceeding anyway?
Are approval thresholds explicit (amount, priority, environment, customer tier) with a named approver, defined outside the prompt?
Do risky writes route to a human with the full context needed to decide, rather than a notification after the fact?
For any past write, can you show in one artifact which evidence was present and which rule produced the decision?
Could you hand an enterprise security team a review-ready dossier for this agent, covering its write-actions, today?
Below 80, the agent may be useful, but at least one control an enterprise security review will ask about is still ad hoc. Rippletide helps teams cross that gap with Safety Cases, unsafe scenario tests, and review-ready evidence.
Keep writes disabled for now.
Observable, not yet provable.
Close, with evidence gaps left.
Package the proof for review.