AI UX PlaygroundNewsletterJoin 2K+ AI designers and PMs on Substack. New teardowns, patterns, and prompts as they drop.

Agents

Sandbox Preview

Dry-run a plan with intended side effects, diffs, or receipts before real execution. Inspect in a safe context, then approve, edit, or cancel the live run.

Interactive demo

Agents

Acme renewal

Ready to run. Review what will happen first.

3 steps planned

Calendar, CRM update, and optional email draft.

Overview

How might we design sandbox preview so people can trust and act on AI output?

When to use

  • Essential for agentic automation, bulk admin tools, and cross-app workflows where users need to trust a manifest of effects before granting execution rights.

When to skip

  • Read-only assistants with no side effects to preview.
  • Trivial single-line edits where a full sandbox is slower than an inline diff.
  • When the sandbox cannot faithfully mirror production permissions, false previews are worse than none.

Rules

  • Previews that omit irreversible effects present in the real plan.

  • Execute buttons that skip sandbox after the first approval forever.

  • Sandboxes that mutate production data “just a little.”

  • Walls of logs with no human-readable summary of blast radius.

Evidence

ProductImplementation
Terraform planInfrastructure dry-run listing creates/updates/destroys before apply.
CI dry runsPipeline simulation or plan jobs before merging agent changes.
Email merge previewsSample personalized messages before bulk send.
Agent plan-then-act modesVisible plan and file diffs prior to apply in coding agents.

FAQ

What is sandbox preview for AI agents?

Sandbox preview is a dry-run of the agent’s plan that shows intended side effects (diffs, recipients, resources) before anything irreversible runs in production.

How does it differ from suggest / confirm / execute?

Suggest/confirm/execute is the autonomy policy. Sandbox preview is the artifact users review during confirm, what would happen if they approve.

What should a preview always include?

Targets, actions, reversible vs irreversible flags, and estimated cost or blast radius. Link to the exact payload when possible.

Is a diff enough?

For code, often yes. For email, purchases, or infra, show the semantic receipt (who/what/how much) not only raw blobs.