Agent Shield and the Guardian
Two layers keep agent mail safe. Agent Shield works at the inbox: it scores what comes in and polices what goes out. The Guardian works at the platform: it watches sending behaviour across all tenants and benches a misbehaving one before mailbox providers notice.
Agent Shield: inbound
Every inbound message is authenticated (SPF, DKIM, DMARC) and scored for spam, phishing and prompt-injection patterns before it is stored. The result is the message’s verdict:
| Verdict | Meaning | Where it lands |
|---|---|---|
clean |
Authenticated, no risk signals | inbox |
suspicious |
Risk heuristics fired (injection phrasing, lookalike sender, odd links) | inbox, flagged |
spam |
Classified as spam | spam |
blocked |
Matches a block rule or known-bad pattern | spam |
unauthenticated |
SPF/DKIM/DMARC failed or absent | inbox, flagged |
The auth object carries the individual SPF/DKIM/DMARC results and shield carries the signals. Threads containing a suspicious or blocked message are tainted, which matters for the outbound rules below.
Treat the verdict as a safety signal, not permission. Content from anything other than clean should be shown to the model as data, never as instructions.
Agent Shield: outbound policy
Set per inbox at creation or with PATCH /v1/inboxes/{id}:
| Field | Values / effect |
|---|---|
dlp |
off, observe, enforce. Enforce holds messages whose body contains secrets, API keys or PII patterns; observe only records them. |
max_send_per_hour |
Sends beyond the cap are held, not dropped |
allow_recipients |
Only listed addresses, domains or their subdomains may be emailed; anything else is refused outright with 403 recipient_not_allowed |
require_approval_on_taint |
Hold any reply on a tainted thread |
The policy applies to every door: REST compose, reply, forward, drafts and scheduled send, SMTP submission and MCP tools. Outbound policy is available from the Developer plan; the human-approval workflow below is a Startup and Enterprise feature (on Developer, require_approval_on_taint is unavailable and dlp: enforce and the hourly cap reject rather than hold).
Approvals
A held send returns 202 with { "held": true, "approval_id": "…", "reason": "dlp" | "tainted_thread" | "rate_cap" } and emits approval.required. A human (or a trusted supervisor process) decides:
| Method | Path | Purpose |
|---|---|---|
GET |
/v1/inboxes/{id}/approvals?status=pending |
Queue of held messages |
GET |
/v1/approvals/{id} |
Inspect one, including the full message |
POST |
/v1/approvals/{id}/approve |
Send it |
POST |
/v1/approvals/{id}/reject |
Reject permanently |
Statuses: pending → approved → sent, or rejected. Deciding twice returns 409. Approvals are also available in the console and are audited.
The Guardian: cross-tenant detection
The Guardian evaluates rolling signals per tenant, domain, sender, credential and tag: submission spikes, bounce rate, complaint rate, authentication failures, recipient-domain concentration and unusual send times. When a rule trips it emits anomaly_detected, raises an alert, notifies your alert contacts and, depending on the rule, automatically pauses the offending scope. Detection-to-pause is typically under a minute.
Only the matching scope is paused. Other tenants on the same account, and even other senders within the same tenant, keep sending.
Pauses
A pause is created automatically by the Guardian or manually with POST /v1/suspensions (scope suspend). Scopes: account, tenant, domain, sender, SMTP credential, API key, tag or recipient domain.
{ "scope": "tenant", "tenant_id": "TENANT_ID", "queue_action": "hold", "reason": "complaint spike" }
queue_action |
New submissions | Already-queued mail |
|---|---|---|
hold (default) |
Accepted and parked (held) |
Parked |
reject |
Refused with 403 (SMTP 4.7.1) |
Parked |
Nothing is deleted. When the investigation is done, resume:
curl -X POST "$API/v1/suspensions/$ID/resume" -H "Authorization: Bearer $ADMIN_KEY" \
-H "Content-Type: application/json" -d '{"action": "release", "note": "false positive, customer confirmed campaign"}'
release re-queues everything that was held; purge cancels it (cancelled events). The action and note are audited.
| Method | Path | Purpose |
|---|---|---|
POST |
/v1/suspensions |
Create a pause |
GET |
/v1/suspensions, /v1/suspensions/{id} |
List / inspect with evidence |
POST |
/v1/suspensions/{id}/resume |
Release or purge |
POST |
/v1/tenants/{id}/pause, /v1/tenants/{id}/resume |
Tenant shortcuts |
Rules, alerts, contacts
| Method | Path | Purpose |
|---|---|---|
GET / PUT / DELETE |
/v1/rules, /v1/rules/{id} |
Read effective rules; override thresholds or actions per account or tenant |
GET |
/v1/alerts |
Open and historical alerts |
POST |
/v1/alerts/{id}/ack |
Acknowledge |
GET / POST / DELETE |
/v1/alert-contacts |
Who gets notified |
Observability and audit
| Method | Path | Purpose |
|---|---|---|
GET |
/v1/metrics |
Time-series counters by tenant, domain, status |
GET |
/v1/reputation |
Tenant and domain reputation snapshot |
GET |
/v1/queue |
Queue depth by status |
GET |
/v1/overview |
Dashboard totals and health |
GET |
/v1/audit |
The append-only, hash-chained audit log (admin) |
Every pause, resume, approval decision, key change, rule change and policy change is written to the audit log, each entry chained to the previous by hash so tampering is detectable. Enterprise plans can export it.