Automation Failure Notifications
StackJack can email you when one of your automations ends badly, so you don't have to watch Run History continuously. You turn notifications on per automation and choose the address that receives…
Written By Christopher Scaminaci
Last updated 3 days ago
StackJack can email you when one of your automations ends badly, so you don't have to watch Run History continuously. You turn notifications on per automation and choose the address that receives them—the email names the automation, says how the run ended, includes the error detail, and links straight to the run's detail page. Digest and daily limits described below mean not every failed run produces its own message.
Turning it on
The setting lives with the automation's other guardrails, in the Advanced Builder's Guardrails & Deploy section (see Editing an Automation):
- Turn on Email me when a run fails.
- Enter the Notification email address — the field appears once the toggle is on. It can be any address — a shared inbox, an on-call alias, a ticketing address — it does not have to be a StackJack team member.
- Save (or deploy) the automation.
The address is validated for basic plausibility when the toggle is on. Switching failure alerts off does not clear the stored address. That address is also the route for production approval requests when StackJack's default-off production-approval gate is enabled and the automation arms that policy, so review it whenever you change approval settings.
The toggle and address sit in the same card as the runtime and credit caps and the autonomous-operation consent switch, so you can confirm the setting there before deploying.
When an email is sent
A failure notification can be staged when a tenant run reaches one of the failure-like terminal states—Failed, Timed out, Ran out of credits, or Interrupted after execution began. A normal Completed run, a run that never started (Skipped), and a run you cancel while it is still Queued do not send a failure email:
This is deliberately a wide definition of "something went wrong": stopping a Running or Awaiting input run is eligible for an alert because the work did not finish. Cancelling a queued run is different—it had no execution slot, AI session, credit reservation, or monitor, and StackJack intentionally records Interrupted without staging a failure email. Eligible alerts are still subject to the opt-in, digest, and daily limits below.
Crash-recovered runs are covered too: if a run was stranded by a service interruption and later recovered into a Failed state, it still triggers the same notification.
A lapsed Agent Runner plan can generate one alert per trigger. If your organization's own Anthropic key was admitted by an Agent Runner plan that has since ended, every run is refused before it starts — but it is recorded Failed, not Skipped, and it stages a failure notification exactly as any other Failed run does. So every scheduled occurrence, webhook delivery and Run now produces both a Failed run and an eligible alert until the situation is resolved; the per-automation 60-minute digest window and the daily cap below are what bound the volume. That distinction is worth naming, because a lapse-refused run also never started — no AI session, no reservation — yet it emails, because the refusal is a fault rather than a capacity outcome, which is what puts it on the Failed row above rather than the Skipped one. Muting the alert does not restore the runs; fix the cause instead — see When an Agent Runner plan ends.
Operator diagnostic runs are excluded. When StackJack staff run your automation for support or diagnostics, that run is tagged as operator-initiated and never sends a tenant failure email — it isn't your automation firing on its own trigger.
Required-tool warnings
There is one warning outcome that uses the same opt-in address without being a failed run: if a Completed or Completed with violations run did not use a tool marked Require, StackJack can send a required tool not used email. These warnings have their own 60-minute per-automation digest window and do not count against the daily failure-alert cap.
A chain-contract violation or post-step-check warning does not send an email by itself. Open Run History to see those outcomes.
Approval-request emails
Production runs do not pause for approval, so StackJack sends no approval-request emails for them. Supervised dry-run tests can pause and receive a live human decision in the Portal — and approving one executes the connector call for real against your live systems.
Those test pauses deliberately do not send approval-request emails, and the reason is the supervision itself: a supervised test run is by definition being watched by the person who started it, and the Portal shows them the approve/deny card live. An email there would be noise.
Because an approved test call is real, an unwatched test pause is not harmless. It sits in Awaiting input with nothing emailed to anyone. Once the pause passes 60 minutes, StackJack attempts to stop it: a cleanup sweep confirms with Anthropic that the session has stopped, and only then does the run end as Failed with the action not taken and the unused reservation refunded. If the sweep cannot confirm the stop, the run deliberately stays Awaiting input and the next sweep tries again, rather than reporting a stop that may not have happened. So elapsed time alone is not proof a run stopped — read its status. Stay on the run detail page while a supervised test is running; if you start one and walk away, expect to come back to a stopped or still-paused run rather than a completed one.
What the email contains
- A headline and subject line that name the actual outcome (failed / timed out / ran out of credits / interrupted) rather than a generic "failed".
- The automation's name.
- The error or stop detail captured for the run (or a neutral "No detail was captured for this run." when a run ended before any detail was recorded).
- A View run details link that opens the run detail page. If the run has an Anthropic session, click View transcript there to fetch the full conversation.
Grouping and rate limits
So a flapping automation — or a bad afternoon across your whole tenant — can't bury your inbox, StackJack groups and caps alert emails:
- Per-automation hourly window. The first failure of an automation is queued for sending immediately and opens a 60-minute window. The email normally goes out within about half a minute of the run failing, so expect a short delay rather than an instant message. Further failures of the same automation inside that window are recorded as suppressed and do not send or update an email. After the window closes, the next failure starts a fresh alert.
- Per-tenant daily cap. Across all your automations, StackJack sends up to about 10 failure alerts per UTC day. When that cap is reached, you get a single "failure alerts muted until tomorrow" notice, and further failures for the rest of the day are recorded silently. Alerts resume automatically at the start of the next UTC day.
These limits only affect failure emails. Required-tool warnings use their separate hourly digest; live approval requests use neither digest nor daily cap. Every run is still recorded in Run History and on the run detail page whether or not an email was sent.
Delivery behavior
- The alert is recorded durably the moment the run ends, in the same write that records the failure or the approval pause. A StackJack restart between the run failing and the email going out does not lose the alert.
- Sending is asynchronous and can lag. Expect roughly half a minute in normal conditions. A restart, a busy queue, or the email provider itself can make it longer. The run detail page is available before the email arrives, so use it as the faster signal.
- A failed send is retried up to five times, with a growing gap between attempts. An alert that still cannot be delivered is retained with its attempt count and the last error, so support can tell you why it never arrived.
- A duplicate email is rare but possible. If StackJack fails after the email provider accepted a message but before the send was recorded, the retry can send it again. Do not treat the number of emails as a count of failures — Run History is the count.
- Notifications are best-effort operational alerts, not a billing or audit record. The authoritative history is always Run History, Run Detail, and the credit ledger.