Triggers and execution guarantees (reference)
Automations start in one of three ways — on demand, on a schedule, or via webhook — plus two special paths (runs started by StackJack staff, and runs invoked by a connected AI assistant). This page…
Written By Christopher Scaminaci
Last updated 3 days ago
Automations start in one of three ways — on demand, on a schedule, or via webhook — plus two special paths (runs started by StackJack staff, and runs invoked by a connected AI assistant). This page explains exactly what execution guarantees each trigger gives you: what happens under load, whether triggers are queued or deduplicated, and how retries behave.
The durable execution contract
StackJack has a fixed pool of execution slots, but an eligible asynchronous trigger is not discarded when every slot is busy. It becomes a durably stored run shown as Queued in the tenant portal. The run ID, automation version, trigger type, and trigger payload are all saved before StackJack accepts the trigger.
While a run is queued:
- It has no Anthropic session, holds no execution slot, and reserves or consumes no credits.
- It survives a restart of StackJack's automation service, because the queue is stored durably rather than held in memory.
- It can be cancelled. Cancellation and the queue drainer compete on the same status check, so either cancellation wins (
Queued → Interrupted) or the drainer wins (Queued → Pending); the work is never launched twice by that race. - StackJack re-checks that the tenant and automation are still runnable before launch. A disabled, archived, or unprovisioned automation is moved to Skipped instead of being resurrected from the queue.
SyncNeededalone is not a blocker: a provisioned automation can run its last successfully synchronized remote configuration.
The defaults are:
The queue drains on-demand work first, then webhooks, then schedules; within each trigger class, older work goes first. This is a priority policy, not a promised start time.
Queue expiry
If a queued run reaches its deadline first, it becomes Skipped. It never starts and costs no credits. Scheduled expiry copy tells you the next scheduled occurrence will try again; expiry does not manufacture an extra retry.
An independent cleanup sweep catches queue rows that are still waiting more than one whole queue lifetime past their own deadline — so two hours for a row on the 60-minute default. Those rows also become Skipped. A run set to wait indefinitely has no deadline and is deliberately exempt: it waits until a slot frees or somebody cancels it.
Which callers can queue
Portal Run now, scheduled triggers, webhook triggers, connected-AI stackjack_run_agent calls, and builder test runs are all queue-eligible — every one of them returns before the run completes. Four lanes are not queued: nested automation-as-tool child runs (always synchronous — they ride the parent's slot), StackJack staff diagnostic runs, an AwaitingInput resume, and fixture-replay evaluation runs. Those paths can still receive a capacity response and should retry only when their response says it is safe.
A queued Run now is now the ordinary outcome for an organization sitting at its own concurrency ceiling, not just for a platform-wide saturation — see Concurrency, Slots, and the Run Queue.
Same-automation overlap
By default, a per-automation lease prevents two live runs of the same automation from starting together. A contending queue-eligible run waits in the same durable queue; it does not use a separate retry mechanism. Turn on Allow this automation to run more than once at a time only when runs are genuinely independent and safe to overlap.
One run at a time is the normal behavior, not a guarantee. Writes that must happen once need their own idempotency or deduplication.
A fixture replay (evaluation) run is the exception. If it cannot hold its automation's one-at-a-time slot, it is refused before any session opens, is not billed, and reports that it was not run. In every other case, when one run at a time cannot be held, fixture recording captures nothing, and ordered-chain and duplicate-call enforcement decline rather than enforcing against unreliable state. The durable queue itself is unaffected — the same queued run still cannot be started twice.
Restart and recovery semantics
Queued work is stored durably and is reconsidered when the queue starts draining again. A run that had already reached Pending or Running when the owning process died follows a different recovery path:
- The cleanup sweep first confirms that any old Anthropic session is stopped and archived. If it cannot confirm that, it leaves the run parked with the old session reference and retries on a later sweep; it does not start a possible duplicate session.
- The first time a tenant run is found stranded, StackJack reuses the same run record, resets it to
Pending, and attempts one fresh session. For StackJack-managed credits, that fresh attempt reuses the original reservation and the ledger settles once per run id. A BYOK run has no StackJack reservation; both the stranded and replacement Anthropic attempts can incur charges on your Anthropic account. If a lease or slot is still unavailable, the run remainsPendingfor a later cleanup attempt; the one-replay budget is consumed only when fresh execution is accepted. - If that accepted recovery attempt is stranded again, or the automation is no longer runnable, the run becomes
Failedand any StackJack-managed reservation is reconciled. StackJack staff diagnostic runs are failed rather than replayed.
This recovery can make one run move from Running back to Pending before it runs again; it is a crash-recovery reset, not an ordinary lifecycle transition. The replacement is a fresh execution and can repeat connector-side effects completed by the stranded session. Use vendor-side idempotency or deduplication for operations that must happen only once. See Run Lifecycle and Statuses for the user-visible statuses.
What StackJack guarantees during an outage
If StackJack's automation service is unavailable, no automation runs until it returns. Queued and in-flight work is kept and picked up afterwards.
Plan around that. An automation whose work is time-critical needs a way for you to notice it did not run — a failure notification does not help here, because a run that never started raises nothing. Check run history after any StackJack service incident.
On-demand (manual) runs
Trigger a run from the automation's page ("Run now"), optionally with free-text context for the agent.
- If you provide context, it is used as the kickoff message.
- If you don't, the automation's configured default context (set in its trigger configuration) is used. Explicit context always wins. Default context applies only to on-demand automations — scheduled and webhook runs never use it.
- A portal launch returns a run ID immediately. It either starts or enters the durable queue described above; the run detail page shows which.
Builder test chats are runs. Every message you send in the builder's test panel starts its own ordinary asynchronous run — it is billed and counted exactly like any other run, and it appears in run history. Because it is an ordinary launch it is queue-eligible: if every slot is busy the test run is queued, and the live stream in the test panel does not open until the run actually starts.
Scheduled runs
Scheduled automations run on a cron expression (5-field, minute precision) with an optional timezone.
What's guaranteed — and what isn't:
- Minimum interval: the schedule must fire strictly less often than once per minute. Every-minute or sub-minute cadences are rejected at save time with guidance to use a webhook trigger instead.
- Timezones: IANA and Windows timezone IDs are accepted; unknown timezones are rejected when you save. (If a timezone somehow becomes unresolvable later, the schedule falls back to UTC rather than stopping.)
- Durable deferral under load: if all execution slots are busy, the occurrence is saved as a queued run and starts when capacity frees. Its deadline is clamped to less than the next scheduled occurrence; if the deadline wins, the run becomes Skipped and costs nothing.
- Silent skips (no run record): a tick skips without leaving a run if the automation has been deactivated, archived, switched to a different trigger type, or if automations have been disabled for your tenant.
- Retries: genuine startup failures (for example a transient platform error) are retried automatically by the scheduler. A queued occurrence is already accepted work, so the scheduler does not enqueue a duplicate retry. Credit refusal, cancellation during launch, and queue expiry are not startup retries; the next cron occurrence tries normally.
- Overlap protection: one-at-a-time execution is the default. A later occurrence waits rather than starting beside an in-flight run, unless you explicitly enable Allow this automation to run more than once at a time. The queue deadline still prevents it from waiting into the next occurrence.
- A staging copy never fires on a schedule. A staging copy inherits its production automation's trigger type and cron verbatim so the two can be compared like for like, but staging copies are excluded from the scheduler outright — the platform cannot register a schedule for one whatever its settings say, so a copy of a scheduled automation can never double up on production's cadence. Run a staging copy on demand to test it.
Schedule drift and repair
If schedule registration fails when you save the automation, the save still succeeds — the automation is flagged with a schedule drift warning in the portal. Drift is repaired automatically the next time the platform restarts, or immediately via the Re-register schedule action on the automation. You can check registration state at any time from the automation's schedule status.
Webhook runs
External systems (your PSA, RMM, Zapier, etc.) trigger the automation by POSTing to its unique webhook URL, optionally passing payload data to the agent.
Key guarantees, in brief (see the webhook ingestion contract for the full contract):
- A delivery accepted for immediate execution or durable queueing receives
202with the new run ID. Do not retry a queued202: StackJack already owns that work. - For new automations, a byte-identical redelivery that reaches deduplication and matches a successfully persisted receipt inside the deduplication window receives
200with the original run ID and creates no second run. Older automations created before deduplication was introduced retain their previous no-dedup behavior. - Per-automation rate limit: 10 invocations per minute. The rate limiter runs before signature verification and deduplication, so an otherwise-identical resend can receive
429with no run ID instead of the duplicate200. - A
500with no run ID is ambiguous: aPendingrow may already exist and later execute. There is no safe unconditional retry or idempotency guarantee for that outcome; retrying can create duplicate work. Use the canonical webhook response table for every retry decision. - StackJack attempts to persist application-level outcomes in the automation's webhook receipt log, but receipt writes are best-effort and do not change the webhook response. A missing receipt is not proof that no request or run exists. Unknown, inactive, and archived automation IDs are deliberately not submitted to the log (scanner-noise defense), and a request rejected by the web server before the controller runs may also have no receipt.
Runs started by StackJack staff
StackJack operators can run an automation on your behalf for support and diagnostics. These runs:
- appear in your run history like any other run; the run detail page displays an information notice that StackJack staff started the run on your behalf,
- never charge your credits (credits consumed shows 0), and
- still record token usage for transparency.
Runs invoked from a connected AI assistant
An automation marked "Expose as MCP tool" can be discovered and run by your own AI assistant (a harness — the app hosting the AI, such as Claude) through your StackJack MCP endpoint. Every trigger type can be exposed: a scheduled or webhook automation invoked this way runs off-cadence and its run records as a manual run, leaving the cron schedule and webhook endpoint untouched (a harness-supplied payload to a webhook automation is fenced exactly like a real inbound webhook body). The stackjack_run_agent tool is asynchronous — it returns a run id immediately rather than waiting for the result, so at saturation the launch is queued and comes back as status: "Queued" with the run id; poll stackjack_get_agent_run rather than retrying (a retried launch without the same idempotencyKey creates a second billed run). It otherwise uses the same billing and guardrail rules as a portal-triggered run. The tool can also re-run a previous run's stored trigger payload (rerunOfRunId) — useful for retrying a failed webhook run after fixing the automation's configuration; the portal's run detail page offers the same action as an asynchronous launch.
Related pages
More in How automations behave
Guardrails and Safety ControlsStill need help? Ask the team