Run Lifecycle and Statuses
Every execution of an automation is a run — a durable record you can inspect in the run history of the automation. This page explains what each status means, how a run moves through its lifecycle,…
Written By Christopher Scaminaci
Last updated 3 days ago
Every execution of an automation is a run — a durable record you can inspect in the run history of the automation. This page explains what each status means, how a run moves through its lifecycle, what happens when things go wrong, and how the platform recovers runs stranded by a crash.
Automations are enabled by default for current tenants. See Automation Availability if the area is missing.
What a run record contains
For every run you can see:
Anthropic preserves the full turn-by-turn session transcript, and StackJack fetches the complete conversation when you open the transcript viewer. The database in your region does not contain the full transcript or full tool-result bodies, but it does contain the run's filtered trigger payload, Anthropic session id and key provenance, summary/error/usage, and pending approval-tool input. The normal run-monitor terminal path also attempts an automatic transcript walk before archive and persists a limited ordered tool-call projection, including names, bounded input arguments, classification/status, and timing. Direct interrupt or monitor-startup failure paths can bypass that capture, and a terminal run with no tool rows may be lazily backfilled when its tool calls are requested. Tool-result bodies and conversation/thinking text are not stored in this projection.
Run statuses
How a run starts
When a run is triggered (manually, on schedule, via webhook, or from a connected AI assistant):
- Eligibility checks. The automation must be active, not archived, have accepted consent, and be fully provisioned with an Anthropic agent id. Your tenant must have automations enabled.
SyncNeededis warning-only: the run can still use the last successfully synchronized remote configuration while newer local edits wait for re-sync. - Lease and execution slot. By default, only one run of an automation executes at a time. If its lease or the shared execution pool is busy, an eligible asynchronous trigger is saved as Queued. Synchronous callers receive an immediate capacity result instead. See Triggers and Execution Guarantees for the exact queue eligibility and defaults.
- Queue claim. When both the lease and a slot are available, the durable row moves Queued → Pending. If it expires or you cancel first, it moves to
SkippedorInterruptedwithout executing. - Credit reservation. For metered (non-BYOK) tenants, credits are reserved only after a slot is held — the lesser of a runtime-based estimate and the automation's per-run credit cap. If your balance can't cover the reservation, the run is recorded as CreditExhausted with the required and available amounts, and nothing is charged.
- Session start. An agent session is created and the run transitions Pending → Running. Live output streams to the portal.
The version id does not freeze every dependency used after launch. A queued trigger records its id at enqueue but later loads the current automation definition when the queue drains; crash recovery keeps the original id while doing the same for a replacement attempt. In deny-list mode, the callable set is materialized from the live catalog rather than stored as an exact historical allow-list.
There is also an important live configuration boundary: a normal save or version restore rewrites the shared per-automation MCP client's AllowedTools and IsDryRun values immediately, and neither is ever refused because a run is in flight. There is no mutation fence — not for an ordinary run, and not for a dry-run, fixture-recording, supervised-approval, deny-list, or strict-chain run. A run that is already Running or paused at AwaitingInput reads the shared client's current values on its next tool call, so a restore that narrows the callable surface, or that flips dry-run off, takes effect underneath it. Let the run finish or interrupt it before restoring a materially different configuration.
When the AI service is overloaded
Sometimes the AI service (Anthropic) is overloaded and refuses to start new work for a while. When that happens as a run starts:
- StackJack tries again a few times within the same few seconds.
- If the service is still overloaded, a webhook or scheduled run can be held instead of failed. It shows as Waiting for the AI service, and the run page shows when the next try is and when StackJack stops trying.
- StackJack tries again automatically, a little less often each time. When the service recovers, the run starts normally.
- If StackJack cannot start the run before the waiting time runs out, the run fails with a message that says so and how many times StackJack tried. The message says the AI service stayed overloaded only when the service refused the last try.
A run you start with Run now is never held: it fails at once, so you can decide what to do. A held run keeps the credits set aside for it. If it never starts, those credits are returned, except for any AI thinking steps that already ran before it tried to start. If the automation runs on your own Anthropic key, no StackJack credits are charged. You can cancel a held run from its run page.
How a run ends
The run monitor watches the agent session and applies the first terminal outcome that occurs:
- Agent finishes →
Completed. - Agent hits its own turn/budget limit →
Completed, with a warning in the error message. - Runtime cap reached → the run is stopped at once and marked
TimedOut, and its agent session is shut down. The agent is not asked to wrap up or summarize first, so its last message may stop mid-task. - Per-run credit cap reached → the run is stopped the same way and marked
CreditExhausted. The credits the run used are charged, including the model turn that crossed the cap. If the platform restarts while a run is in progress, the run can go further past the cap before the check resumes, and it is charged for the credits it used. - Model-side failure (retries exhausted, session terminated) →
Failed. - You stop the run →
Interrupted. - A Required/Ordered chain missed a callable tool after otherwise successful work → the completed run is relabeled
CompletedWithViolationswithout reopening or re-executing it.
After an ordinary monitored completion, StackJack re-fetches final token usage and active seconds from Anthropic before settlement. If that final provider read is unavailable on an interruption or recovery path, settlement uses the best persisted counters it can safely reconcile; the run page labels the StackJack current-rate calculation accordingly. For BYOK, Anthropic's invoice remains authoritative.
Status transitions are strictly ordered. A terminal run is not restarted or rewritten by ordinary lifecycle processing. The one deliberate refinement is Completed → CompletedWithViolations: it changes only the completion label and violations metadata after settlement; it is not another execution.
Stopping a run
You can cancel a run that is still Queued or stop one that is Running from the run detail page.
- For a run shown as Queued, click Cancel. The row moves directly from
QueuedtoInterrupted; there is no session, reservation, or remote work to unwind. - For a running execution, click Interrupt. There is no confirmation dialog or reason prompt — the stop takes effect immediately, and the reason is recorded as user-requested.
- The running row is durably marked Interrupted before StackJack cancels the live monitor and stops and archives the remote session. If that remote cleanup cannot be confirmed immediately, a background sweep keeps retrying it.
Billing for an interrupted live run reflects actual usage up to the stop point; the unused portion of the credit reservation is refunded automatically. A queued cancellation costs nothing. Stopping a run that has already finished is rejected — the run's recorded outcome never changes after it reaches a terminal status.
Human-in-the-loop approvals
A run can pause and wait for a person to approve a tool call. This happens only in a supervised test run, where both of these are true of the automation:
- its Live approval in test runs setting is on, and
- it is in dry-run (test) mode.
Approving a paused call runs it for real, in a test run as much as a production one. Dry-run mode blocks writes that the agent decides to make on its own; it does not block one you approve. This is the only way a connector write escapes dry-run mode, and it is worth knowing before you click.
Production runs do not pause for approval. The Ask for approval again on production runs setting is saved with the automation, but production pausing is not available yet.
In such a run, only destructive tools pause; read and ordinary write tools proceed without interrupting you, so a test run doesn't stop on every step. When one does pause:
- The run moves to AwaitingInput and its detail page shows a "Paused — approve a pending tool call" card naming the tool and the input the agent proposed. The tenant and global execution slots and the same-automation lease are all released, so a paused run is not counted against your organization's cap while it waits.
- Choose Approve & run or Deny. A deny can carry an optional reason, which is passed back to the agent so it can adapt. Approving executes the tool for real against your live systems — a supervised test run is still a test run, but the approved call is not simulated. A destructive tool additionally requires this automation's destructive-action acknowledgment; without it, StackJack still blocks the call after you approve it. Denying it, or letting the deadline pass, leaves it unexecuted and recorded as “would have called.”
- Resuming needs capacity again, and can be refused. Resume reacquires the global execution slot, and — for an organization on an Agent Runner plan — a slot from that organization's own purchased ceiling as well. If either is full, the approval is refused, the run stays AwaitingInput, and nothing is sent to the model. Because a pause releases the slots, an organization can accumulate several paused runs and then be unable to resume them all at once. The 60-minute approval deadline keeps running while you retry, so free a slot rather than retrying blindly. Two exceptions keep their older behavior: an organization with no Agent Runner plan is refused only when the shared platform pool is full, and a run a StackJack staff member started — or one a staff member resumes — is not measured against the purchased ceiling. See Concurrency, Slots, and the Run Queue. Resume deliberately does not retake the per-automation lease, because a resuming run is already that automation's in-flight run and would contend with itself.
- The run returns to Running and continues. The approval is recorded with its approver, and it authorizes exactly one real execution: StackJack writes a short-lived, single-use grant that is consumed and deleted the first time a matching call arrives. That grant is scoped to the automation's MCP client and the tool's name — the run id, the tool-use id, and the approver are stored on it as an audit record but are not compared when it is consumed, and the tool's input is not part of the key. So the grant authorizes the next call to that tool name from that automation, not specifically the call you looked at. It expires on its own within minutes if no such call arrives.
Two things to know about the edges:
- A run that hits a tool approval without being eligible under the supervised-test policy or the enabled production policy ends as
Failed, with a message pointing you at the approval setting or at adjusting the automation's tool policy to auto-allow. With the shipping production flag off, promoted live automations therefore do not sit paused waiting for a human. - The approval window is one hour. If nobody approves or denies within 60 minutes of the pause, cleanup tries to stop and confirm the provider session. Only a confirmed stop lets the platform end the run as
Failed, bill the usage actually incurred, and refund the remainder of the reservation. If stop confirmation is uncertain, the run staysAwaitingInputwith its session reference and cleanup retries on a later sweep.
Occasionally StackJack cannot capture which tool is pending; the page says so plainly and the run can't be approved from there. Adjust that tool to auto-allow before trying again, or contact support. If no decision arrives, the one-hour cleanup rule still applies.
Crash recovery: what you may observe
The durable queue survives a restart and resumes draining when StackJack's automation service returns. A run that had already reached Pending or Running uses a requeue-once recovery policy:
- On its first detected stranding, StackJack confirms the old remote session is stopped, then reuses the same run ID and credit reservation for one fresh attempt. You may see the run return to Pending and then Running.
- If the old session cannot be confirmed stopped, the run stays parked in its current state and cleanup retries later. StackJack does not start a second session while the first might still be live.
- If the accepted recovery attempt strands again, or the automation is no longer runnable, the run becomes Failed. Actual usage is settled and the unused reservation is refunded.
- A StackJack staff diagnostic run is failed rather than replayed.
The fresh attempt keeps the original AgentVersionId but reloads the current live automation definition and current remote ids. It can also repeat connector-side effects that the old session completed before StackJack lost contact. Use vendor-supported idempotency keys or another deduplication boundary for operations that must happen only once.
Queued rows follow their own deadline: they become Skipped, not Failed, if they wait too long. Approval pauses follow the one-hour rule above. A separate settlement sweep reconciles terminal runs whose ledger finalization was interrupted, so recovery does not rely on the UI staying open.