I am currently architecting a system that handles long-running distributed workflows involving agentic planning and external execution. I’m focusing on the “temporal validity gap”—the TOCTOU (Time-of-Check to Time-of-Use) problem where state (IAM permissions, resource policies, or external environment) changes in the window between the initial authorization/planning phase and the actual execution (e.g., due to human-in-the-loop delays or retry cycles).
While Temporal’s durable execution is a perfect fit for maintaining the workflow state, I’m looking for the best architectural patterns to handle execution context validity at the moment of resume:
In scenarios where a workflow waits for human approval or external events (potentially for minutes or hours), how do you recommend ensuring that the execution step re-validates the initial authorization context?
Is there a preferred pattern for “Side Effect” re-validation to prevent executing tasks with stale credentials or outdated policy assumptions?
How do you typically handle compensating transactions or safe aborts if an activity discovers that the environment state no longer matches the original intent of the workflow?
I’m interested in hearing how you balance strict consistency versus performance in these long-running flows. Are there specific patterns you use to bind the “decision” to a snapshot of the system state?
Thanks for any insights or pointers to existing best practices!
I wouldn’t use re-validation as a separate step here, since that just moves the TOCTOU window. Put the check in the same Activity that performs the action, mint a short-TTL, action-scoped credential there, and pass the original intent snapshot (resource ID, action, plan hash) as a precondition that the target system enforces.
workflow.sideEffect is a footgun for this because its result is recorded in history and replayed verbatim, so an old “authorized” result can remain frozen into the workflow. If the Activity detects drift, throw an ApplicationError with nonRetryable set, and let the Workflow catch it to take the abort or compensation branch.
Thanks for the detailed response — this is exactly the kind of pattern I was hoping to hear about.
Completely agree on point 1: revalidation as a separate step just relocates the TOCTOU window rather than closing it. Minting the credential inside the same Activity that performs the action, scoped to the intent snapshot, keeps the check-and-use atomic in the way that actually matters.
On workflow.sideEffect — good catch, and it highlights something more general: in most agent/workflow frameworks, “authorization” gets treated as a point-in-time event rather than a state binding that has to hold at commit time. Temporal’s replay semantics make that concrete (a stale “authorized” result literally gets frozen into history), but the same failure shows up anywhere a plan and its execution are separated by latency — retries without revalidation, stale cached tool outputs, absent inter-step revalidation on longer plans, etc.
One addition worth considering: the precondition check probably shouldn’t be bound to a global state snapshot. A single hash-of-everything invalidates on changes that are completely irrelevant to the operation and produces excessive false positives. Better to have each Activity declare which state dimensions are actually material to its correctness (resource ID, relevant policy version, specific IAM scope) and bind the credential/precondition only to that declared scope — otherwise you either miss real drift outside a too-narrow check, or get spurious reauth storms from a too-broad one.
On compensation — agreed on the ApplicationError(nonRetryable) + Workflow-level catch pattern, though worth flagging explicitly: this guarantees a safe stop, not a rollback of already-executed side effects. Operations with irreversible external effects (sent notifications, already-applied cloud changes) need their own compensation semantics defined per-operation — there’s no generic precondition-check mechanism that solves that part for you.
Curious whether you’ve run into the credential-minting Activity itself becoming a bottleneck under high-frequency, short-lived Activities — did you land on any pattern for amortizing that cost?