Choose a fixed workflow when the steps are known; use agents where decisions vary.
Most useful GenAI systems are not a single model call. They read a document, extract fields, check them against a database, ask a person when something looks wrong, and write a record at the end. Orchestration is the code and infrastructure that decides which step runs next, passes data between steps, retries what failed, and remembers where a run got to. It is ordinary distributed-systems engineering, but model calls make it harder: they are slow, they fail in new ways (rate limits, timeouts, malformed output), they cost money per attempt, and the same input can produce a different result on a second run.
The central mental model is a spectrum of control. At one end is a fixed workflow: the steps and their order are written in code, and the model fills in content inside each step. At the other end is an agent, where the model chooses the next step at run time. Most production systems sit nearer the workflow end and use agents only inside a bounded step. Around that sits the runtime: state that records progress, queues and events that move work between processes, and durable execution that lets a run survive a crash, a deploy or a three-day wait for a human approver.
The Basic questions cover the building blocks: what a workflow is, how it differs from an agent, DAGs, routing, fan-out and fan-in, asynchronous execution and state. The Advanced questions cover the infrastructure underneath: queues, event streams such as Kafka, making runs resumable, event-driven designs, choosing between a visual tool like n8n and code, and when to build or adopt an orchestration engine. Three added questions cover durable execution, human approval steps and changing a workflow while runs are in flight, which are the problems teams usually meet in their first month of production.
An AI workflow is a sequence of steps, defined in code or configuration ahead of time, where some steps call a model and others run ordinary logic such as validation, lookups and database writes.
A workflow fixes the control flow (which step runs, in what order, under what conditions) and lets the model contribute content inside individual steps. An invoice pipeline might call a model to extract the supplier, date and line items, then use plain code to check that the line items sum to the total, look up the supplier in the vendor table, and save the record. The model never decides whether the database write happens; the code does.
This split matters because models are good at reading messy input and bad at being consistent. Putting the deterministic parts in code gives you behaviour you can unit test, predictable cost (you know how many model calls a run makes), and clear failure points. When the extraction step returns a total that does not match the line items, the workflow knows exactly which step failed and can retry it, route the document to a person, or reject it.
Most workflows follow a few recurring shapes: a chain (each step feeds the next), a router (classify the input, then take one of several paths), fan-out and fan-in (split work into parallel branches and merge the results), and a check-and-retry step (validate model output and ask again with the error if it fails). Larger systems combine these shapes into a graph.
A workflow is the right default when you can write the steps down before seeing the input. That covers most document processing, data enrichment, report generation and triage. The cost is rigidity: an input the designers did not anticipate either fails a check or takes the wrong path, so you need a fallback route, usually a human queue, for cases that do not fit.
In a workflow, code decides the sequence of steps and the model fills in content. In an agent, the model decides the next step at run time by choosing tools in a loop until it judges the task done.
The difference is who owns the control flow. A workflow is a program with model calls inside it; you can read the code and list every path a run might take. An agent is a loop: the model sees the goal and the results so far, picks a tool or an answer, the harness executes it, and the result goes back to the model. The set of possible paths is open-ended, which is the point: the agent can handle tasks whose steps cannot be written down in advance, such as researching an unfamiliar question or debugging a failing test.
That flexibility has a price. Each decision is a model call that can be wrong, and errors compound: if each step is right 95% of the time, a ten-step trajectory is fully right only about 60% of the time. Agents use more tokens, take longer, are harder to test because the path varies, and need guards a workflow does not: step limits, budgets, loop detection and permission checks on every tool.
In practice the choice is not binary. Common hybrids are a workflow with one agentic step (a fixed pipeline where only the research stage runs an agent with a step cap), and an agent whose tools are themselves small workflows (the agent calls create_refund, which runs validation and approval in code). Many teams find that a task they first built as an agent becomes a workflow once they see which paths it actually takes in production.
A useful test: if a domain expert can write the procedure as a numbered list that covers most inputs, build a workflow. If the expert says it depends on what you find, and the branches are too many to enumerate, consider an agent, and constrain it as tightly as the task allows.
A DAG (directed acyclic graph) is a set of tasks connected by dependency arrows that never loop back. Any task whose inputs are ready can run, which makes the order, the parallelism and the restart point explicit.
Each node is a task and each arrow means "this task needs that task's output". Directed means the arrows have a direction; acyclic means you can never follow them back to where you started. Those two properties let an orchestrator compute a valid execution order (a topological sort), run independent nodes at the same time, and know exactly which downstream tasks to rerun when one node changes or fails.
For AI pipelines a DAG is a natural fit when the steps are known. A report generator might fetch sales data and fetch support tickets (independent, so in parallel), summarise each (each depends on its fetch), then write the combined report (depends on both summaries). If the support summary fails, only that branch and the final step need rerunning; the sales branch's cached output is reused.
The acyclic rule is also the main limitation. A DAG cannot express "retry the draft until the reviewer approves it" as a graph edge, because that is a cycle. Data orchestrators handle it by keeping retries inside a node, or by unrolling a fixed number of attempts. Agent frameworks that need real loops use a more general graph or state machine that allows cycles, with an explicit iteration cap. If your process genuinely loops, choose a tool built for cycles rather than fighting a DAG engine.
DAGs are the model behind many data-pipeline schedulers and behind build systems. Their strength is that the graph is inspectable before anything runs: you can see the critical path (the longest chain of dependencies, which sets the minimum run time) and spot accidental serial steps that could be parallel.
Conditional routing sends each input down a different path based on a rule, a classifier's label or a check on an earlier step's result, such as sending low-confidence extractions to human review.
A router is a decision point in a workflow. It looks at something known about the current run and picks the next step. The decision can come from plain code (the document is over 50 pages, so use the chunked path), from a cheap model call that classifies the input (billing, technical or cancellation), or from the result of a previous step (validation failed, so route to review). The paths after the router are ordinary workflow steps.
Routing is how a workflow handles variety without becoming an agent. A support assistant can classify a ticket into one of six intents and run a specialised prompt and tool set for each, which is usually more accurate than one giant prompt that tries to handle everything. It also controls cost: route easy inputs to a small, fast model and only hard ones to a larger model.
The quality of a route depends on the quality of the signal. Rules on hard facts (amount over a limit, missing field, file type) are reliable. A model classifier is only as good as its accuracy on your traffic, so measure it on a labelled sample, and give it an explicit "other" or "unsure" label rather than forcing a choice. Self-reported confidence from a model, such as "rate your certainty from 1 to 10", is poorly calibrated; prefer checks you can compute, such as validation passing, agreement between two extractions, or a retrieval score.
Every router needs a default path. Inputs that match no rule, or that the classifier marks unsure, should go somewhere safe, usually a human queue or a general fallback, rather than to whichever branch happens to be first in the code.
Fan-out splits work into independent branches that run in parallel, such as one model call per document. Fan-in waits for the branches and combines their results into one output, such as a single summary.
The pattern is the workflow form of map and reduce. Fan-out takes one input and creates many independent tasks: one per document, per chunk of a long file, per candidate answer, or per specialised reviewer. Fan-in collects the results and merges them, either with code (concatenate, vote, sum) or with another model call (summarise these ten summaries). Because the branches do not depend on each other, the wall-clock time is roughly the slowest branch rather than the sum of all branches.
In AI systems the pattern shows up constantly: summarising a 300-page contract by sections, running three different checks on one answer at once, generating five candidate outputs and choosing the best, or querying several retrieval sources in parallel. It is often the single biggest latency win available in a workflow.
The hard parts are limits and partial failure. Firing 500 model calls at once will hit provider rate limits, so fan-out needs a concurrency cap (a semaphore or a worker pool). Some branches will fail or time out, so the fan-in step must decide what to do: fail the whole run, proceed with the successful branches and note the gaps, or retry only the failed ones. For a summary, proceeding with 48 of 50 sections and saying so may be fine; for a financial reconciliation, a missing branch should fail the run.
Fan-in by model call has its own limit: the merged inputs must fit the context window, and quality drops when a model is asked to combine very many pieces at once. Large fan-outs usually merge in a tree, combining groups of five or ten, then combining those results.
Asynchronous execution starts a piece of work, returns immediately with a handle such as a job id, and delivers the result later by polling, a webhook, a stream or a notification, instead of holding the caller open until it finishes.
A synchronous request holds a connection open while the server does the work. That is fine for a 2-second chat reply, but many AI tasks take minutes: analysing a long PDF, generating a report, running an agent through twenty tool calls. Load balancers, browsers and mobile networks drop connections long before that, and a server holding thousands of open requests wastes memory on waiting.
The asynchronous pattern separates accepting work from doing it. The API validates the request, writes a job record, puts a message on a queue and returns 202 Accepted with a job id, all in milliseconds. A worker picks up the job, runs it, and writes the result. The client finds out by polling a status endpoint, by receiving a webhook, by listening on a server-sent events or WebSocket stream, or by an email or in-app notification. The job record, not the open connection, is the source of truth.
There are two meanings of async worth keeping apart. Async I/O inside one process (for example asyncio in Python) lets one worker wait on many model calls at once; it improves throughput but the caller still waits. Asynchronous jobs decouple the caller from the work entirely; the caller can disconnect and come back. Long AI tasks usually need both.
The trade-off is complexity. You now need a job store, status transitions, a way to report progress, cleanup of old results, and handling for jobs that a crashed worker never finished. Use it when work regularly exceeds a few tens of seconds, when load is bursty, or when the user does not need to watch the work happen.
Workflow state is the persisted record of one run: which steps have finished, what each produced, what is pending, and any data needed to continue. It lets the run be inspected, resumed after a crash and audited afterwards.
Every run of a workflow has a position (which step it is on) and data (inputs, intermediate outputs, decisions made). If that lives only in a process's memory, a crash, a deploy or a scale-down loses it, and the run must start over, repeating model calls you already paid for and side effects you already caused. Persisting state to a durable store such as a database row, a document or an orchestration engine's history is what makes a workflow robust.
A typical state record holds a run id, the workflow version, the overall status (queued, running, waiting, done, failed), and a per-step entry with status, attempt count, timestamps, and either the output or a pointer to it. Large outputs such as extracted text or generated files usually go to object storage, with only the reference in the state record. Including the model, prompt version and token counts per step pays for itself the first time someone asks why a run produced a strange result.
State changes should be transitions with rules, not arbitrary edits. A run cannot go from done back to running; a step cannot be marked done without an output. Writing transitions with a condition (update the row only if the status is still running) prevents two workers from both claiming the same step.
Distinguish workflow state from conversation state and from business data. The workflow state says the extraction step finished; the invoice record it created lives in the business tables. Keeping them separate lets you delete old run histories without losing business records, and lets you rerun a workflow without duplicating them.
Use a queue when work is slow, bursty or failure-prone and does not need to finish inside the user's request. The queue absorbs spikes, lets workers process at a controlled rate, and gives retries and dead-lettering a natural home.
A queue is a buffer between producers that create work and consumers (workers) that do it. The producer adds a message and moves on; workers pull messages at their own pace. Three properties make this valuable for AI workloads. Load levelling: a spike of 5,000 uploads becomes a backlog that workers drain at the rate the model provider allows, instead of 5,000 simultaneous calls that hit rate limits. Isolation: a slow or failing provider backs up the queue rather than taking down the web tier. Retry: a message that fails is returned to the queue and tried again later.
Most queues give at-least-once delivery: a worker takes a message, which becomes invisible for a visibility timeout (or lease); if the worker does not acknowledge it in time, the message reappears for another worker. That is how crashed work gets retried, and it also means a message can be processed twice, for example when a slow worker finishes just after its lease expired. So consumers must be idempotent: processing the same message twice must not send two emails or create two records. Set the visibility timeout longer than the slowest realistic job, or extend the lease while working.
Messages that fail repeatedly should not cycle forever. After a set number of attempts, move them to a dead-letter queue (DLQ) where someone can inspect them. For AI work, a poisoned message is often an input the model cannot handle, such as a corrupted PDF or a document over the context limit, and retrying it ten times only burns tokens.
Do not reach for a queue when the work is fast and the user is waiting for the answer anyway; a 2-second chat reply gains nothing from a queue except latency and moving parts. A queue also does not give you multi-step orchestration on its own. Chaining five queues by hand to build a workflow works, but state, retries and visibility end up scattered; past two or three steps, a workflow engine or a state table usually serves better.
Kafka helps when many independent systems need to read the same durable, ordered stream of events at high volume, and replay it later. For handing jobs to workers, a simple queue is usually enough and far easier to run.
Kafka is a distributed log, not a queue. Producers append events to a topic, which is split into partitions; events are kept for a retention period (days, weeks or indefinitely) whether or not anyone has read them. Each consumer group tracks its own position (offset) in each partition. That means five different systems can all read the same order events independently, at their own pace, and a new system added next month can replay history from the beginning. A queue, by contrast, usually deletes a message once one consumer acknowledges it.
This makes Kafka (and similar log-based systems) a good fit for event streaming: an order service publishes order.created and the billing, fraud, analytics and AI enrichment services each consume it. For GenAI systems, common uses are feeding a stream of new documents into embedding and indexing pipelines, capturing every model call and tool action as an event for audit and analytics, and rebuilding a derived store (such as a vector index) by replaying the log after you change the embedding model.
Ordering is guaranteed only within a partition, and events go to partitions by key. If all events for one customer share a key, they are processed in order for that customer; across customers there is no global order. Parallelism is capped by partition count, since one partition is consumed by at most one member of a group at a time. A single slow model call on one partition stalls everything behind it on that partition, which is a poor match for long, variable-length AI jobs.
That is why Kafka is rarely the right tool for distributing model work itself. Per-message retry, delays and dead-lettering are things you build on top, and running a cluster (or paying for a managed one) is real operational load. If the need is "give these 1,000 jobs to workers and retry failures", use a queue. If the need is "many systems react to the same facts, and we need the history", a log is worth it. Many architectures use both: Kafka carries business events, and a consumer turns the relevant ones into jobs on a work queue.
Persist each step's output keyed by run and step, check that record before running a step, and make every side effect idempotent. A restarted run then skips finished steps and safely repeats only the one that was in progress.
A run can be interrupted at any moment: a worker crashes, a deploy restarts the pod, the provider times out, a rate limit trips. Without resumability the only option is to start over, which repeats model calls (cost and latency), may produce different outputs the second time, and may repeat side effects such as emails or payments. Resumability rests on two ideas: checkpoint step outputs and make side effects safe to repeat.
Checkpointing means that before a step runs, the workflow looks up whether (run_id, step_name) already has a stored result. If it does, it uses that result and moves on. If not, it runs the step and stores the result before continuing. This is memoisation at step granularity. It also fixes the non-determinism problem: once the extraction output is stored, every later attempt of the run sees the same extraction, so downstream steps get consistent input. Choose step boundaries around expensive or external actions; a step that is cheap and pure can simply rerun.
The unavoidable gap is the step that was mid-flight at the crash. The workflow cannot know whether the email was sent just before the process died. So side effects need an idempotency key, a stable identifier derived from the run and step (not a random value generated per attempt), passed to the external system so it can recognise and ignore a duplicate. Many payment and messaging APIs accept such a key; for your own database, a unique constraint on the key, or an insert-if-absent, does the same job. Where an external system has no such feature, record an intent before the call and check the external system for the effect on retry.
Two further details make resume work in practice. Version the workflow and store the version in the run record, so a resumed run continues on the code it started with or a compatible one. And write a clear rule for when not to resume: if a run has been waiting for days, the data it read may be stale, so some steps should be re-validated rather than trusted from the checkpoint. Durable execution engines automate much of this, but the idempotency of external effects remains your responsibility.
In event-driven orchestration, something that happened (a file uploaded, a record changed, a step finished) triggers the next operation, rather than a central process calling each step in turn and waiting for it.
There are two broad ways to coordinate multi-step work. In orchestration, a central coordinator holds the plan: it calls classify, waits, then calls extract, then calls index. In choreography, there is no coordinator: each service listens for events and reacts. The upload service emits file.uploaded, the classifier reacts and emits file.classified, the extractor reacts to that, and so on. Event-driven designs can use either; the common thread is that work starts because an event arrived, not because something was polling or a person pressed a button.
The appeal is loose coupling and natural scaling. Services do not call each other directly, so a new consumer (say, an AI step that tags documents for compliance) can be added by subscribing to an existing event without changing the producer. Each consumer scales with its own backlog. Event triggers are also how AI work gets attached to existing systems: a new support ticket, a changed CRM record or a commit to a repository can each start a workflow.
The cost is visibility. In pure choreography, the end-to-end process exists only implicitly, spread across subscriptions. Answering "where is document 1029 and why has it not been indexed?" means correlating events across services, which needs a correlation id carried on every event and good tracing. Failure handling is also spread out: if the extractor fails, who notices that the process stalled? Events can arrive twice or out of order, so consumers need idempotency and must tolerate an event about something they have not seen yet.
A practical hybrid is common: use events to start workflows and to announce their results, and use an orchestrator (a workflow engine or a state table) to run the multi-step process in between. The event says a contract was uploaded; the orchestrator owns the six steps of review, including retries and human approval, and emits contract.reviewed at the end. You get loose coupling at the edges and a single place to see each run's status.
Use a visual tool such as n8n for integration-heavy flows that connect existing apps, change often, and benefit from people outside engineering seeing them. Use code when you need complex state, real testing, version control discipline, custom error handling or long-running agent logic.
Visual workflow tools such as n8n (and similar low-code automation platforms) let you build a flow by connecting nodes on a canvas: a trigger, a few app integrations, a model call, a branch, an action. Their strength is the library of prebuilt connectors and the speed of wiring them together. A flow that watches a form, enriches the lead with a model call, posts to a chat channel and creates a CRM record can be live in an afternoon, and an operations person can read and adjust it.
Code orchestration means the workflow is written in a general-purpose language, either directly or on top of an orchestration framework or durable execution engine. It wins as soon as the logic gets interesting: nested loops and conditions, typed data passed between steps, unit tests for each step, eval suites run in CI, code review on every change, and precise control of retries, idempotency and concurrency. It also wins for long-running agents, where the loop, tool permissions and budgets need careful engineering that a canvas makes awkward.
The usual failure is not the initial choice but the drift. A visual flow that started as five nodes grows to sixty, with expressions embedded in node settings, logic copied between flows and no tests. At that point it is code in a format that is hard to diff, review or refactor. The reverse failure also exists: an engineering team spends weeks writing integration glue for a simple notification flow that a visual tool would have handled.
Consider who maintains it, how much it changes, what breaks if it fails, and how it is governed. Many visual tools can be self-hosted and support exporting flows as JSON for version control, which helps, but check how credentials are stored and who can edit production flows. A sound split is common: prototype and run lightweight integrations visually, and move any flow that becomes business-critical, stateful or complex into code, often keeping the visual tool as the trigger layer that calls a coded service.
Build your own only when your control flow cannot be expressed cleanly in existing engines and your team can own a stateful runtime for years. Otherwise adopt an engine and build the domain layer on top.
Orchestration looks simple at first: a loop, a database table, a few retries. The hard parts appear later. You need durable state that survives crashes, leases so two workers do not take the same step, timers for waits that last days, retries with backoff, cancellation, visibility into stuck runs, migration of in-flight runs when the code changes, and correct behaviour when the database or network fails halfway through a transition. Proven engines encode years of fixes for these edge cases. A homegrown version usually rediscovers them one incident at a time.
The landscape has several categories, each with a different shape. Durable execution engines (for example Temporal, or cloud services such as AWS Step Functions and Azure Durable Functions) run long-lived workflows with retries, timers and waits. Data pipeline schedulers (Airflow, Dagster, Prefect and similar) run DAGs on schedules. Agent graph frameworks model loops and tool-calling agents with checkpointed state. Visual automation tools connect apps. Matching the category to the problem matters more than which product you choose within it.
Building your own is justified in a narrower set of cases: the core of your product is the orchestration itself (you sell a workflow platform); you have unusual constraints (an on-premise deployment where no engine can be installed, or strict latency where engine overhead per step is too high); or your workflows are simple and will stay simple, and a state table plus a queue covers them with less operational load than adopting an engine. The last case is real, but it should be a deliberate decision with a list of features you are choosing not to have.
If you do build, keep the scope minimal and the design boring. Store run and step state in your main database with conditional updates, use an existing queue for delivery, make steps idempotent, and add a sweeper that finds runs whose lease expired. Write the domain concepts (approval states, business rules, audit records) as your own code regardless; the engine is plumbing, the domain logic is yours either way.
Durable execution is a runtime model where a workflow's progress is recorded as it runs, so after any crash or deploy the workflow resumes exactly where it stopped, with completed steps' results restored rather than rerun.
In a durable execution engine you write the workflow as ordinary-looking code: call extract, then validate, then wait for approval, then save. Behind the scenes the engine records an event history for each run: step scheduled, step completed with this result, timer started, signal received. Work that touches the outside world, such as model calls, API requests and database writes, runs as separate units (often called activities) that the engine retries according to a policy. If the process running the workflow dies, another worker loads the history and replays the workflow code; completed activities return their recorded results instantly, and execution continues from the first unfinished step.
This gives you, by default, much of what resumable workflows otherwise need hand-built: checkpointing, retries with backoff, timeouts, durable timers (sleep for three days without a process sitting idle), waiting for an external signal such as a human approval, and a queryable history of every run. For AI systems with long agent runs, document batches and approval steps, that removes a large class of infrastructure work.
The main constraint is determinism of the workflow code. Because the code is replayed, it must make the same decisions given the same history. Anything non-deterministic, such as reading the clock, generating random ids, calling a model or reading a database, must happen inside an activity whose result is recorded, not in the workflow body. A model call in the workflow body would return a different answer on replay and corrupt the run. Engines provide deterministic substitutes for time and randomness, and some detect violations during replay.
Other trade-offs: activities still run at least once, so external effects still need idempotency keys; each recorded step adds latency and storage, which matters for sub-second interactive paths; long histories need compaction (some engines call it continue-as-new); and changing workflow code while runs are in flight needs explicit versioning. The precise programming model varies by engine, so read its determinism rules before writing the first workflow.
Model the approval as a durable pause: persist the run, notify an approver with the exact action proposed, wait for a signal with an expiry, then re-check the facts before acting. Never hold a process or connection open while a person decides.
An approval step sits between a proposed action and its execution: the model drafted a refund, an email to a customer, a database change, and a person must accept, edit or reject it before it happens. The person might answer in 30 seconds or on Monday. So the workflow must suspend without consuming resources, survive deploys while it waits, and resume when the decision arrives. A durable execution engine provides this as a wait-for-signal with a timeout; without one, store the run as waiting in a state table, and have the approval handler enqueue a resume message.
What the approver sees determines whether the approval means anything. Show the concrete action, not the model's reasoning alone: the exact refund amount and account, the exact email text, the diff of the record. Include the evidence the model used and anything that triggered the review (amount over limit, low agreement between extractions). Let the approver edit the proposal, and record the edit, since approved-with-changes is valuable training and eval data.
Three details are easy to miss. First, expiry: an approval that never arrives must not leave a run waiting forever; after a deadline, escalate, auto-reject or notify. Second, re-validation on resume: between proposal and approval, the world may change (the order was cancelled, the balance moved), so recheck preconditions before executing. Third, binding the approval to the action: store a hash of the approved payload and execute only that payload, so a later step cannot change the amount after approval. Record who approved, when, and what they saw, for audit.
Approval also has a cost: people become the bottleneck, and if most requests are approved without change, approvers start rubber-stamping. Route only the risky cases to review (by amount, reversibility or a failed check), measure the approval and edit rates, and lower the review threshold only when the data shows the automated path is reliable.
Version the workflow and record the version on every run. Let in-flight runs finish on the version they started with, or migrate them only through an explicit, tested compatibility path, and keep old step handlers until no run needs them.
Short workflows rarely notice this problem: a deploy happens, the few runs in progress fail or finish, and new runs use the new code. Long workflows do. A contract review waiting four days for approval, a batch of 10,000 documents processing over a weekend, or an agent run paused for user input will all be in flight when you ship a change. If the new code assumes a step that the old run never executed, or reads a state field with a new name, the resumed run fails or, worse, quietly does the wrong thing.
The base mechanism is versioning. Store the workflow version (and the prompt and schema versions each step used) in the run record. On resume, dispatch to the code for that version. Additive changes are the safe kind: a new optional field with a default, a new step only on a path old runs will not reach. Breaking changes, such as reordering steps, removing a step, renaming state fields or changing a step's output schema, need either a new workflow version that only new runs use, or a migration that converts old state to the new shape and is tested against real stored runs.
Durable execution engines make this sharper, because replay re-executes the workflow code against recorded history. Changing the order or number of steps in code changes what replay expects, and the engine will report non-determinism for runs that started on the old code. Engines offer version markers or patching APIs that let one workflow function branch on which version a run started with, plus worker routing so old runs go to workers with old code. The mechanics differ by engine, so read its guidance before the first breaking change rather than after.
Prompts and models are part of the version too. Changing a step's prompt mid-run can produce output in a slightly different format from the one earlier steps produced, which breaks a later merge. Treat the prompt and model as configuration pinned per run version, and run your eval suite against the new version before routing new runs to it. Plan for retirement: keep old handlers deployed until a query shows zero runs on that version, then remove them.