Catalog / Business processes / Revenue operations / Signal-to-System / Choose a Tool
reference process · revenue-operations · 18 activities · 14 on the roster
Choose a Tool
Somebody in the revenue organization reports work the current tools cannot carry, and this process ends with one tool chosen and written into the stack record. Every run names the process that would run inside the tool before any candidate is looked at, because a tool bought without a process to run in it is what this stream exists to prevent. The requirements and the criteria are fixed and dated first, the market is searched with one set of questions, and the security and procurement reviews run alongside the scoring rather than after it.
The walk
The document
One run
Adoption
The activities
What happens in a run
18 activities from a need the current tools cannot meet is raised to the tool is chosen, reviewed and recorded against a process, 3 of them gates a person has to sign. Drag the diagram to move along it.
Choose a Tool a need the current tools cannot meet is raised → the tool is chosen, reviewed and recorded against a process
no process would run inside it every candidate fails a requirement findings the candidate has to answer it cannot hold the definitions in force the trial shows it cannot do the work the cost rules every finalist out terms the organization cannot accept no finalist is worth taking the choice needs more evidence a need the current tools cannot meet is raised 1 Take in the Need say what the revenue organization cannot do today 2 Name the Process It Serves the process that would run inside the tool 3 Check the Stack for It read the stack record for something that does it 4 Write down the Requiremen… what the tool has to do, ranked and testable 5 Set the Scoring Criteria how a candidate is judged, fixed beforehand 6 Brief the Evaluation every agent hears the same requirements at once 7 Map What It Must Connect … the systems either side of it, and what flows 8 Search the Market find every candidate worth looking at 9 Cut to a Shortlist drop the candidates that miss a requirement 10 Score the Shortlist a verdict on every requirement, per candidate 11 Review the Security how it holds data and what it would be given a person signs · never an agent 12 Check It against the Stan… whether it can hold the definitions in force 13 Trial the Finalists on Re… the same work, the same period, measured once 14 Work out the Whole Cost licence, setup, the people, and the cost to leave 15 Run the Procurement Review price, term, exit, and what the supplier must meet a person signs · never an agent 16 Compare the Finalists scores, trial, security and cost side by side 17 Choose the Tool the named decider picks one and says why a person signs · never an agent 18 Write It into the Stack R… the process served, the owner, the cost, the date the tool is chosen, reviewed and recorded against a process
The roster this process needs
Hover a name to see the activities it holds. A dashed one is a person, and stays one.
The document
The document, with the blanks marked
Everything in amber is yours to fill in: who owns it, when it takes effect, which platforms, which numbers, and who holds each activity. Everything else is the process, and it is the same wherever it is run.
Use in LangGraph
Use in Agent Framework
Use in CrewAI
Use in Google ADK
PROCESS: choose a tool id: <team> /choose-a-tool v1
from: ref/rev/choose-a-tool v1
owner: <who> effective: <date>
trigger: a need the current tools cannot meet is raised in <your
planning process>
watch: record=<need> system=<your planning process>
change=<a need the current tools cannot meet is
raised>
concurrency: runs may overlap - <how many> searches open at once, and
two searches for the same need are merged into one
goal: one tool chosen against requirements fixed before any candidate
was seen, cleared by the security and the procurement reviews,
and recorded against the process it will run
phases:
take-in-need - human: read the need as the requester states it and
write down what the revenue organization cannot do
today
owner: stack-planner after: trigger
automation: <level>
name-process - human: the process that would run inside the tool,
named and pointed at. A need with no process behind
it goes back to the requester
owner: stack-planner after: take-in-need
by: <days> automation: <level>
check-stack - system: read the stack record for something already
licensed that does the job or comes near it
owner: stack-planner after: name-process
by: <days> automation: <level>
requirements - human: what the tool has to do, ranked, each one
written so a candidate can be tested against it
owner: stack-planner after: check-stack
by: <days> automation: <level>
set-criteria - human: how a candidate is judged and what a pass
takes, fixed at a version and dated
owner: supplier-check after: requirements
by: <days> automation: <level>
brief-evaluation - convenes briefing: every agent hears the same
requirements, the criteria and the due date at once
owner: stack-planner after: set-criteria
automation: <level>
map-connections - human: the systems that would feed the tool, the
systems that would read it, and what would flow
owner: integration-keeper after: brief-evaluation
by: <days> automation: <level>
search-market - runs collect-and-report: one set of questions in one
format to every candidate, every answer kept against
the candidate that gave it
owner: researcher after: brief-evaluation
by: <weeks> automation: <level>
shortlist - human: the candidates that miss a requirement are
dropped, each with the requirement it failed
owner: stack-planner
after: map-connections + search-market
by: <days> automation: <level>
score-shortlist - runs assessment: every requirement gets a verdict
for every candidate, against the list at the version
it was fixed at
owner: supplier-check after: shortlist
by: <days> automation: <level>
review-security - human: how the candidate holds the data, what it
would be given, and who at the supplier can read it
owner: <your security role> after: score-shortlist
by: <days> automation: never
check-standards - runs assessment: whether the tool can hold the
definitions, the records and the retention rules in
force
owner: standards-keeper after: score-shortlist
by: <days> automation: <level>
trial-finalists - runs assessment: each finalist takes the same work
the team actually has, over the same period,
measured the same way
owner: stack-planner after: score-shortlist
by: <weeks> automation: <level>
whole-cost - human: the licence, the setup, the people it takes
to run, and what leaving would cost
owner: <your finance role> after: trial-finalists
by: <days> automation: <level>
procurement-review - convenes approval: the price, the term, the notice
a cancellation needs, and what the supplier has to
meet. Procurement and legal sign as named signers
owner: <your procurement role>
after: whole-cost
by: <weeks> automation: never
compare - human: the finalists side by side against one list,
with what taking each one would rule out
owner: stack-planner
after: review-security + check-standards +
procurement-review
by: <days> automation: <level>
choose-tool - convenes decide-and-announce: <who> picks one, says
why, and every agent waiting on it hears at the same
time
owner: stack-planner after: compare
by: <days> automation: never
record-entry - human: the entry lands with the process it serves,
its owner, its cost, its term and its renewal date
owner: stack-planner after: choose-tool
automation: <level>
run-scoped:
status - runs roll-call owner: stack-planner
every: <cadence>
from: brief-evaluation until: choose-tool
money - runs allocate-and-reconcile owner: <your finance role>
from: whole-cost until: run close
handoffs:
take-in-need -> name-process [the-need]: what the revenue
organization cannot do today, in one line the requirements are
written from
name-process -> check-stack [process-served]: the process that would
run inside the tool, and the activities in it the tool would carry
check-stack -> requirements [near-misses-in-stack]: what the stack
already holds that comes near, and what each of those falls short
on
requirements -> set-criteria [ranked-requirements]: the requirements
ranked, each written so a candidate can be tested against it
set-criteria -> brief-evaluation [fixed-criteria]: the requirements
and the criteria at a version, dated before any candidate was
looked at
brief-evaluation -> map-connections [process-systems]: the process
the tool serves and the systems that process already reads and
writes
brief-evaluation -> search-market [evaluation-brief]: the
requirements as every agent heard them, and the date the choice is
due
map-connections -> shortlist [required-connections]: what would have
to flow in and out, and the connections a candidate would have to
support
search-market -> shortlist [candidates-found]: every candidate found,
with where it came from and what it answered, and a candidate that
did not answer recorded as missing
shortlist -> score-shortlist [shortlisted-candidates]: the candidates
left, and the requirement each dropped candidate failed
score-shortlist -> trial-finalists [requirement-verdicts]: a verdict
per requirement per candidate with the evidence behind it, and
every requirement that could not be tested marked as not tested
review-security -> compare [security-findings]: how the candidate
holds the data, what it would be given, and every finding it has to
answer
check-standards -> compare [standards-verdict]: whether the tool can
hold the definitions and the retention rules, with the rule behind
every failure
trial-finalists -> whole-cost [trial-results]: what each finalist did
on the same work, with the figures and the range around each one
whole-cost -> procurement-review [cost-of-ownership]: the licence,
the setup, the people it takes to run, and what leaving would cost
procurement-review -> compare [granted-terms]: the terms as the
supplier will grant them, and every term the organization will not
accept
compare -> choose-tool [side-by-side]: the finalists set side by side
against one list, and what taking each one would rule out
choose-tool -> record-entry [the-choice]: the chosen tool, who
decided, on what date, for what reasons, and anyone who disagreed
deviations:
name-process -> take-in-need [no-process-behind-it]: no process would
run inside the tool, so the need goes back to the requester as a
way of working rather than as a thing to buy
score-shortlist -> requirements [all-candidates-fail]: every
candidate fails a requirement, so the requirements are written
again with the decision and the scores behind it recorded
review-security -> score-shortlist [findings-to-answer]: the security
review raises findings the candidate has to answer, so the scoring
is worked again while the evaluation is still open
check-standards -> shortlist [cannot-hold-definitions]: the tool
cannot hold the definitions in force, so the candidate drops off
and the shortlist is cut again
trial-finalists -> shortlist [trial-fails]: the trial shows a
finalist cannot do the work, so the shortlist is cut again
whole-cost -> search-market [cost-rules-all-out]: the whole cost
rules every finalist out, so the market is searched again
procurement-review -> shortlist [terms-refused]: the organization
cannot accept the terms, so the shortlist is cut again and the next
candidate is taken
compare -> search-market [no-finalist-worth-taking]: no finalist is
worth taking, so the search reopens rather than the least bad one
being chosen
choose-tool -> trial-finalists [more-evidence-needed]: the choice
needs more evidence, so the finalists are trialled again on the
same work
bindings:
roster: <who holds each role - agents claiming the abstract agents
above, and named people for the requester, the systems owner,
the data owner, the security reviewer, procurement, legal,
finance and the revenue operations lead>
systems: the stack record (write), the requirements record (write),
the scoring record (write), the decision record (write),
the budget record (write), a trial environment (read),
the process register (read)
data: the requirements and the criteria at the version they were
fixed at, <your security and privacy standards> <version> ,
<your data retention rules> , the definitions in force,
the stack record as it stands
policy:
no run goes past name-process without a process named that the tool
would run inside
the requirements and the criteria are written down and dated before
any candidate is looked at
every candidate is scored against the same list, and a requirement
that could not be tested is recorded as not tested
a change to the requirements after candidates have been seen goes in
as a new version, and every score made against the earlier version
is marked
a candidate with a relationship to the organization or to anyone
deciding is declared before scoring starts
no agreement is entered into by an agent. Named people sign it and
<your finance role> commits the money
the security review and the procurement review are human gates and
are never delegated to an agent
no entry goes into the stack record without a named owner, a renewal
date and the process it serves
measures:
cycle time: <target> from the need to the recorded entry
coverage: <share> of the requirements carrying a tested verdict on the
tool that was chosen
quality gate: nothing is chosen without a scored comparison behind it,
a cleared security review and a named person who decided
Copy
Take it somewhere
Use this process in LangGraph
close
Paste this into an assistant that can read the web, such as Claude, ChatGPT or Cursor. It reads the specification and the current LangGraph documentation, then writes two files: the graph, and a note on what did not survive the translation. Read the note first. What a runtime cannot express is the part worth arguing about, and this process is a draft to argue with.
copy the prompt
338 lines · the document is inside it, so nothing else is needed
Convert the business process below into a runnable LangGraph graph:
one Python file with a TypedDict state, a StateGraph, nodes, edges,
conditional edges and a checkpointer.
The document is a reference process written to the Agent Processes
specification. Read the specification before you start, because it defines
terms that look ordinary and are not:
https://agentcatalog.com/spec/agent-processes
Sections 6 (the phase graph), 6.5.1 (exception edges), 6.7.1 (deviations),
7 (automation) and 8 (handoffs) are the ones this conversion turns on.
Then read the current documentation for the primitives you will need, rather
than relying on what you remember of the API:
https://docs.langchain.com/oss/python/langgraph/interrupts
interrupt() and Command(resume=), which is how a gate stops a run
https://reference.langchain.com/python/langgraph/graph/state/StateGraph
StateGraph, add_edge, add_conditional_edges, defer
WHAT THE DOCUMENT ASKS FOR
These hold wherever the process lands, and they matter more than style.
1. Each phase under `phases:` becomes one step, and keeps its name.
2. `after:` gives the edges. `after: a + b` is a join and waits for BOTH.
Reading it as "either" is the defect the specification calls out by name.
3. Every handoff carries a key in square brackets. Each key becomes one field
on the run's state, named exactly as the key with hyphens turned into
underscores, and the sentence beside it becomes that field's comment. The key
is the stable name; the sentence is prose that may be rewritten.
4. A phase MUST NOT begin before its inbound handoff exists. Where that is
checkable, check it in the step rather than assuming it.
5. `automation: never` is a gate a person signs. The run stops there and does
not continue until a person's decision comes back. Do not turn one into a
notification, a log line, or an automatic transition, whatever the queue
looks like.
6. Each line under `deviations:` is a backward or sideways edge, returning to
the phase named on the right. The key in brackets names it, and that name
belongs in the code.
7. A phase whose `after:` reads like "X or Y, whichever could not finish" is an
exception edge: it is entered when those phases FAIL, not when they succeed.
Do not wire it as an ordinary successor.
8. Anything in angle brackets is a blank the adopting organization fills in.
Leave each one as a named constant at the top of the file with a TODO. Do not
invent a value, a threshold or a date.
9. Record the document's `from:` line at the top of the file, so it says which
reference process and which version it was generated from.
10. Run-scoped lines under `run-scoped:` are work that runs alongside the whole
process rather than at one point in it, and a run may not close while one is
unfinished. Say in the code what you did about them, including if the answer
is that the runtime has nowhere to put them.
HOW THAT LOOKS IN LANGGRAPH
11. A phase is a node added with `add_node`, under the phase's own name.
12. The state is a TypedDict. Each handoff key is one field on it.
13. A join is the trap. `add_edge(["a", "b"], "c")` looks right and releases
once: when a backward edge re-enters ONE arm, the joined node never runs
again, and the run ends early reporting success rather than raising. Mark the
joined node `defer=True` and re-check inside it that both inbound handoffs
exist.
14. A gate is `interrupt()` inside the node, resumed with `Command(resume=...)`.
The platform lets anything at all call resume, so require the resumed value to
name a person and a date and refuse anything else. Say in the fidelity note
that this proves only that whoever resumed typed a name, because
`Command(resume=True)` from a scheduled job is indistinguishable from a person
signing.
15. A deviation is `add_conditional_edges` with a routing function named after
the key in brackets.
16. Pass a durable checkpointer rather than taking the in-memory default. The
gates wait days, and the default loses every paused run on restart.
17. Leave every phase body unimplemented, raising until somebody registers an
implementation. The automation level is a blank, so writing a body would
answer on the adopter's behalf whether an agent may do that work.
Produce a second file alongside it, `FIDELITY.md`, and treat it as the more
important of the two. The code is for whoever builds this. The fidelity note
is for whoever has to decide whether this platform suits the process at all,
and that is usually a different person who will never read the code.
It has three parts.
**What came across.** Briefly: how many phases became steps, how many handoff
keys became state fields, which gates stop the run, which deviations became
edges. Counts and names, not reassurance.
**What did not, and what was done instead.** One entry per gap. For each one,
say what the document requires, what the platform can actually express, what
you did in its place, and what breaks if somebody later removes your
workaround. This last part matters most: a workaround nobody understands is a
workaround somebody deletes.
**What a person still has to decide.** The blanks are not a translation
failure, they are the point of a reference process, so list what has to be
filled in before this could run against anything real, and say which of those
choices the platform constrains.
Write it in plain English for somebody who has not read the specification, and
do not soften it. A translation of a reference process is a draft to argue
with, not a build artifact, and the honest account of what was lost is the most
useful thing you will produce.
Here is the process document.
```
PROCESS: choose a tool id: <team>/choose-a-tool v1
from: ref/rev/choose-a-tool v1
owner: <who> effective: <date>
trigger: a need the current tools cannot meet is raised in <your
planning process>
watch: record=<need> system=<your planning process>
change=<a need the current tools cannot meet is
raised>
concurrency: runs may overlap - <how many> searches open at once, and
two searches for the same need are merged into one
goal: one tool chosen against requirements fixed before any candidate
was seen, cleared by the security and the procurement reviews,
and recorded against the process it will run
phases:
take-in-need - human: read the need as the requester states it and
write down what the revenue organization cannot do
today
owner: stack-planner after: trigger
automation: <level>
name-process - human: the process that would run inside the tool,
named and pointed at. A need with no process behind
it goes back to the requester
owner: stack-planner after: take-in-need
by: <days> automation: <level>
check-stack - system: read the stack record for something already
licensed that does the job or comes near it
owner: stack-planner after: name-process
by: <days> automation: <level>
requirements - human: what the tool has to do, ranked, each one
written so a candidate can be tested against it
owner: stack-planner after: check-stack
by: <days> automation: <level>
set-criteria - human: how a candidate is judged and what a pass
takes, fixed at a version and dated
owner: supplier-check after: requirements
by: <days> automation: <level>
brief-evaluation - convenes briefing: every agent hears the same
requirements, the criteria and the due date at once
owner: stack-planner after: set-criteria
automation: <level>
map-connections - human: the systems that would feed the tool, the
systems that would read it, and what would flow
owner: integration-keeper after: brief-evaluation
by: <days> automation: <level>
search-market - runs collect-and-report: one set of questions in one
format to every candidate, every answer kept against
the candidate that gave it
owner: researcher after: brief-evaluation
by: <weeks> automation: <level>
shortlist - human: the candidates that miss a requirement are
dropped, each with the requirement it failed
owner: stack-planner
after: map-connections + search-market
by: <days> automation: <level>
score-shortlist - runs assessment: every requirement gets a verdict
for every candidate, against the list at the version
it was fixed at
owner: supplier-check after: shortlist
by: <days> automation: <level>
review-security - human: how the candidate holds the data, what it
would be given, and who at the supplier can read it
owner: <your security role> after: score-shortlist
by: <days> automation: never
check-standards - runs assessment: whether the tool can hold the
definitions, the records and the retention rules in
force
owner: standards-keeper after: score-shortlist
by: <days> automation: <level>
trial-finalists - runs assessment: each finalist takes the same work
the team actually has, over the same period,
measured the same way
owner: stack-planner after: score-shortlist
by: <weeks> automation: <level>
whole-cost - human: the licence, the setup, the people it takes
to run, and what leaving would cost
owner: <your finance role> after: trial-finalists
by: <days> automation: <level>
procurement-review - convenes approval: the price, the term, the notice
a cancellation needs, and what the supplier has to
meet. Procurement and legal sign as named signers
owner: <your procurement role>
after: whole-cost
by: <weeks> automation: never
compare - human: the finalists side by side against one list,
with what taking each one would rule out
owner: stack-planner
after: review-security + check-standards +
procurement-review
by: <days> automation: <level>
choose-tool - convenes decide-and-announce: <who> picks one, says
why, and every agent waiting on it hears at the same
time
owner: stack-planner after: compare
by: <days> automation: never
record-entry - human: the entry lands with the process it serves,
its owner, its cost, its term and its renewal date
owner: stack-planner after: choose-tool
automation: <level>
run-scoped:
status - runs roll-call owner: stack-planner
every: <cadence>
from: brief-evaluation until: choose-tool
money - runs allocate-and-reconcile owner: <your finance role>
from: whole-cost until: run close
handoffs:
take-in-need -> name-process [the-need]: what the revenue
organization cannot do today, in one line the requirements are
written from
name-process -> check-stack [process-served]: the process that would
run inside the tool, and the activities in it the tool would carry
check-stack -> requirements [near-misses-in-stack]: what the stack
already holds that comes near, and what each of those falls short
on
requirements -> set-criteria [ranked-requirements]: the requirements
ranked, each written so a candidate can be tested against it
set-criteria -> brief-evaluation [fixed-criteria]: the requirements
and the criteria at a version, dated before any candidate was
looked at
brief-evaluation -> map-connections [process-systems]: the process
the tool serves and the systems that process already reads and
writes
brief-evaluation -> search-market [evaluation-brief]: the
requirements as every agent heard them, and the date the choice is
due
map-connections -> shortlist [required-connections]: what would have
to flow in and out, and the connections a candidate would have to
support
search-market -> shortlist [candidates-found]: every candidate found,
with where it came from and what it answered, and a candidate that
did not answer recorded as missing
shortlist -> score-shortlist [shortlisted-candidates]: the candidates
left, and the requirement each dropped candidate failed
score-shortlist -> trial-finalists [requirement-verdicts]: a verdict
per requirement per candidate with the evidence behind it, and
every requirement that could not be tested marked as not tested
review-security -> compare [security-findings]: how the candidate
holds the data, what it would be given, and every finding it has to
answer
check-standards -> compare [standards-verdict]: whether the tool can
hold the definitions and the retention rules, with the rule behind
every failure
trial-finalists -> whole-cost [trial-results]: what each finalist did
on the same work, with the figures and the range around each one
whole-cost -> procurement-review [cost-of-ownership]: the licence,
the setup, the people it takes to run, and what leaving would cost
procurement-review -> compare [granted-terms]: the terms as the
supplier will grant them, and every term the organization will not
accept
compare -> choose-tool [side-by-side]: the finalists set side by side
against one list, and what taking each one would rule out
choose-tool -> record-entry [the-choice]: the chosen tool, who
decided, on what date, for what reasons, and anyone who disagreed
deviations:
name-process -> take-in-need [no-process-behind-it]: no process would
run inside the tool, so the need goes back to the requester as a
way of working rather than as a thing to buy
score-shortlist -> requirements [all-candidates-fail]: every
candidate fails a requirement, so the requirements are written
again with the decision and the scores behind it recorded
review-security -> score-shortlist [findings-to-answer]: the security
review raises findings the candidate has to answer, so the scoring
is worked again while the evaluation is still open
check-standards -> shortlist [cannot-hold-definitions]: the tool
cannot hold the definitions in force, so the candidate drops off
and the shortlist is cut again
trial-finalists -> shortlist [trial-fails]: the trial shows a
finalist cannot do the work, so the shortlist is cut again
whole-cost -> search-market [cost-rules-all-out]: the whole cost
rules every finalist out, so the market is searched again
procurement-review -> shortlist [terms-refused]: the organization
cannot accept the terms, so the shortlist is cut again and the next
candidate is taken
compare -> search-market [no-finalist-worth-taking]: no finalist is
worth taking, so the search reopens rather than the least bad one
being chosen
choose-tool -> trial-finalists [more-evidence-needed]: the choice
needs more evidence, so the finalists are trialled again on the
same work
bindings:
roster: <who holds each role - agents claiming the abstract agents
above, and named people for the requester, the systems owner,
the data owner, the security reviewer, procurement, legal,
finance and the revenue operations lead>
systems: the stack record (write), the requirements record (write),
the scoring record (write), the decision record (write),
the budget record (write), a trial environment (read),
the process register (read)
data: the requirements and the criteria at the version they were
fixed at, <your security and privacy standards> <version>,
<your data retention rules>, the definitions in force,
the stack record as it stands
policy:
no run goes past name-process without a process named that the tool
would run inside
the requirements and the criteria are written down and dated before
any candidate is looked at
every candidate is scored against the same list, and a requirement
that could not be tested is recorded as not tested
a change to the requirements after candidates have been seen goes in
as a new version, and every score made against the earlier version
is marked
a candidate with a relationship to the organization or to anyone
deciding is declared before scoring starts
no agreement is entered into by an agent. Named people sign it and
<your finance role> commits the money
the security review and the procurement review are human gates and
are never delegated to an agent
no entry goes into the stack record without a named owner, a renewal
date and the process it serves
measures:
cycle time: <target> from the need to the recorded entry
coverage: <share> of the requirements carrying a tested verdict on the
tool that was chosen
quality gate: nothing is chosen without a scored comparison behind it,
a cleared security review and a named person who decided
```
Take it somewhere
Use this process in Microsoft Agent Framework
close
Paste this into an assistant that can read the web, such as Claude, ChatGPT or Cursor. It reads the specification and the current Agent Framework documentation, then writes two files: the workflow, and a note on what did not survive the translation. Read the note first. What a runtime cannot express is the part worth arguing about, and this process is a draft to argue with.
copy the prompt
393 lines · the document is inside it, so nothing else is needed
Convert the business process below into a runnable Microsoft Agent Framework workflow:
one Python file with Executor classes, a WorkflowBuilder, typed edges,
request_info gates and durable checkpoint storage.
The document is a reference process written to the Agent Processes
specification. Read the specification before you start, because it defines
terms that look ordinary and are not:
https://agentcatalog.com/spec/agent-processes
Sections 6 (the phase graph), 6.5.1 (exception edges), 6.7.1 (deviations),
7 (automation) and 8 (handoffs) are the ones this conversion turns on.
Then read the current documentation for the primitives you will need, rather
than relying on what you remember of the API:
https://learn.microsoft.com/en-us/agent-framework/workflows/human-in-the-loop
ctx.request_info, @response_handler, and answering a parked run later
https://learn.microsoft.com/en-us/agent-framework/workflows/checkpoints
what a checkpoint holds, and allowed_checkpoint_types
https://learn.microsoft.com/en-us/agent-framework/concepts/workflows/edges
add_edge with condition, add_fan_in_edges, add_switch_case_edge_group
https://learn.microsoft.com/en-us/agent-framework/concepts/workflows/state
ctx.set_state and ctx.get_state as they actually are today
WHAT THE DOCUMENT ASKS FOR
These hold wherever the process lands, and they matter more than style.
1. Each phase under `phases:` becomes one step, and keeps its name.
2. `after:` gives the edges. `after: a + b` is a join and waits for BOTH.
Reading it as "either" is the defect the specification calls out by name.
3. Every handoff carries a key in square brackets. Each key becomes one field
on the run's state, named exactly as the key with hyphens turned into
underscores, and the sentence beside it becomes that field's comment. The key
is the stable name; the sentence is prose that may be rewritten.
4. A phase MUST NOT begin before its inbound handoff exists. Where that is
checkable, check it in the step rather than assuming it.
5. `automation: never` is a gate a person signs. The run stops there and does
not continue until a person's decision comes back. Do not turn one into a
notification, a log line, or an automatic transition, whatever the queue
looks like.
6. Each line under `deviations:` is a backward or sideways edge, returning to
the phase named on the right. The key in brackets names it, and that name
belongs in the code.
7. A phase whose `after:` reads like "X or Y, whichever could not finish" is an
exception edge: it is entered when those phases FAIL, not when they succeed.
Do not wire it as an ordinary successor.
8. Anything in angle brackets is a blank the adopting organization fills in.
Leave each one as a named constant at the top of the file with a TODO. Do not
invent a value, a threshold or a date.
9. Record the document's `from:` line at the top of the file, so it says which
reference process and which version it was generated from.
10. Run-scoped lines under `run-scoped:` are work that runs alongside the whole
process rather than at one point in it, and a run may not close while one is
unfinished. Say in the code what you did about them, including if the answer
is that the runtime has nowhere to put them.
HOW THAT LOOKS IN MICROSOFT AGENT FRAMEWORK
11. Write for Python, and read those pages before writing a line. This API has
moved: `set_shared_state`, `RequestInfoExecutor`, `RequestInfoMessage` and
`send_responses_streaming` are all in the training data and none of them exist
any more. The .NET workflow API differs in kind rather than in spelling, so a
file written for one does not port by renaming.
12. A phase is a class deriving from `Executor` whose `super().__init__(id=...)`
takes the phase name verbatim. The id is not cosmetic: a checkpoint stores a
signature over the topology and the executor ids, so an id built from a run, a
timestamp or a counter cannot be resumed into.
13. Give each phase a `@handler` method and move work on with
`await ctx.send_message(...)`. A handler that returns without sending is a dead
end: the branch stops, the run converges, and it reports success. So a phase
you are deliberately leaving unimplemented must still send a placeholder
onward. Do not stub with `raise NotImplementedError`, which fails the run
instead of leaving it runnable.
14. `after: a` is `builder.add_edge(a, b)`. `after: a + b` is
`builder.add_fan_in_edges([a, b], target)`, and the target's handler must be
annotated `list[T]`, because a fan-in delivers one aggregated list rather than
the separate messages. A handler typed for the single value is dropped as a
mismatch with nothing raised.
15. The join is the trap here, and it is the opposite of LangGraph's. The
barrier re-arms: it clears its buffer when it fires and then demands a fresh
message from every source. So a deviation that re-enters ONE arm parks the
run forever waiting for an arm that will not run again, and the workflow ends
IDLE reporting success. Wherever a deviation re-enters one arm of a join,
replace the barrier with an ordinary edge from each arm into a small executor
that records each arrival with `ctx.set_state` and only forwards when every
expected key is present, and say in a comment that putting `add_fan_in_edges`
back reintroduces the stall.
16. A phase that is both a join target and a deviation target needs two
handlers, one annotated `list[T]` for the barrier and one annotated `T` for the
backward message. Write only the list handler and every backward edge into it
is discarded as a type mismatch, silently.
17. `automation: never` is `await ctx.request_info(request_data=...,
response_type=...)` inside the phase, answered by a `@response_handler` on the
same executor whose annotations match those exact types. The run parks at
`IDLE_WITH_PENDING_REQUESTS` and the host answers with
`workflow.run(stream=True, responses={request_id: value})`. If no handler
matches the pair, the framework logs a warning and parks anyway, so the gate
reads as working right up until somebody asks why the approval did not take.
18. A gate is only a gate if the wait survives a restart, so pass
`checkpoint_storage=FileCheckpointStorage(...)` to the builder. Checkpointing
is off by default and `InMemoryCheckpointStorage` reads as configured while
persisting nothing. Register every handoff payload type in
`allowed_checkpoint_types`, or the first restore raises. Anything an executor
keeps as an instance attribute is absent after a restore unless you export it
from `on_checkpoint_save` and read it back in `on_checkpoint_restore`, and it
comes back empty rather than missing.
19. Each line under `deviations:` is `builder.add_edge(source, earlier,
condition=fn)` with `fn` named after the key. Cycles are legal and unchecked,
but raise `max_iterations` well above its default of 100, because several live
cycles will exhaust a budget sized for a straight line and fail with a message
about convergence that reads like a broken graph. Keep the ordinary forward
edge unconditional and add each deviation beside it: a condition that returns
false is dropped with no event, so a forward path expressed as a condition
dies silently on every normal run, which is most of them.
20. Handoff values go in `ctx.set_state(key, value)` and come back from
`ctx.get_state(key)`, untyped and unchecked. A write is visible to its writer
at once and to everyone else only in the next superstep, and two writers of one
key in a superstep keep the last write. Never read a key in the same superstep
another phase wrote it.
21. There are no timers, no deadlines and no scheduled wakes. Nothing in
`run-scoped:` becomes an executor and `by:` has no expression at all, so write
them as comments naming where they start and stop, and say plainly in the
fidelity note that a run can close over an unfinished run-scoped line. Do not
fake a deadline with a sleep inside a handler, which blocks the whole superstep
barrier, and never let an expiring wait release a gate.
Produce a second file alongside it, `FIDELITY.md`, and treat it as the more
important of the two. The code is for whoever builds this. The fidelity note
is for whoever has to decide whether this platform suits the process at all,
and that is usually a different person who will never read the code.
It has three parts.
**What came across.** Briefly: how many phases became steps, how many handoff
keys became state fields, which gates stop the run, which deviations became
edges. Counts and names, not reassurance.
**What did not, and what was done instead.** One entry per gap. For each one,
say what the document requires, what the platform can actually express, what
you did in its place, and what breaks if somebody later removes your
workaround. This last part matters most: a workaround nobody understands is a
workaround somebody deletes.
**What a person still has to decide.** The blanks are not a translation
failure, they are the point of a reference process, so list what has to be
filled in before this could run against anything real, and say which of those
choices the platform constrains.
Write it in plain English for somebody who has not read the specification, and
do not soften it. A translation of a reference process is a draft to argue
with, not a build artifact, and the honest account of what was lost is the most
useful thing you will produce.
Here is the process document.
```
PROCESS: choose a tool id: <team>/choose-a-tool v1
from: ref/rev/choose-a-tool v1
owner: <who> effective: <date>
trigger: a need the current tools cannot meet is raised in <your
planning process>
watch: record=<need> system=<your planning process>
change=<a need the current tools cannot meet is
raised>
concurrency: runs may overlap - <how many> searches open at once, and
two searches for the same need are merged into one
goal: one tool chosen against requirements fixed before any candidate
was seen, cleared by the security and the procurement reviews,
and recorded against the process it will run
phases:
take-in-need - human: read the need as the requester states it and
write down what the revenue organization cannot do
today
owner: stack-planner after: trigger
automation: <level>
name-process - human: the process that would run inside the tool,
named and pointed at. A need with no process behind
it goes back to the requester
owner: stack-planner after: take-in-need
by: <days> automation: <level>
check-stack - system: read the stack record for something already
licensed that does the job or comes near it
owner: stack-planner after: name-process
by: <days> automation: <level>
requirements - human: what the tool has to do, ranked, each one
written so a candidate can be tested against it
owner: stack-planner after: check-stack
by: <days> automation: <level>
set-criteria - human: how a candidate is judged and what a pass
takes, fixed at a version and dated
owner: supplier-check after: requirements
by: <days> automation: <level>
brief-evaluation - convenes briefing: every agent hears the same
requirements, the criteria and the due date at once
owner: stack-planner after: set-criteria
automation: <level>
map-connections - human: the systems that would feed the tool, the
systems that would read it, and what would flow
owner: integration-keeper after: brief-evaluation
by: <days> automation: <level>
search-market - runs collect-and-report: one set of questions in one
format to every candidate, every answer kept against
the candidate that gave it
owner: researcher after: brief-evaluation
by: <weeks> automation: <level>
shortlist - human: the candidates that miss a requirement are
dropped, each with the requirement it failed
owner: stack-planner
after: map-connections + search-market
by: <days> automation: <level>
score-shortlist - runs assessment: every requirement gets a verdict
for every candidate, against the list at the version
it was fixed at
owner: supplier-check after: shortlist
by: <days> automation: <level>
review-security - human: how the candidate holds the data, what it
would be given, and who at the supplier can read it
owner: <your security role> after: score-shortlist
by: <days> automation: never
check-standards - runs assessment: whether the tool can hold the
definitions, the records and the retention rules in
force
owner: standards-keeper after: score-shortlist
by: <days> automation: <level>
trial-finalists - runs assessment: each finalist takes the same work
the team actually has, over the same period,
measured the same way
owner: stack-planner after: score-shortlist
by: <weeks> automation: <level>
whole-cost - human: the licence, the setup, the people it takes
to run, and what leaving would cost
owner: <your finance role> after: trial-finalists
by: <days> automation: <level>
procurement-review - convenes approval: the price, the term, the notice
a cancellation needs, and what the supplier has to
meet. Procurement and legal sign as named signers
owner: <your procurement role>
after: whole-cost
by: <weeks> automation: never
compare - human: the finalists side by side against one list,
with what taking each one would rule out
owner: stack-planner
after: review-security + check-standards +
procurement-review
by: <days> automation: <level>
choose-tool - convenes decide-and-announce: <who> picks one, says
why, and every agent waiting on it hears at the same
time
owner: stack-planner after: compare
by: <days> automation: never
record-entry - human: the entry lands with the process it serves,
its owner, its cost, its term and its renewal date
owner: stack-planner after: choose-tool
automation: <level>
run-scoped:
status - runs roll-call owner: stack-planner
every: <cadence>
from: brief-evaluation until: choose-tool
money - runs allocate-and-reconcile owner: <your finance role>
from: whole-cost until: run close
handoffs:
take-in-need -> name-process [the-need]: what the revenue
organization cannot do today, in one line the requirements are
written from
name-process -> check-stack [process-served]: the process that would
run inside the tool, and the activities in it the tool would carry
check-stack -> requirements [near-misses-in-stack]: what the stack
already holds that comes near, and what each of those falls short
on
requirements -> set-criteria [ranked-requirements]: the requirements
ranked, each written so a candidate can be tested against it
set-criteria -> brief-evaluation [fixed-criteria]: the requirements
and the criteria at a version, dated before any candidate was
looked at
brief-evaluation -> map-connections [process-systems]: the process
the tool serves and the systems that process already reads and
writes
brief-evaluation -> search-market [evaluation-brief]: the
requirements as every agent heard them, and the date the choice is
due
map-connections -> shortlist [required-connections]: what would have
to flow in and out, and the connections a candidate would have to
support
search-market -> shortlist [candidates-found]: every candidate found,
with where it came from and what it answered, and a candidate that
did not answer recorded as missing
shortlist -> score-shortlist [shortlisted-candidates]: the candidates
left, and the requirement each dropped candidate failed
score-shortlist -> trial-finalists [requirement-verdicts]: a verdict
per requirement per candidate with the evidence behind it, and
every requirement that could not be tested marked as not tested
review-security -> compare [security-findings]: how the candidate
holds the data, what it would be given, and every finding it has to
answer
check-standards -> compare [standards-verdict]: whether the tool can
hold the definitions and the retention rules, with the rule behind
every failure
trial-finalists -> whole-cost [trial-results]: what each finalist did
on the same work, with the figures and the range around each one
whole-cost -> procurement-review [cost-of-ownership]: the licence,
the setup, the people it takes to run, and what leaving would cost
procurement-review -> compare [granted-terms]: the terms as the
supplier will grant them, and every term the organization will not
accept
compare -> choose-tool [side-by-side]: the finalists set side by side
against one list, and what taking each one would rule out
choose-tool -> record-entry [the-choice]: the chosen tool, who
decided, on what date, for what reasons, and anyone who disagreed
deviations:
name-process -> take-in-need [no-process-behind-it]: no process would
run inside the tool, so the need goes back to the requester as a
way of working rather than as a thing to buy
score-shortlist -> requirements [all-candidates-fail]: every
candidate fails a requirement, so the requirements are written
again with the decision and the scores behind it recorded
review-security -> score-shortlist [findings-to-answer]: the security
review raises findings the candidate has to answer, so the scoring
is worked again while the evaluation is still open
check-standards -> shortlist [cannot-hold-definitions]: the tool
cannot hold the definitions in force, so the candidate drops off
and the shortlist is cut again
trial-finalists -> shortlist [trial-fails]: the trial shows a
finalist cannot do the work, so the shortlist is cut again
whole-cost -> search-market [cost-rules-all-out]: the whole cost
rules every finalist out, so the market is searched again
procurement-review -> shortlist [terms-refused]: the organization
cannot accept the terms, so the shortlist is cut again and the next
candidate is taken
compare -> search-market [no-finalist-worth-taking]: no finalist is
worth taking, so the search reopens rather than the least bad one
being chosen
choose-tool -> trial-finalists [more-evidence-needed]: the choice
needs more evidence, so the finalists are trialled again on the
same work
bindings:
roster: <who holds each role - agents claiming the abstract agents
above, and named people for the requester, the systems owner,
the data owner, the security reviewer, procurement, legal,
finance and the revenue operations lead>
systems: the stack record (write), the requirements record (write),
the scoring record (write), the decision record (write),
the budget record (write), a trial environment (read),
the process register (read)
data: the requirements and the criteria at the version they were
fixed at, <your security and privacy standards> <version>,
<your data retention rules>, the definitions in force,
the stack record as it stands
policy:
no run goes past name-process without a process named that the tool
would run inside
the requirements and the criteria are written down and dated before
any candidate is looked at
every candidate is scored against the same list, and a requirement
that could not be tested is recorded as not tested
a change to the requirements after candidates have been seen goes in
as a new version, and every score made against the earlier version
is marked
a candidate with a relationship to the organization or to anyone
deciding is declared before scoring starts
no agreement is entered into by an agent. Named people sign it and
<your finance role> commits the money
the security review and the procurement review are human gates and
are never delegated to an agent
no entry goes into the stack record without a named owner, a renewal
date and the process it serves
measures:
cycle time: <target> from the need to the recorded entry
coverage: <share> of the requirements carrying a tested verdict on the
tool that was chosen
quality gate: nothing is chosen without a scored comparison behind it,
a cleared security review and a named person who decided
```
Take it somewhere
Use this process in CrewAI Flows
close
Paste this into an assistant that can read the web, such as Claude, ChatGPT or Cursor. It reads the specification and the current CrewAI documentation, then writes two files: the flow, and a note on what did not survive the translation. Read the note first. What a runtime cannot express is the part worth arguing about, and this process is a draft to argue with.
copy the prompt
398 lines · the document is inside it, so nothing else is needed
Convert the business process below into a runnable CrewAI Flow:
one Python file with a Pydantic state model, one Flow subclass, @start,
@listen, @router and a durable human feedback provider.
The document is a reference process written to the Agent Processes
specification. Read the specification before you start, because it defines
terms that look ordinary and are not:
https://agentcatalog.com/spec/agent-processes
Sections 6 (the phase graph), 6.5.1 (exception edges), 6.7.1 (deviations),
7 (automation) and 8 (handoffs) are the ones this conversion turns on.
Then read the current documentation for the primitives you will need, rather
than relying on what you remember of the API:
https://docs.crewai.com/en/concepts/flows
Flow, @start, @listen, @router, and_, or_, state, kickoff, plot
https://docs.crewai.com/en/learn/human-feedback-in-flows
@human_feedback, the provider protocol, from_pending and resume
https://docs.crewai.com/en/guides/flows/mastering-flow-state
@persist and what persistence actually promises, which is less than it sounds
WHAT THE DOCUMENT ASKS FOR
These hold wherever the process lands, and they matter more than style.
1. Each phase under `phases:` becomes one step, and keeps its name.
2. `after:` gives the edges. `after: a + b` is a join and waits for BOTH.
Reading it as "either" is the defect the specification calls out by name.
3. Every handoff carries a key in square brackets. Each key becomes one field
on the run's state, named exactly as the key with hyphens turned into
underscores, and the sentence beside it becomes that field's comment. The key
is the stable name; the sentence is prose that may be rewritten.
4. A phase MUST NOT begin before its inbound handoff exists. Where that is
checkable, check it in the step rather than assuming it.
5. `automation: never` is a gate a person signs. The run stops there and does
not continue until a person's decision comes back. Do not turn one into a
notification, a log line, or an automatic transition, whatever the queue
looks like.
6. Each line under `deviations:` is a backward or sideways edge, returning to
the phase named on the right. The key in brackets names it, and that name
belongs in the code.
7. A phase whose `after:` reads like "X or Y, whichever could not finish" is an
exception edge: it is entered when those phases FAIL, not when they succeed.
Do not wire it as an ordinary successor.
8. Anything in angle brackets is a blank the adopting organization fills in.
Leave each one as a named constant at the top of the file with a TODO. Do not
invent a value, a threshold or a date.
9. Record the document's `from:` line at the top of the file, so it says which
reference process and which version it was generated from.
10. Run-scoped lines under `run-scoped:` are work that runs alongside the whole
process rather than at one point in it, and a run may not close while one is
unfinished. Say in the code what you did about them, including if the answer
is that the runtime has nowhere to put them.
HOW THAT LOOKS IN CREWAI FLOWS
11. Build a Flow, not a Crew. A Crew is a team of roles with no graph, no join,
no persistence handle and no gate, and cannot express this document at all.
Each phase is one method on a single `Flow` subclass, keeping its name with
hyphens turned into underscores. A phase that genuinely needs role-based agents
builds its own Crew inside its own method body, and the Flow stays the graph.
12. Declare the state as a Pydantic model bound as the type parameter,
`class NegotiateTheAgreement(Flow[NegotiationState])`, and read and write
`self.state.field`. Never use the untyped dict form: the handoff keys are this
process's memory, and an untyped dict turns a misspelt key into a handoff that
is silently absent. Keep the auto-injected `id` field, which is what every
resume depends on.
13. `after: a` is `@listen(a)`. `after: a + b` is `@listen(and_(a, b))`. Never
write `or_` where the document writes `+`.
14. Check every inbound handoff at the top of the method and raise if one is
missing. `@listen` says when a method may run and says nothing about what is
in hand when it does.
15. The join is the trap, and it is proven rather than theoretical. `and_()`
empties its accumulator the moment it fires, so re-entering BOTH arms works
forever, and re-entering ONE arm after it has fired leaves it holding a single
trigger and waiting for the other for good. The cascade drains, CrewAI prints
that the flow completed, and `kickoff()` returns normally. So for any joining
phase that a deviation can send work back into, do not use `and_()` at all:
make the join a `@router` that both arms trigger, which reads the state fields
and emits its label only when every inbound handoff is present. A router is
re-evaluated against durable state every time and is never suppressed by the
once-fired set.
16. Do not write a phase as `@listen(or_(and_(a, b), "some_label"))`. There is
one accumulator per listener, shared across every branch of its condition and
wiped when any branch satisfies, so the label firing while the join is half
full erases the arm that had already arrived.
17. `automation: never` is `@human_feedback(message=..., provider=...)` stacked
under the method's `@listen`, with a provider whose `request_feedback` raises
`HumanFeedbackPending`. The run then persists, returns that object from
`kickoff()`, and a different process answers later with `from_pending(flow_id,
persistence)` and `resume(text)`. Do not take the default `ConsoleProvider`,
which calls `input()`: that gate exists only while somebody is watching a
terminal, and a run started by a scheduler either hangs or takes an empty
string.
18. Silence must not approve, and the platform's default is that it does. With
`emit=[...]` set, an empty resume collapses to `default_outcome`, or to the
first label when that is unset, with no model consulted and nobody named. Treat
an empty or unrecognised answer as a refusal in your own router. And write the
approver's name and the time onto the state yourself, because
`HumanFeedbackResult` carries the text, the outcome and a timestamp but has no
field for the person, which the specification requires.
19. Each line under `deviations:` is a `@router` named after the key, returning
a label named after the same key, with the target subscribing as
`@listen(or_(normal_trigger, "the_label"))`. Route rather than listen
directly, because a router is re-evaluated on every cycle while a top-level
`or_` listener is suppressed after it first fires.
20. An exception edge is a `try` and `except` around the failing phase's body,
recording the failure on the state and emitting an exception label from a
router. Wiring it as `@listen(or_(x, y))` fires when those phases SUCCEED, so
the clearing phase would run on every healthy run.
21. An unimplemented phase is a method with the document's own sentence as its
docstring and a bare `pass`, which is safe because listeners still fire on a
`None` return. An unimplemented `@router` is not safe: returning `None` emits
no label, every phase below it disappears from the run including the gates, and
the flow reports success. A stub router must return a hard-coded label with a
TODO beside it, or raise.
22. Turn persistence on with `@persist(SQLiteFlowPersistence(...))`, then treat
every phase that performs a real act as something that will run twice.
Persistence saves the state fields and nothing else, so a restart rehydrates
the data and runs the graph again from `@start`: a flow that had already sent a
written refusal sends a second one. Guard each acting phase with a state field
it checks and sets.
23. There are no timers, no deadlines and no cadences, and `run-scoped:` has no
counterpart at all. Put each `by:` value as a named constant, name the
run-scoped lines in the module docstring as unimplemented obligations, and say
in the fidelity note that no deadline in this document is enforced by anything.
These absences produce no diagnostic whatsoever, which is exactly why they have
to be written down.
Produce a second file alongside it, `FIDELITY.md`, and treat it as the more
important of the two. The code is for whoever builds this. The fidelity note
is for whoever has to decide whether this platform suits the process at all,
and that is usually a different person who will never read the code.
It has three parts.
**What came across.** Briefly: how many phases became steps, how many handoff
keys became state fields, which gates stop the run, which deviations became
edges. Counts and names, not reassurance.
**What did not, and what was done instead.** One entry per gap. For each one,
say what the document requires, what the platform can actually express, what
you did in its place, and what breaks if somebody later removes your
workaround. This last part matters most: a workaround nobody understands is a
workaround somebody deletes.
**What a person still has to decide.** The blanks are not a translation
failure, they are the point of a reference process, so list what has to be
filled in before this could run against anything real, and say which of those
choices the platform constrains.
Write it in plain English for somebody who has not read the specification, and
do not soften it. A translation of a reference process is a draft to argue
with, not a build artifact, and the honest account of what was lost is the most
useful thing you will produce.
Here is the process document.
```
PROCESS: choose a tool id: <team>/choose-a-tool v1
from: ref/rev/choose-a-tool v1
owner: <who> effective: <date>
trigger: a need the current tools cannot meet is raised in <your
planning process>
watch: record=<need> system=<your planning process>
change=<a need the current tools cannot meet is
raised>
concurrency: runs may overlap - <how many> searches open at once, and
two searches for the same need are merged into one
goal: one tool chosen against requirements fixed before any candidate
was seen, cleared by the security and the procurement reviews,
and recorded against the process it will run
phases:
take-in-need - human: read the need as the requester states it and
write down what the revenue organization cannot do
today
owner: stack-planner after: trigger
automation: <level>
name-process - human: the process that would run inside the tool,
named and pointed at. A need with no process behind
it goes back to the requester
owner: stack-planner after: take-in-need
by: <days> automation: <level>
check-stack - system: read the stack record for something already
licensed that does the job or comes near it
owner: stack-planner after: name-process
by: <days> automation: <level>
requirements - human: what the tool has to do, ranked, each one
written so a candidate can be tested against it
owner: stack-planner after: check-stack
by: <days> automation: <level>
set-criteria - human: how a candidate is judged and what a pass
takes, fixed at a version and dated
owner: supplier-check after: requirements
by: <days> automation: <level>
brief-evaluation - convenes briefing: every agent hears the same
requirements, the criteria and the due date at once
owner: stack-planner after: set-criteria
automation: <level>
map-connections - human: the systems that would feed the tool, the
systems that would read it, and what would flow
owner: integration-keeper after: brief-evaluation
by: <days> automation: <level>
search-market - runs collect-and-report: one set of questions in one
format to every candidate, every answer kept against
the candidate that gave it
owner: researcher after: brief-evaluation
by: <weeks> automation: <level>
shortlist - human: the candidates that miss a requirement are
dropped, each with the requirement it failed
owner: stack-planner
after: map-connections + search-market
by: <days> automation: <level>
score-shortlist - runs assessment: every requirement gets a verdict
for every candidate, against the list at the version
it was fixed at
owner: supplier-check after: shortlist
by: <days> automation: <level>
review-security - human: how the candidate holds the data, what it
would be given, and who at the supplier can read it
owner: <your security role> after: score-shortlist
by: <days> automation: never
check-standards - runs assessment: whether the tool can hold the
definitions, the records and the retention rules in
force
owner: standards-keeper after: score-shortlist
by: <days> automation: <level>
trial-finalists - runs assessment: each finalist takes the same work
the team actually has, over the same period,
measured the same way
owner: stack-planner after: score-shortlist
by: <weeks> automation: <level>
whole-cost - human: the licence, the setup, the people it takes
to run, and what leaving would cost
owner: <your finance role> after: trial-finalists
by: <days> automation: <level>
procurement-review - convenes approval: the price, the term, the notice
a cancellation needs, and what the supplier has to
meet. Procurement and legal sign as named signers
owner: <your procurement role>
after: whole-cost
by: <weeks> automation: never
compare - human: the finalists side by side against one list,
with what taking each one would rule out
owner: stack-planner
after: review-security + check-standards +
procurement-review
by: <days> automation: <level>
choose-tool - convenes decide-and-announce: <who> picks one, says
why, and every agent waiting on it hears at the same
time
owner: stack-planner after: compare
by: <days> automation: never
record-entry - human: the entry lands with the process it serves,
its owner, its cost, its term and its renewal date
owner: stack-planner after: choose-tool
automation: <level>
run-scoped:
status - runs roll-call owner: stack-planner
every: <cadence>
from: brief-evaluation until: choose-tool
money - runs allocate-and-reconcile owner: <your finance role>
from: whole-cost until: run close
handoffs:
take-in-need -> name-process [the-need]: what the revenue
organization cannot do today, in one line the requirements are
written from
name-process -> check-stack [process-served]: the process that would
run inside the tool, and the activities in it the tool would carry
check-stack -> requirements [near-misses-in-stack]: what the stack
already holds that comes near, and what each of those falls short
on
requirements -> set-criteria [ranked-requirements]: the requirements
ranked, each written so a candidate can be tested against it
set-criteria -> brief-evaluation [fixed-criteria]: the requirements
and the criteria at a version, dated before any candidate was
looked at
brief-evaluation -> map-connections [process-systems]: the process
the tool serves and the systems that process already reads and
writes
brief-evaluation -> search-market [evaluation-brief]: the
requirements as every agent heard them, and the date the choice is
due
map-connections -> shortlist [required-connections]: what would have
to flow in and out, and the connections a candidate would have to
support
search-market -> shortlist [candidates-found]: every candidate found,
with where it came from and what it answered, and a candidate that
did not answer recorded as missing
shortlist -> score-shortlist [shortlisted-candidates]: the candidates
left, and the requirement each dropped candidate failed
score-shortlist -> trial-finalists [requirement-verdicts]: a verdict
per requirement per candidate with the evidence behind it, and
every requirement that could not be tested marked as not tested
review-security -> compare [security-findings]: how the candidate
holds the data, what it would be given, and every finding it has to
answer
check-standards -> compare [standards-verdict]: whether the tool can
hold the definitions and the retention rules, with the rule behind
every failure
trial-finalists -> whole-cost [trial-results]: what each finalist did
on the same work, with the figures and the range around each one
whole-cost -> procurement-review [cost-of-ownership]: the licence,
the setup, the people it takes to run, and what leaving would cost
procurement-review -> compare [granted-terms]: the terms as the
supplier will grant them, and every term the organization will not
accept
compare -> choose-tool [side-by-side]: the finalists set side by side
against one list, and what taking each one would rule out
choose-tool -> record-entry [the-choice]: the chosen tool, who
decided, on what date, for what reasons, and anyone who disagreed
deviations:
name-process -> take-in-need [no-process-behind-it]: no process would
run inside the tool, so the need goes back to the requester as a
way of working rather than as a thing to buy
score-shortlist -> requirements [all-candidates-fail]: every
candidate fails a requirement, so the requirements are written
again with the decision and the scores behind it recorded
review-security -> score-shortlist [findings-to-answer]: the security
review raises findings the candidate has to answer, so the scoring
is worked again while the evaluation is still open
check-standards -> shortlist [cannot-hold-definitions]: the tool
cannot hold the definitions in force, so the candidate drops off
and the shortlist is cut again
trial-finalists -> shortlist [trial-fails]: the trial shows a
finalist cannot do the work, so the shortlist is cut again
whole-cost -> search-market [cost-rules-all-out]: the whole cost
rules every finalist out, so the market is searched again
procurement-review -> shortlist [terms-refused]: the organization
cannot accept the terms, so the shortlist is cut again and the next
candidate is taken
compare -> search-market [no-finalist-worth-taking]: no finalist is
worth taking, so the search reopens rather than the least bad one
being chosen
choose-tool -> trial-finalists [more-evidence-needed]: the choice
needs more evidence, so the finalists are trialled again on the
same work
bindings:
roster: <who holds each role - agents claiming the abstract agents
above, and named people for the requester, the systems owner,
the data owner, the security reviewer, procurement, legal,
finance and the revenue operations lead>
systems: the stack record (write), the requirements record (write),
the scoring record (write), the decision record (write),
the budget record (write), a trial environment (read),
the process register (read)
data: the requirements and the criteria at the version they were
fixed at, <your security and privacy standards> <version>,
<your data retention rules>, the definitions in force,
the stack record as it stands
policy:
no run goes past name-process without a process named that the tool
would run inside
the requirements and the criteria are written down and dated before
any candidate is looked at
every candidate is scored against the same list, and a requirement
that could not be tested is recorded as not tested
a change to the requirements after candidates have been seen goes in
as a new version, and every score made against the earlier version
is marked
a candidate with a relationship to the organization or to anyone
deciding is declared before scoring starts
no agreement is entered into by an agent. Named people sign it and
<your finance role> commits the money
the security review and the procurement review are human gates and
are never delegated to an agent
no entry goes into the stack record without a named owner, a renewal
date and the process it serves
measures:
cycle time: <target> from the need to the recorded entry
coverage: <share> of the requirements carrying a tested verdict on the
tool that was chosen
quality gate: nothing is chosen without a scored comparison behind it,
a cleared security review and a named person who decided
```
Take it somewhere
Use this process in Google ADK
close
Paste this into an assistant that can read the web, such as Claude, ChatGPT or Cursor. It reads the specification and the current ADK documentation, then writes two files: the workflow, and a note on what did not survive the translation. Read the note first. What a runtime cannot express is the part worth arguing about, and this process is a draft to argue with.
copy the prompt
389 lines · the document is inside it, so nothing else is needed
Convert the business process below into a runnable Google ADK workflow:
one Python file with a Workflow, nodes, routed edges, a JoinNode,
RequestInput gates and a persisting session service.
The document is a reference process written to the Agent Processes
specification. Read the specification before you start, because it defines
terms that look ordinary and are not:
https://agentcatalog.com/spec/agent-processes
Sections 6 (the phase graph), 6.5.1 (exception edges), 6.7.1 (deviations),
7 (automation) and 8 (handoffs) are the ones this conversion turns on.
Then read the current documentation for the primitives you will need, rather
than relying on what you remember of the API:
https://adk.dev/graphs/routes/
nodes, tuple chains, Event(route=), JoinNode, back-edges
https://adk.dev/graphs/human-input/
RequestInput and the rerun_on_resume handoff
https://adk.dev/runtime/resume/
ResumabilityConfig, resuming by invocation id, at-least-once tools
https://adk.dev/graphs/data-handling/
Event.output against state, and the selector syntax in instructions
WHAT THE DOCUMENT ASKS FOR
These hold wherever the process lands, and they matter more than style.
1. Each phase under `phases:` becomes one step, and keeps its name.
2. `after:` gives the edges. `after: a + b` is a join and waits for BOTH.
Reading it as "either" is the defect the specification calls out by name.
3. Every handoff carries a key in square brackets. Each key becomes one field
on the run's state, named exactly as the key with hyphens turned into
underscores, and the sentence beside it becomes that field's comment. The key
is the stable name; the sentence is prose that may be rewritten.
4. A phase MUST NOT begin before its inbound handoff exists. Where that is
checkable, check it in the step rather than assuming it.
5. `automation: never` is a gate a person signs. The run stops there and does
not continue until a person's decision comes back. Do not turn one into a
notification, a log line, or an automatic transition, whatever the queue
looks like.
6. Each line under `deviations:` is a backward or sideways edge, returning to
the phase named on the right. The key in brackets names it, and that name
belongs in the code.
7. A phase whose `after:` reads like "X or Y, whichever could not finish" is an
exception edge: it is entered when those phases FAIL, not when they succeed.
Do not wire it as an ordinary successor.
8. Anything in angle brackets is a blank the adopting organization fills in.
Leave each one as a named constant at the top of the file with a TODO. Do not
invent a value, a threshold or a date.
9. Record the document's `from:` line at the top of the file, so it says which
reference process and which version it was generated from.
10. Run-scoped lines under `run-scoped:` are work that runs alongside the whole
process rather than at one point in it, and a run may not close while one is
unfinished. Say in the code what you did about them, including if the answer
is that the runtime has nowhere to put them.
HOW THAT LOOKS IN GOOGLE ADK
11. Build a `Workflow` from `google.adk.workflow`, and pin `google-adk>=2.0` in
a comment. Do not use `SequentialAgent`, `ParallelAgent` or `LoopAgent`: they
are deprecated in favour of the graph, and they carry their own defects around
state and control flow. The documentation moved to adk.dev, and anything you
remember about nesting agents rather than drawing a graph is out of date.
12. Each phase is one node keeping its name. Take the node kind from how the
phase resolves rather than from taste: a `human:` or `system:` phase is a plain
Python function node, and a `runs` or `convenes` phase is an `Agent`.
13. `after:` gives the edges, written as tuple chains in `edges=[...]`, and the
trigger is the `"START"` keyword. Take the order only from `after:` lines and
never from the order the phases are listed in.
14. `after: a + b` is a `JoinNode`, and it must be guarded, because this is the
worst trap of any runtime here. The join fires when every static predecessor is
marked COMPLETED, nothing ever un-completes a node, and stored outputs are
never cleared. So after a deviation re-runs one arm, the join fires the instant
that arm finishes and hands the next phase LAST PASS'S value for every arm that
did not re-run. It does not stall, it proceeds with stale data, and nothing
logs. Stamp each arm's output with a pass counter or a content hash, and have
the phase after the join compare the stamps and refuse to run when they
disagree.
15. Each line under `deviations:` is a routed back-edge: a router after the
phase on the left returning `Event(route=...)`, named after the key in
brackets, with one arm going back to the phase on the right and one going
forward. An unconditional cycle raises at construction, which is the one place
this model checks your work. Nothing budgets a routed cycle, so add your own
count and stop rather than looping forever.
16. Give every router an explicit `DEFAULT_ROUTE` arm, and route it to a phase
that stops and asks a person. A route value matching no key writes a log
warning, ends that branch, and lets the run finish reporting success with the
rest of the process never having happened.
17. `automation: never` is a `RequestInput` node of its own, never an `Agent`
asking a question. Decorate it `@node(rerun_on_resume=False)` and yield
`RequestInput(message=..., payload=..., response_schema=...)`, so the run
stops, persists, and delivers the person's answer to the node's successor as
its typed input. A resumed workflow runs its tools at least once, so any
irreversible act needs its own duplicate guard.
18. Make the gates durable or say plainly that they are not. Wrap the graph in
`App(..., resumability_config=ResumabilityConfig(is_resumable=True))` and pass
a persisting session service, never the in-memory one. Note in the file that
the command line and the web UI cannot resume a run, so whoever releases these
gates needs an operator surface that somebody has to write.
19. Every `Agent` in the graph gets `mode="single_turn"` and no `sub_agents`.
A non-empty `sub_agents` list silently adds a transfer tool, and a model that
uses it runs a different agent in this node's place while the graph's outgoing
edge fires on schedule regardless: the topology is honoured perfectly and the
work belongs to somebody else.
20. Model failure as a route, not as an exception. A node that raises does not
propagate: the failure is caught, recorded, and shuts the workflow down without
raising to the caller. So a phase that can fail catches its own failure and
returns `Event(route="could-not-finish")`, and the exception phase hangs off
that arm.
21. Keep every blank as a named module-level constant and never interpolate one
into an `instruction=` string. Angle brackets and curly braces are ADK's own
data selector syntax inside instructions, so a blank pasted verbatim stops
being a blank and becomes a selector.
22. `by:` and `not-before:` have no expression, and `@node(timeout=)` is not
one: it is an in-process wall clock that cancels the node and, because failures
are swallowed, ends the run silently rather than recording a missed deadline.
Nothing in `run-scoped:` has an expression either, and it must not be faked as
an ordinary node, because a node has to be reached and has to finish before
anything downstream starts, which is the opposite of what those lines mean.
Leave both out of the graph and name them in the fidelity note.
Produce a second file alongside it, `FIDELITY.md`, and treat it as the more
important of the two. The code is for whoever builds this. The fidelity note
is for whoever has to decide whether this platform suits the process at all,
and that is usually a different person who will never read the code.
It has three parts.
**What came across.** Briefly: how many phases became steps, how many handoff
keys became state fields, which gates stop the run, which deviations became
edges. Counts and names, not reassurance.
**What did not, and what was done instead.** One entry per gap. For each one,
say what the document requires, what the platform can actually express, what
you did in its place, and what breaks if somebody later removes your
workaround. This last part matters most: a workaround nobody understands is a
workaround somebody deletes.
**What a person still has to decide.** The blanks are not a translation
failure, they are the point of a reference process, so list what has to be
filled in before this could run against anything real, and say which of those
choices the platform constrains.
Write it in plain English for somebody who has not read the specification, and
do not soften it. A translation of a reference process is a draft to argue
with, not a build artifact, and the honest account of what was lost is the most
useful thing you will produce.
Here is the process document.
```
PROCESS: choose a tool id: <team>/choose-a-tool v1
from: ref/rev/choose-a-tool v1
owner: <who> effective: <date>
trigger: a need the current tools cannot meet is raised in <your
planning process>
watch: record=<need> system=<your planning process>
change=<a need the current tools cannot meet is
raised>
concurrency: runs may overlap - <how many> searches open at once, and
two searches for the same need are merged into one
goal: one tool chosen against requirements fixed before any candidate
was seen, cleared by the security and the procurement reviews,
and recorded against the process it will run
phases:
take-in-need - human: read the need as the requester states it and
write down what the revenue organization cannot do
today
owner: stack-planner after: trigger
automation: <level>
name-process - human: the process that would run inside the tool,
named and pointed at. A need with no process behind
it goes back to the requester
owner: stack-planner after: take-in-need
by: <days> automation: <level>
check-stack - system: read the stack record for something already
licensed that does the job or comes near it
owner: stack-planner after: name-process
by: <days> automation: <level>
requirements - human: what the tool has to do, ranked, each one
written so a candidate can be tested against it
owner: stack-planner after: check-stack
by: <days> automation: <level>
set-criteria - human: how a candidate is judged and what a pass
takes, fixed at a version and dated
owner: supplier-check after: requirements
by: <days> automation: <level>
brief-evaluation - convenes briefing: every agent hears the same
requirements, the criteria and the due date at once
owner: stack-planner after: set-criteria
automation: <level>
map-connections - human: the systems that would feed the tool, the
systems that would read it, and what would flow
owner: integration-keeper after: brief-evaluation
by: <days> automation: <level>
search-market - runs collect-and-report: one set of questions in one
format to every candidate, every answer kept against
the candidate that gave it
owner: researcher after: brief-evaluation
by: <weeks> automation: <level>
shortlist - human: the candidates that miss a requirement are
dropped, each with the requirement it failed
owner: stack-planner
after: map-connections + search-market
by: <days> automation: <level>
score-shortlist - runs assessment: every requirement gets a verdict
for every candidate, against the list at the version
it was fixed at
owner: supplier-check after: shortlist
by: <days> automation: <level>
review-security - human: how the candidate holds the data, what it
would be given, and who at the supplier can read it
owner: <your security role> after: score-shortlist
by: <days> automation: never
check-standards - runs assessment: whether the tool can hold the
definitions, the records and the retention rules in
force
owner: standards-keeper after: score-shortlist
by: <days> automation: <level>
trial-finalists - runs assessment: each finalist takes the same work
the team actually has, over the same period,
measured the same way
owner: stack-planner after: score-shortlist
by: <weeks> automation: <level>
whole-cost - human: the licence, the setup, the people it takes
to run, and what leaving would cost
owner: <your finance role> after: trial-finalists
by: <days> automation: <level>
procurement-review - convenes approval: the price, the term, the notice
a cancellation needs, and what the supplier has to
meet. Procurement and legal sign as named signers
owner: <your procurement role>
after: whole-cost
by: <weeks> automation: never
compare - human: the finalists side by side against one list,
with what taking each one would rule out
owner: stack-planner
after: review-security + check-standards +
procurement-review
by: <days> automation: <level>
choose-tool - convenes decide-and-announce: <who> picks one, says
why, and every agent waiting on it hears at the same
time
owner: stack-planner after: compare
by: <days> automation: never
record-entry - human: the entry lands with the process it serves,
its owner, its cost, its term and its renewal date
owner: stack-planner after: choose-tool
automation: <level>
run-scoped:
status - runs roll-call owner: stack-planner
every: <cadence>
from: brief-evaluation until: choose-tool
money - runs allocate-and-reconcile owner: <your finance role>
from: whole-cost until: run close
handoffs:
take-in-need -> name-process [the-need]: what the revenue
organization cannot do today, in one line the requirements are
written from
name-process -> check-stack [process-served]: the process that would
run inside the tool, and the activities in it the tool would carry
check-stack -> requirements [near-misses-in-stack]: what the stack
already holds that comes near, and what each of those falls short
on
requirements -> set-criteria [ranked-requirements]: the requirements
ranked, each written so a candidate can be tested against it
set-criteria -> brief-evaluation [fixed-criteria]: the requirements
and the criteria at a version, dated before any candidate was
looked at
brief-evaluation -> map-connections [process-systems]: the process
the tool serves and the systems that process already reads and
writes
brief-evaluation -> search-market [evaluation-brief]: the
requirements as every agent heard them, and the date the choice is
due
map-connections -> shortlist [required-connections]: what would have
to flow in and out, and the connections a candidate would have to
support
search-market -> shortlist [candidates-found]: every candidate found,
with where it came from and what it answered, and a candidate that
did not answer recorded as missing
shortlist -> score-shortlist [shortlisted-candidates]: the candidates
left, and the requirement each dropped candidate failed
score-shortlist -> trial-finalists [requirement-verdicts]: a verdict
per requirement per candidate with the evidence behind it, and
every requirement that could not be tested marked as not tested
review-security -> compare [security-findings]: how the candidate
holds the data, what it would be given, and every finding it has to
answer
check-standards -> compare [standards-verdict]: whether the tool can
hold the definitions and the retention rules, with the rule behind
every failure
trial-finalists -> whole-cost [trial-results]: what each finalist did
on the same work, with the figures and the range around each one
whole-cost -> procurement-review [cost-of-ownership]: the licence,
the setup, the people it takes to run, and what leaving would cost
procurement-review -> compare [granted-terms]: the terms as the
supplier will grant them, and every term the organization will not
accept
compare -> choose-tool [side-by-side]: the finalists set side by side
against one list, and what taking each one would rule out
choose-tool -> record-entry [the-choice]: the chosen tool, who
decided, on what date, for what reasons, and anyone who disagreed
deviations:
name-process -> take-in-need [no-process-behind-it]: no process would
run inside the tool, so the need goes back to the requester as a
way of working rather than as a thing to buy
score-shortlist -> requirements [all-candidates-fail]: every
candidate fails a requirement, so the requirements are written
again with the decision and the scores behind it recorded
review-security -> score-shortlist [findings-to-answer]: the security
review raises findings the candidate has to answer, so the scoring
is worked again while the evaluation is still open
check-standards -> shortlist [cannot-hold-definitions]: the tool
cannot hold the definitions in force, so the candidate drops off
and the shortlist is cut again
trial-finalists -> shortlist [trial-fails]: the trial shows a
finalist cannot do the work, so the shortlist is cut again
whole-cost -> search-market [cost-rules-all-out]: the whole cost
rules every finalist out, so the market is searched again
procurement-review -> shortlist [terms-refused]: the organization
cannot accept the terms, so the shortlist is cut again and the next
candidate is taken
compare -> search-market [no-finalist-worth-taking]: no finalist is
worth taking, so the search reopens rather than the least bad one
being chosen
choose-tool -> trial-finalists [more-evidence-needed]: the choice
needs more evidence, so the finalists are trialled again on the
same work
bindings:
roster: <who holds each role - agents claiming the abstract agents
above, and named people for the requester, the systems owner,
the data owner, the security reviewer, procurement, legal,
finance and the revenue operations lead>
systems: the stack record (write), the requirements record (write),
the scoring record (write), the decision record (write),
the budget record (write), a trial environment (read),
the process register (read)
data: the requirements and the criteria at the version they were
fixed at, <your security and privacy standards> <version>,
<your data retention rules>, the definitions in force,
the stack record as it stands
policy:
no run goes past name-process without a process named that the tool
would run inside
the requirements and the criteria are written down and dated before
any candidate is looked at
every candidate is scored against the same list, and a requirement
that could not be tested is recorded as not tested
a change to the requirements after candidates have been seen goes in
as a new version, and every score made against the earlier version
is marked
a candidate with a relationship to the organization or to anyone
deciding is declared before scoring starts
no agreement is entered into by an agent. Named people sign it and
<your finance role> commits the money
the security review and the procurement review are human gates and
are never delegated to an agent
no entry goes into the stack record without a named owner, a renewal
date and the process it serves
measures:
cycle time: <target> from the need to the recorded entry
coverage: <share> of the requirements carrying a tested verdict on the
tool that was chosen
quality gate: nothing is chosen without a scored comparison behind it,
a cleared security review and a named person who decided
```
One run
A simulation of one run
The activities are on the left, whoever is doing the active one is on the right, and the record of the run builds up as it goes.
This run is built from the same rows as the diagram above: the left column is the activity list, the captions are the activity lines, the cast is the roster, and the labels on the wires are what the handoffs say actually passes.
Adoption
What you fill in
48 blanks to fill. Everything else is the process.
this process
from: ref/rev/choose-a-tool v1
Copy this line into your own document. It never claims this process is running anywhere; it records which draft yours started from, and it is what lets the catalog tell you when this one changes.
The header. Your own id, an owner, and the date it takes effect. One line records where it came from, and that line is what lets the catalog tell you when this reference process changes.
The roster. Which agent takes each activity, and which person takes each of the human ones. The process already names what it needs, so this is a lookup rather than a design exercise.
The numbers. Dates, budgets, cadences, and the targets in the measures block. Nothing here can be a reference value, because a target nobody chose is a target nobody meets.