Route Coding Tasks to the Right Agent with Jev

Not every coding task needs the most expensive model. In this guide a Buddy pipeline asks Jev which agent should handle a task, runs only that agent, and pushes its change to a new branch after the tests pass.

The demo app is a small Flask booking service for a bike repair shop. We sent it three tasks:

Task Jev answer Confidence What ran
Fix typo in booking confirmation email subject haiku 0.98 Claude Haiku, branch agent-haiku-8
Clean up the booking code opus 0.29 A human picked sonnet, branch agent-sonnet-9
Add partial refunds for cancelled bookings risky 0.99 Approval, then Claude Opus, branch agent-risky-10

What you need

The pipeline uses four integrations:

  • TypeSafe for the Jev action, which classifies the task,
  • Anthropic for Claude Code with Haiku, Sonnet and Opus,
  • Cursor for changes limited to templates and CSS,
  • OpenAI for Codex, which answers questions about the code.

Add them in Integrations before you create the pipeline. The YAML below refers to them by their IDs: typesafe, anthropic, cursor and openai.

How the pipeline decides

Jev returns one key from the list below. Each key has a matching agent:

Key When Agent
question The task needs no file changes Codex in read-only mode
ui Only templates/ or static/ change Cursor
haiku A trivial change of a few lines Claude Haiku
sonnet A backend change with a clear scope Claude Sonnet
opus A vague or cross-cutting change Claude Opus
risky Payments, refunds, auth or migrations Approval, then Claude Opus

Two rules sit on top of the answer:

  • confidence below 0.6: a human picks the route,
  • risky, or a risky probability above 0.3: a human approves the run before any agent starts.

Jev picks who works. It never grants permissions and it can only add a gate, never remove one.

Image loading...Workflow tab of the route-coding-task pipeline: Classify task with Jev, Set route, two gates, five agent actions, Check agent changes, Run tests and Push agent branch, every action after Set route carries an IF badge

Building the pipeline

1. Classify the task

The task comes from the TASK variable. Jev also reads README.md, so it knows which folders hold payments and which hold templates.

yaml
- pipeline: route-coding-task refs: - refs/heads/main variables: - key: ROUTE value: none settable: ENABLED - key: TASK value: Fix typo in booking confirmation email subject actions: - action: Classify task with Jev type: JEV integration: typesafe question: Which agent and model should handle this coding task in a Flask bike service booking app? question_type: CHOICE state: - "Task: $TASK" - file: README.md criteria: question: "The task asks a question about the code and needs no file changes (explain, find, review)." ui: "The change is limited to templates/ or static/ (HTML, CSS, front-end JavaScript)." haiku: "A trivial change of one or a few lines, such as a typo, copy text, a rename or a constant." sonnet: A regular backend change with a clear scope in one area of the app. opus: A vague or cross-cutting change that needs design work across several modules. risky: "The change touches payments, deposits, refunds, authentication or database migrations."

2. Save the route

Jev's answer goes to ROUTE, a settable variable that a human can still change later. The action also saves checksums of the source files, so we can check after the agent whether anything changed. Add .buddy/ to .gitignore, so the checksum file does not end up on the agent branch.

yaml
- action: Set route type: BUILD docker_image_name: library/alpine docker_image_tag: latest commands: |- echo "Jev result: $BUDDY_ACTION_JEV_RESULT" echo "Confidence: $BUDDY_ACTION_JEV_CONFIDENCE" echo "Risky probability: $BUDDY_ACTION_JEV_PROBABILITY_RISKY" export ROUTE="$BUDDY_ACTION_JEV_RESULT" mkdir -p .buddy find app templates static tests migrations -type f -not -path "*/__pycache__/*" -exec md5sum {} + | sort > .buddy/before.md5

3. Ask a human when Jev is unsure

When confidence is below 0.6, the run stops and shows a select with the routes. The choice overwrites ROUTE. With cancel the remaining actions are skipped and the run finishes without a branch.

yaml
- action: Ask human for route type: WAIT_FOR_VARIABLES trigger_conditions: - trigger_condition: VAR_LESS_THAN trigger_variable_key: BUDDY_ACTION_JEV_CONFIDENCE trigger_variable_value: "0.6" variables: - key: ROUTE defaults: |- sonnet opus haiku ui question cancel description: "Jev picked $BUDDY_ACTION_JEV_RESULT with confidence $BUDDY_ACTION_JEV_CONFIDENCE. Pick the route for: $TASK"

"Clean up the booking code" got opus with confidence 0.29, so the run waited for a decision. We picked sonnet.

Image loading...Run #9 in Buddy: Jev picks opus with confidence 0.29, Ask human for route waits with a ROUTE select set to sonnet

Image loading...ROUTE select with the options sonnet, opus, haiku, ui, question and cancel

4. Approve risky changes

yaml
- action: Approve risky change type: WAIT_FOR_APPLY trigger_conditions: - trigger_condition: OR trigger_operands: - trigger_condition: VAR_IS trigger_variable_key: ROUTE trigger_variable_value: risky - trigger_condition: VAR_GREATER_THAN trigger_variable_key: BUDDY_ACTION_JEV_PROBABILITY_RISKY trigger_variable_value: "0.3" description: "Jev flagged this task as risky (p=$BUDDY_ACTION_JEV_PROBABILITY_RISKY). Approve before an agent touches the code: $TASK"

The second condition catches tasks where risky was not the top answer but still had a real share. The gate runs after the human choice, so a risky probability above 0.3 asks for approval even when someone picked another route. The refunds task got risky with confidence 0.99 and waited for approval.

Image loading...Run #10 in Buddy: Jev picks risky at 0.99, Approve risky change waits with Approve and Stop buttons before Claude Opus starts

5. One action per agent

Each agent action runs only for its route. Claude Code gets --permission-mode acceptEdits, so it can edit files. Shell commands are blocked, because nobody approves them in a pipeline.

yaml
- action: Quick fix with Claude Haiku type: CLAUDE_CODE integration: anthropic trigger_conditions: - trigger_condition: VAR_IS trigger_variable_key: ROUTE trigger_variable_value: haiku prompts: - "Task: $TASK. Find the code and edit the file yourself with the smallest change that solves it. Do not delegate to subagents, do not run tests and do not commit." model: haiku effort: low claude_args: "--permission-mode acceptEdits"

The other routes follow the same pattern:

Action Condition Settings
Answer with Codex (read-only) ROUTE is question type: CODEX_CLI, integration: openai, sandbox_mode: READ_ONLY
UI change with Cursor ROUTE is ui type: CURSOR, integration: cursor, model: auto-smart
Backend change with Claude Sonnet ROUTE is sonnet model: sonnet, effort: medium
Complex change with Claude Opus ROUTE is opus or risky model: opus

Codex on the question route answers in the action log. The run skips the tests and the push, because nothing in the code changes.

Warning

Without acceptEdits the agent changes nothing.

In our first run Haiku found the typo, asked for permission to write the file and ended with a green status. The pipeline went on as if the fix was done. The check in the next step catches this case.

6. Check that the agent changed something

yaml
- action: Check agent changes type: BUILD docker_image_name: library/alpine docker_image_tag: latest trigger_conditions: - trigger_condition: VAR_IS_NOT trigger_variable_key: ROUTE trigger_variable_value: question - trigger_condition: VAR_IS_NOT trigger_variable_key: ROUTE trigger_variable_value: cancel commands: |- find app templates static tests migrations -type f -not -path "*/__pycache__/*" -exec md5sum {} + | sort > /tmp/after.md5 if cmp -s .buddy/before.md5 /tmp/after.md5; then echo "The agent did not change any files" exit 1 fi comm -13 .buddy/before.md5 /tmp/after.md5 | awk '{print "Changed: " $2}'

The log lists new and modified files. comm -13 skips files the agent deleted, while cmp still catches them. Haiku changed app/emails.py, Sonnet changed app/bookings/service.py, and Opus changed five files for the refunds, including a new tests/test_refunds.py.

7. Run the tests and push a branch

yaml
- action: Run tests type: BUILD docker_image_name: library/python docker_image_tag: "3.12" trigger_conditions: - trigger_condition: VAR_IS_NOT trigger_variable_key: ROUTE trigger_variable_value: question - trigger_condition: VAR_IS_NOT trigger_variable_key: ROUTE trigger_variable_value: cancel commands: |- pip install -q --root-user-action=ignore -r requirements.txt pytest -q - action: Push agent branch type: PUSH trigger_conditions: - trigger_condition: VAR_IS_NOT trigger_variable_key: ROUTE trigger_variable_value: question - trigger_condition: VAR_IS_NOT trigger_variable_key: ROUTE trigger_variable_value: cancel targets: - project_repository target_branch: agent-$ROUTE-$BUDDY_RUN_ID comment: "$ROUTE: $TASK"

If the tests fail, the run stops before the push and no branch is created. The agent never pushes to main. A human reviews the branch and merges it.

Image loading...Run #9 in Buddy: Sonnet, Check agent changes, Run tests and Push agent branch are green, the other agents are skipped

Warning

The pipeline needs write access to the repository.

By default pipelines can only read the project repository, so the push fails. Open Targets, select Project repository, go to the Permissions tab and set the pipeline to Read-write. You can leave All pipelines as Read-only and grant write access only to route-coding-task.

Image loading...Permissions tab of the Project repository target: All pipelines set to Read-only, bike-service / route-coding-task set to Read-write

The refunds run shows why a human merge stays the last gate. Opus added the refund tiers, but it made up the policy on its own (100% at 48 hours or more before the booking, 50% between 24 and 48 hours, nothing later) and asked for a confirmation in its summary. It also noted that the cancel endpoint has no authentication.

Image loading...Run #10 in Buddy: Claude Opus summary with the proposed refund tiers and open questions, followed by the push of branch agent-risky-10

Things to watch

  • Confidence is not accuracy. A confident answer can still be wrong. The test run and the human merge are the real gates.
  • Write criteria as conditions. "Touches payments, refunds or migrations" works better than "risky". Overlapping criteria split the probability and lower the confidence.
  • Tune the thresholds. We set 0.6 and 0.3 by hand after a few runs. If people often change the route Jev picked, raise the confidence threshold.
  • Give Jev the map of the project. Without README.md in the state, Jev only sees the task title and cannot tell where payments live.
Jarek Dylewski

Jarek Dylewski

Customer Support

A journalist and an SEO specialist trying to find himself in the unforgiving world of coders. Gamer, a non-fiction literature fan and obsessive carnivore. Jarek uses his talents to convert the programming lingo into a cohesive and approachable narration.

Oct 7, 2026
Share