Flaky Test or Bug? Let the Pipeline Decide
A test failed. Jev decides if it is flaky, a real bug or a case for a human. Based on the answer the pipeline re-runs the tests, notifies on-call or starts Claude to find the cause. The routing uses ordinary trigger_conditions, so nothing has to be parsed from the model's response.
What Jev does
Jev is a model from TypeSafe. You ask it one question and give it context and a list of possible answers. It does not browse the repository and does not run commands. It picks an answer and says how sure it is.
With the CHOICE type the result lands in two variables:
BUDDY_ACTION_JEV_RESULT- the key of the selected option,BUDDY_ACTION_JEV_CONFIDENCE- confidence from 0 to 1. TypeSafe computes it from the whole probability distribution, so it is not the same as the probability of the selected option.
The other types and variables are described in the documentation.
Pipeline
The application is a small Vite store. It has tests for the cart, for the discount at 100 USD and for fetching exchange rates. For this guide every failure is caused on purpose: with FLAKY_FAIL=1 or RATES_URL in the test command, with one changed line in src/cart.js or with a separate branch that has no test database.
After a failed test Jev classifies the failure and one branch runs:
flakywith confidence of at least0.7: re-run the tests,infrastructurewith confidence of at least0.7: notify on-call on Slack,real_bugwith confidence of at least0.7: Claude Code diagnosis,- confidence below
0.7: human decision.
Build and deploy run after green tests, after a successful re-run or after a human approval. In every other case, including other, the run ends with the Failed status.
Image loading...
Tests save the log and the result
yaml- action: Run tests type: BUILD docker_image_name: node docker_image_tag: "22" commands: |- rm -f test-output.log test-results.xml triage.md npm ci --no-audit --no-fund if npm test > test-output.log 2>&1; then export OUTPUT_TESTS=passed; else export OUTPUT_TESTS=failed; fi
The action stays green and the test result goes to OUTPUT_TESTS. rm -f removes files from the previous run, because the pipeline filesystem is kept between runs. This matters most for triage.md, which Claude writes in a later action. Without the cleanup an old diagnosis would stay in a run where Claude did not start at all.
Jev answers why the tests failed
yaml- action: Classify failure type: JEV integration: typesafe trigger_conditions: - trigger_condition: VAR_IS trigger_variable_key: OUTPUT_TESTS trigger_variable_value: failed question: Why did the test run fail? question_type: CHOICE state: - "Commit: $BUDDY_RUN_COMMIT_MESSAGE" - file: .buddy/triage-policy.md - file: test-output.log criteria: flaky: Tests failed on a timeout, connection reset, port already in use, or another condition that a re-run would most likely fix; the assertions themselves did not fail infrastructure: A required service, database, or package registry was unreachable during test startup; no test assertions were reached real_bug: At least one assertion failed with a wrong value, a thrown exception in application code, or a missing export; a re-run would fail the same way other: Anything that does not fit the options above
triage-policy.md from the repository says what each category means in this project:
markdown# Test triage policy How to read a failed test run in this project. ## Flaky (re-run is enough) - `ECONNRESET`, `ETIMEDOUT`, `ECONNREFUSED`, `fetch failed`, `socket hang up` - "timed out after N ms" from the test runner - `EADDRINUSE` on the test server port - the same test passed on the previous run of the same commit ## Infrastructure (tests did not run) - `npm ci` failed: registry unreachable, lockfile out of sync, `E404`, `ENOSPC` - Docker image pull failed - Node binary missing or wrong version ## Real bug (a re-run will fail the same way) - `AssertionError`, `expected ... to equal ...`, wrong values - `TypeError`, `ReferenceError` or another exception thrown from `src/` - `Cannot find module` pointing at a file in `src/` - a test that fails deterministically with the same message twice ## Not our problem - failures only in `test/experimental/` are known and tracked, treat as `other`
The criteria in the action describe signals, not labels. "Flaky test" on its own tells Jev nothing.
Routing on the answer
Buddy joins the conditions of a single action with AND:
yaml- action: Re-run tests type: BUILD trigger_conditions: - trigger_condition: VAR_IS trigger_variable_key: BUDDY_ACTION_JEV_RESULT trigger_variable_value: flaky - trigger_condition: VAR_GREATER_THAN_OR_EQUAL trigger_variable_key: BUDDY_ACTION_JEV_CONFIDENCE trigger_variable_value: "0.7" commands: npm test
Notify on-call has the same conditions with the value infrastructure. Diagnose with Claude runs on real_bug:
yaml- action: Diagnose with Claude type: CLAUDE_CODE integration: anthropic claude_args: --max-turns 25 dangerously_skip_permissions: true
The prompt tells Claude to find the cause in src/ and write the test name, the file, the line and a diff of the fix to triage.md. The file stays in the pipeline filesystem, so it is worth passing it on right away: as a commit comment through GitHub CLI or in a Slack notification. dangerously_skip_permissions: true is required, because nobody approves the agent's commands in a pipeline. The limits in the prompt are only instructions, not a hard block, so do not give this action secrets that Claude does not need.
Ask a human (WAIT_FOR_APPLY) runs when confidence is below 0.7, regardless of the category.
Build or Fail the run
Build app has an OR condition: green tests, a successful Re-run tests or an approved Ask a human. After the build Deploy to production uploads dist/ to a sandbox and Restart app restarts the application there.
The last action, Fail the run, runs when Build app is skipped and ends with exit 1. Without it a run with a real bug would end with the Successful status, because all the other actions pass or are skipped.
Four runs, one pipeline
Timeout on the exchange rate request
FLAKY_FAIL=1 cuts the timeout to 50 ms, so the test ends with TimeoutError. Jev picks flaky with confidence 0.98. The re-run with the default timeout passes and the application goes to production. In a real project the re-run can fail the same way as the first attempt. Then Build app is skipped and the run ends with the Failed status.
Image loading...
The discount does not apply at exactly 100 USD
src/cart.js has > instead of >=, so a cart worth 100 USD does not get the 10% discount. Jev picks real_bug with confidence 1.00. In nine turns Claude writes the file, the line and a one-line diff to triage.md. The run ends with the Failed status.
Image loading...
HTTP 503 from the rates service
RATES_URL points to https://httpbin.org/status/503, so src/rates.js throws an exception. The log shows an error from src/ and a failure of an external service at the same time. Jev gives flaky 0.54 and real_bug 0.45 with confidence 0.39, so the run waits for a human. The reviewer sees the 503 in the log and clicks Approve. Build and deploy go through.
Image loading...
The test database does not respond
test/setup.js connects to a database that does not exist in the container, so every test ends at startup with ECONNREFUSED. Jev picks infrastructure with confidence 1.00. Notify on-call runs and the run ends with the Failed status.
Image loading...
Things to watch
otherwith high confidence starts no branch. The run ends with theFailedstatus and no notification. If you want to know about such cases, add the valueotherto the conditions ofNotify on-call.- Keep the policy and the criteria in sync. The file lists
ECONNREFUSEDunder flaky. The action criteria put an unreachable database at test startup under infrastructure. In run #12 Jev followed the criteria. When the two disagree, the result depends on which one Jev weighs more, so fix the policy before you rely on it. - Give the agent a turn limit and a specific task. With
--max-turns 12Claude got stuck in planning. With 25 it finished in nine turns. - Check the agent's diagnosis. In one of the test runs Claude found the real bug but added a second one that did not exist.
What next
- TypeSafe integration - how to connect Jev to your workspace,
- Jev action in YAML - all fields of the action,
- Conditional executions - all types of action conditions,
- Dependabot PRs: Semver Is Not a Risk Model - Jev rates the risk of dependency updates.
Jarek Dylewski
Customer Support
A journalist and an SEO specialist trying to find himself in the unforgiving world of coders. Gamer, a non-fiction literature fan and obsessive carnivore. Jarek uses his talents to convert the programming lingo into a cohesive and approachable narration.