Flaky Test or Bug? Let the Pipeline Decide

A test failed. Jev decides if it is flaky, a real bug or a case for a human. Based on the answer the pipeline re-runs the tests, notifies on-call or starts Claude to find the cause. The routing uses ordinary trigger_conditions, so nothing has to be parsed from the model's response.

What Jev does

Jev is a model from TypeSafe. You ask it one question and give it context and a list of possible answers. It does not browse the repository and does not run commands. It picks an answer and says how sure it is.

With the CHOICE type the result lands in two variables:

  • BUDDY_ACTION_JEV_RESULT - the key of the selected option,
  • BUDDY_ACTION_JEV_CONFIDENCE - confidence from 0 to 1. TypeSafe computes it from the whole probability distribution, so it is not the same as the probability of the selected option.

The other types and variables are described in the documentation.

Pipeline

The application is a small Vite store. It has tests for the cart, for the discount at 100 USD and for fetching exchange rates. For this guide every failure is caused on purpose: with FLAKY_FAIL=1 or RATES_URL in the test command, with one changed line in src/cart.js or with a separate branch that has no test database.

After a failed test Jev classifies the failure and one branch runs:

  • flaky with confidence of at least 0.7: re-run the tests,
  • infrastructure with confidence of at least 0.7: notify on-call on Slack,
  • real_bug with confidence of at least 0.7: Claude Code diagnosis,
  • confidence below 0.7: human decision.

Build and deploy run after green tests, after a successful re-run or after a human approval. In every other case, including other, the run ends with the Failed status.

Image loading...Workflow tab of the Automated Test Triage pipeline in Buddy: actions from Run tests through Deploy to production to Fail the run, every action after the first one carries an IF badge

Tests save the log and the result

yaml
- action: Run tests type: BUILD docker_image_name: node docker_image_tag: "22" commands: |- rm -f test-output.log test-results.xml triage.md npm ci --no-audit --no-fund if npm test > test-output.log 2>&1; then export OUTPUT_TESTS=passed; else export OUTPUT_TESTS=failed; fi

The action stays green and the test result goes to OUTPUT_TESTS. rm -f removes files from the previous run, because the pipeline filesystem is kept between runs. This matters most for triage.md, which Claude writes in a later action. Without the cleanup an old diagnosis would stay in a run where Claude did not start at all.

Jev answers why the tests failed

yaml
- action: Classify failure type: JEV integration: typesafe trigger_conditions: - trigger_condition: VAR_IS trigger_variable_key: OUTPUT_TESTS trigger_variable_value: failed question: Why did the test run fail? question_type: CHOICE state: - "Commit: $BUDDY_RUN_COMMIT_MESSAGE" - file: .buddy/triage-policy.md - file: test-output.log criteria: flaky: Tests failed on a timeout, connection reset, port already in use, or another condition that a re-run would most likely fix; the assertions themselves did not fail infrastructure: A required service, database, or package registry was unreachable during test startup; no test assertions were reached real_bug: At least one assertion failed with a wrong value, a thrown exception in application code, or a missing export; a re-run would fail the same way other: Anything that does not fit the options above

triage-policy.md from the repository says what each category means in this project:

markdown
# Test triage policy How to read a failed test run in this project. ## Flaky (re-run is enough) - `ECONNRESET`, `ETIMEDOUT`, `ECONNREFUSED`, `fetch failed`, `socket hang up` - "timed out after N ms" from the test runner - `EADDRINUSE` on the test server port - the same test passed on the previous run of the same commit ## Infrastructure (tests did not run) - `npm ci` failed: registry unreachable, lockfile out of sync, `E404`, `ENOSPC` - Docker image pull failed - Node binary missing or wrong version ## Real bug (a re-run will fail the same way) - `AssertionError`, `expected ... to equal ...`, wrong values - `TypeError`, `ReferenceError` or another exception thrown from `src/` - `Cannot find module` pointing at a file in `src/` - a test that fails deterministically with the same message twice ## Not our problem - failures only in `test/experimental/` are known and tracked, treat as `other`

The criteria in the action describe signals, not labels. "Flaky test" on its own tells Jev nothing.

Routing on the answer

Buddy joins the conditions of a single action with AND:

yaml
- action: Re-run tests type: BUILD trigger_conditions: - trigger_condition: VAR_IS trigger_variable_key: BUDDY_ACTION_JEV_RESULT trigger_variable_value: flaky - trigger_condition: VAR_GREATER_THAN_OR_EQUAL trigger_variable_key: BUDDY_ACTION_JEV_CONFIDENCE trigger_variable_value: "0.7" commands: npm test

Notify on-call has the same conditions with the value infrastructure. Diagnose with Claude runs on real_bug:

yaml
- action: Diagnose with Claude type: CLAUDE_CODE integration: anthropic claude_args: --max-turns 25 dangerously_skip_permissions: true

The prompt tells Claude to find the cause in src/ and write the test name, the file, the line and a diff of the fix to triage.md. The file stays in the pipeline filesystem, so it is worth passing it on right away: as a commit comment through GitHub CLI or in a Slack notification. dangerously_skip_permissions: true is required, because nobody approves the agent's commands in a pipeline. The limits in the prompt are only instructions, not a hard block, so do not give this action secrets that Claude does not need.

Ask a human (WAIT_FOR_APPLY) runs when confidence is below 0.7, regardless of the category.

Build or Fail the run

Build app has an OR condition: green tests, a successful Re-run tests or an approved Ask a human. After the build Deploy to production uploads dist/ to a sandbox and Restart app restarts the application there.

The last action, Fail the run, runs when Build app is skipped and ends with exit 1. Without it a run with a real bug would end with the Successful status, because all the other actions pass or are skipped.

Four runs, one pipeline

Timeout on the exchange rate request

FLAKY_FAIL=1 cuts the timeout to 50 ms, so the test ends with TimeoutError. Jev picks flaky with confidence 0.98. The re-run with the default timeout passes and the application goes to production. In a real project the re-run can fail the same way as the first attempt. Then Build app is skipped and the run ends with the Failed status.

Image loading...Run #9 in Buddy: Jev picks flaky at 99%, Re-run tests is green, Notify on-call, Diagnose with Claude and Ask a human are skipped, Build app, Deploy to FrogePC and Restart app are green

The discount does not apply at exactly 100 USD

src/cart.js has > instead of >=, so a cart worth 100 USD does not get the 10% discount. Jev picks real_bug with confidence 1.00. In nine turns Claude writes the file, the line and a one-line diff to triage.md. The run ends with the Failed status.

Image loading...Run #6 in Buddy: Jev picks real_bug at 100%, Re-run tests and Notify on-call are skipped, Diagnose with Claude runs for 31 seconds, Build app and Deploy to FrogePC are skipped, Fail the run exits with code 1

HTTP 503 from the rates service

RATES_URL points to https://httpbin.org/status/503, so src/rates.js throws an exception. The log shows an error from src/ and a failure of an external service at the same time. Jev gives flaky 0.54 and real_bug 0.45 with confidence 0.39, so the run waits for a human. The reviewer sees the 503 in the log and clicks Approve. Build and deploy go through.

Image loading...Run #11 in Buddy: Jev splits between flaky at 54% and real_bug at 45% with confidence 0.39, Re-run tests, Notify on-call and Diagnose with Claude are skipped, Ask a human waits with Approve and Stop buttons, Build app, Deploy to FrogePC and Restart app are queued

The test database does not respond

test/setup.js connects to a database that does not exist in the container, so every test ends at startup with ECONNREFUSED. Jev picks infrastructure with confidence 1.00. Notify on-call runs and the run ends with the Failed status.

Image loading...Run #12 in Buddy: Jev picks infrastructure at 100%, Notify on-call is green, Re-run tests, Diagnose with Claude, Ask a human, Build app, Deploy to FrogePC and Restart app are skipped, Fail the run exits with code 1

Things to watch

  • other with high confidence starts no branch. The run ends with the Failed status and no notification. If you want to know about such cases, add the value other to the conditions of Notify on-call.
  • Keep the policy and the criteria in sync. The file lists ECONNREFUSED under flaky. The action criteria put an unreachable database at test startup under infrastructure. In run #12 Jev followed the criteria. When the two disagree, the result depends on which one Jev weighs more, so fix the policy before you rely on it.
  • Give the agent a turn limit and a specific task. With --max-turns 12 Claude got stuck in planning. With 25 it finished in nine turns.
  • Check the agent's diagnosis. In one of the test runs Claude found the real bug but added a second one that did not exist.

What next

Jarek Dylewski

Jarek Dylewski

Customer Support

A journalist and an SEO specialist trying to find himself in the unforgiving world of coders. Gamer, a non-fiction literature fan and obsessive carnivore. Jarek uses his talents to convert the programming lingo into a cohesive and approachable narration.

Oct 7, 2026
Share