Sutura Case Lab

Deterministic flaky failure

Recorded live result Flaky, no patch

No patch: the failure was not consistent

The failure did not reproduce consistently, so any patch would be guesswork.

Historical recorded result · A timing race fails two runs in five. Sutura must classify the flake and refuse to invent a patch. · subject e724f3b22de79d6ab3f40cffa96de7776c256ce9 · 2026-09-17T10:18:51.331Z

Stabilize the test before asking for a repair.

This result is a recorded live benchmark evaluation of the released Sutura version, not a run started from this page.

Outcome matches the expected outcome.

Failed commit and CI evidence

Nano diagnosis and confidence

Failure class
test-bug
Confidence
49%
Signals
  • SUTURA_TRIAGE_ATTEMPT must be a non-negative integer
  • mechanical:infra
  • llm:test-bug
  • llm-command-mismatch

Clean audit branch and Ultra verdict

Not run. No candidate survived for adversarial audit.

Final outcome

Flaky, no patch Recorded live result

Sutura never merges a generated repair. A flaky, no patch result still needs human review.

Check a patch this page did not produce

This run did not record an exact failing commit and policy commit, so both are shown as placeholders. Verification is always tied to exact commits. A patch cannot ask for a more permissive policy by carrying one, and no green log is accepted in place of execution.

sutura verify \
  --case-dir <checkout> \
  --source-sha <failing commit> \
  --policy-base-sha <trusted policy commit> \
  --candidate-diff <patch file> \
  --failing-command diagnosed \
  --format json

The same route runs as a GitHub Action and needs no repository write access. Uploading a patch stays a command-line and Action route; this page never runs one.

Execution and cost detail
Release
v0.3.1 · Action e724f3b22de79d6ab3f40cffa96de7776c256ce9
Controller
e724f3b22de79d6ab3f40cffa96de7776c256ce9
Request
recorded-flaky-failure · 2026-09-17T14:30:57.366Z

ConTree search tree and branch status

Triage: intermittent · reproduced 2/5 · stop reason maximum-attempts · method sprt-p20-p80-a05-b05-v1

No repair search branches were opened.

Super proposals and candidate patches

No candidate patch was proposed.

Rejected patches and rejection reasons

No candidate was rejected.

Token cost, provider cost, latency, and sandbox operations

Inference cost
USD 0.000228
Sandbox cost
USD 0.062244
Elapsed
104.5 s
CPU time
5.72 s
Peak memory
230 MiB
Sandbox stages
13 (5 with an operation id)
RoleCallsInput tokensOutput tokensCost
nano11942464USD 0.000228

Recorded from docs/demo/placebo-v0.3.1-live-2026-09-17.json (result hash 4ff3693ba6bb…) at 2026-09-17T10:18:51.331Z, subject e724f3b22de79d6ab3f40cffa96de7776c256ce9.

Result hash b15a52ce7727fd2caa167dc5d37362b5f61df2ed30858ad1b3bbf32caa5e29d9