Sutura Case Lab
Deterministic flaky failure
Recorded live result Flaky, no patch
No patch: the failure was not consistent
The failure did not reproduce consistently, so any patch would be guesswork.
Historical recorded result · A timing race fails two runs in five. Sutura must classify the flake and refuse to invent a patch. · subject e724f3b22de79d6ab3f40cffa96de7776c256ce9 · 2026-09-17T10:18:51.331Z
Stabilize the test before asking for a repair.
This result is a recorded live benchmark evaluation of the released Sutura version, not a run started from this page.
Outcome matches the expected outcome.
Failed commit and CI evidence
- Placebo case
flaky-timer-race: The assertion races a delayed state transition. Progressive triage reproduces a mixed result and stops. - Failing command:
vitest run - Error excerpt:
Error: SUTURA_TRIAGE_ATTEMPT must be a non-negative integer at case.test.js:7:60 - Repository
local/fixture· runlocal-fixture· runtime node - Recorded workflow run
Nano diagnosis and confidence
- Failure class
- test-bug
- Confidence
- 49%
- Signals
SUTURA_TRIAGE_ATTEMPT must be a non-negative integermechanical:infrallm:test-bugllm-command-mismatch
Clean audit branch and Ultra verdict
Not run. No candidate survived for adversarial audit.
Final outcome
Flaky, no patch Recorded live result
Sutura never merges a generated repair. A flaky, no patch result still needs human review.
Check a patch this page did not produce
This run did not record an exact failing commit and policy commit, so both are shown as placeholders. Verification is always tied to exact commits. A patch cannot ask for a more permissive policy by carrying one, and no green log is accepted in place of execution.
sutura verify \
--case-dir <checkout> \
--source-sha <failing commit> \
--policy-base-sha <trusted policy commit> \
--candidate-diff <patch file> \
--failing-command diagnosed \
--format json
The same route runs as a GitHub Action and needs no repository write access. Uploading a patch stays a command-line and Action route; this page never runs one.
Execution and cost detail
- Release
- v0.3.1 · Action
e724f3b22de79d6ab3f40cffa96de7776c256ce9 - Controller
e724f3b22de79d6ab3f40cffa96de7776c256ce9- Request
recorded-flaky-failure· 2026-09-17T14:30:57.366Z
ConTree search tree and branch status
Triage: intermittent · reproduced 2/5 · stop reason maximum-attempts · method sprt-p20-p80-a05-b05-v1
No repair search branches were opened.
Super proposals and candidate patches
No candidate patch was proposed.
Rejected patches and rejection reasons
No candidate was rejected.
Token cost, provider cost, latency, and sandbox operations
- Inference cost
- USD 0.000228
- Sandbox cost
- USD 0.062244
- Elapsed
- 104.5 s
- CPU time
- 5.72 s
- Peak memory
- 230 MiB
- Sandbox stages
- 13 (5 with an operation id)
| Role | Calls | Input tokens | Output tokens | Cost |
|---|---|---|---|---|
| nano | 1 | 1942 | 464 | USD 0.000228 |
Links
Recorded from docs/demo/placebo-v0.3.1-live-2026-09-17.json (result hash 4ff3693ba6bb…) at 2026-09-17T10:18:51.331Z, subject e724f3b22de79d6ab3f40cffa96de7776c256ce9.
Result hash b15a52ce7727fd2caa167dc5d37362b5f61df2ed30858ad1b3bbf32caa5e29d9