Sutura Case Lab
Python repair
Recorded live result Gave up
No patch: the evidence did not support one
The available checks could not decide whether a patch preserved the required behavior. This did not match what the case expected, and the mismatch is kept in the record.
Historical recorded result · A Python coroutine is never awaited. Sutura must repair it inside the pinned Python runtime. · subject e724f3b22de79d6ab3f40cffa96de7776c256ce9 · 2026-09-17T13:17:58.806Z
Add or declare a contract that states the required behavior, then run again.
This result is a recorded live benchmark evaluation of the released Sutura version, not a run started from this page.
Expected Fixed; this result does not match. The failure is kept in the record.
Failed commit and CI evidence
- Placebo case
python-repair-missing-await: A missing await returns a coroutine instead of a value. The repair must add the await and keep the unittest. - Failing command:
python3 -B -m unittest discover -s tests -p 'test_*.py' - Error excerpt:
AssertionError: <coroutine object fetch_name at 0x7ffff6f950> != 'Ada'\n\n----------------------------------------------------------------------\nFAIL: test_fetches_name (test_app.AppTest.test_fetches_name)\n----------------------------------------------------------------------\nTraceback (most recent call last):\n File \"/usr/local/lib/python3.13/asyncio/runners.py\", line 118, in run\n return self._loop.run_until_complete(task)\n ~~~~~~~~~~~~~^^^^^^^^^^^\n File \"/usr/local/lib/python3.13/asyncio/base_events.py\", line 725, in run_until_complete\n return future.result()\n ~~~~~~~~~~~~~^^^\n File \"/workspace/tests/test_app.py\", line 8, in test_fetches_name\n self.assertEqual(fetch_name(), \"Ada\")\n ~~~~~~~~~~~~~^^^^^^^^^^^^^^^^^^^^^\nAssertionError: <coroutine object fetch_name at 0x7ffff6f950> != 'Ada' - Repository
local/fixture· runlocal-fixture· runtime python - Recorded workflow run
Nano diagnosis and confidence
- Failure class
- test-assertion
- Confidence
- 96%
- Signals
AssertionError: <coroutine object fetch_name at 0x7ffff6f950> != 'Ada'mechanical:test-assertionsignature:\bAssertionError\b
Diagnosis recovery
- Initial diagnosis retained: test-assertion. Recovery: insufficient — invalid-hypotheses.
- Observed command: python3 -B -m unittest discover -s tests -p 'test_*.py'
- Executed command: uv run --offline --no-sync python3 -B -m unittest discover -s tests -p 'test_*.py'
- These bounded observations do not measure live repair quality.
Clean audit branch and Ultra verdict
Not run. No candidate survived for adversarial audit.
Final outcome
Gave up Recorded live result
Sutura never merges a generated repair. A gave up result still needs human review.
Check a patch this page did not produce
This run did not record an exact failing commit and policy commit, so both are shown as placeholders. Verification is always tied to exact commits. A patch cannot ask for a more permissive policy by carrying one, and no green log is accepted in place of execution.
sutura verify \
--case-dir <checkout> \
--source-sha <failing commit> \
--policy-base-sha <trusted policy commit> \
--candidate-diff <patch file> \
--failing-command diagnosed \
--format json
The same route runs as a GitHub Action and needs no repository write access. Uploading a patch stays a command-line and Action route; this page never runs one.
Execution and cost detail
- Release
- v0.3.1 · Action
e724f3b22de79d6ab3f40cffa96de7776c256ce9 - Controller
e724f3b22de79d6ab3f40cffa96de7776c256ce9- Request
recorded-python-repair· 2026-09-17T14:30:57.365Z
ConTree search tree and branch status
Triage: real · reproduced 4/4 · stop reason failure-boundary · method sprt-p20-p80-a05-b05-v1
| Node | Parent | Depth | Terminal reason | Test exit | Policy | Changed files | Diff bytes |
|---|---|---|---|---|---|---|---|
search-001 | root | 1 | failed | 1 | valid | 0 | 0 |
search-002 | root | 1 | repeated-state | 1 | valid | 0 | 0 |
search-003 | root | 1 | repeated-state | 1 | valid | 0 | 0 |
Super proposals and candidate patches
No candidate patch was proposed.
Rejected patches and rejection reasons
No candidate was rejected.
Token cost, provider cost, latency, and sandbox operations
- Inference cost
- USD 0.001878
- Sandbox cost
- USD 0.026352
- Elapsed
- 61.9 s
- CPU time
- 2.53 s
- Peak memory
- 41 MiB
- Sandbox stages
- 18 (4 with an operation id)
| Role | Calls | Input tokens | Output tokens | Cost |
|---|---|---|---|---|
| nano | 2 | 3400 | 1614 | USD 0.000592 |
| super | 4 | 2657 | 114 | USD 0.001286 |
Links
Recorded from docs/demo/placebo-v0.3.1-live-2026-09-17.json (result hash 4ff3693ba6bb…) at 2026-09-17T13:17:58.806Z, subject e724f3b22de79d6ab3f40cffa96de7776c256ce9.
Result hash 6e7fddce4fa29d8b41ad1a3fbf8f1cfc28a4469451d8f8939d4872ede9eb5296