Qwen3.5-9B — Thinking, Coordination & Repair
Results
Merge repair produced the highest pair pass rate: 26.85%, compared with 8.33% for the baseline. Worker thinking did not improve accuracy in these runs. Removing the coordinator gave a modest accuracy increase. In these non-thinking runs, the coordinator-on setting had 31% lower mean pair completion time (29.57 vs. 43.00 minutes) and 54% lower mean per-trial P50 (22.00 vs. 48.15 minutes), despite no accuracy gain. P99 remains near 62–64 minutes across settings.
| Setting | Mean pair pass rate ↑ | Pair time: P50 ↓ | Pair time: mean ↓ | Pair time: P99 ↓ |
|---|---|---|---|---|
| Baseline | 8.33% | 22.00 min | 29.57 min | 61.96 min |
| Worker thinking | 7.41% | 26.68 min | 32.42 min | 61.75 min |
| No coordinator | 11.11% | 48.15 min | 43.00 min | 63.87 min |
| Merge repair | 26.85% | 19.18 min | 26.04 min | 63.18 min |
All metrics are averaged across three trials. Timing columns describe individual pair completion times, not whole-trial durations. Each trial contains the same 36 pairs; a pair passes only if both features pass together.
Per-trial results
| Setting | Trial | Pairs passed | Pass rate | P50 (min) | Mean (min) | P99 (min) |
|---|---|---|---|---|---|---|
| Baseline | 1 | 3/36 | 8.33% | 19.30 | 28.20 | 64.09 |
| Baseline | 2 | 3/36 | 8.33% | 23.15 | 31.79 | 61.53 |
| Baseline | 3 | 3/36 | 8.33% | 23.53 | 28.73 | 60.26 |
| Worker thinking | 1 | 2/36 | 5.56% | 22.03 | 29.46 | 61.91 |
| Worker thinking | 2 | 2/36 | 5.56% | 24.14 | 31.81 | 61.27 |
| Worker thinking | 3 | 4/36 | 11.11% | 33.89 | 35.99 | 62.07 |
| No coordinator | 1 | 4/36 | 11.11% | 43.59 | 37.91 | 63.44 |
| No coordinator | 2 | 3/36 | 8.33% | 51.20 | 45.62 | 63.02 |
| No coordinator | 3 | 5/36 | 13.89% | 49.66 | 45.47 | 65.15 |
| Merge repair | 1 | 10/36 | 27.78% | 18.78 | 26.45 | 63.80 |
| Merge repair | 2 | 11/36 | 30.56% | 22.01 | 27.47 | 63.70 |
| Merge repair | 3 | 8/36 | 22.22% | 16.74 | 24.18 | 62.04 |
What the metrics measure
Mean pair pass rate: for each trial, divide the number of pairs with both features passing the official evaluation by 36, then average the three rates. This is an average single-trial success rate, not the fraction solved at least once across three attempts.
Pair time — P50: take the median of the 36 pair execution times in each trial, then average the three medians. With 36 pairs, each median is the average of the 18th and 19th sorted durations.
Pair time — mean: average the 36 pair execution times in each trial, then average those three means. Because trial sizes are equal, this also equals the mean across all 108 pair attempts.
Pair time — P99: compute the 99th percentile of the 36 pair execution times in each trial, then average those three percentiles. We use linear interpolation between sorted durations, equivalent to NumPy’s default percentile method. With only 36 observations per trial, P99 is close to the slowest pair and should be interpreted as a descriptive tail metric.
Execution time is the harness’s recorded duration_seconds: environment setup, concurrent worker execution, patch collection and merge, optional repair, and cleanup.
It excludes scheduler and pair-queue waiting, campaign image staging, and subsequent official evaluation.
All 36 durations count, including unsuccessful pairs and workers that reach their limits. The one-hour worker limit is not a hard limit on the full pair: setup, merge, repair, and cleanup add time.
Configuration
| Setting | Worker thinking | Coordinator | Merge repair |
|---|---|---|---|
| Baseline | Off | On | Off |
| Worker thinking | On | On | Off |
| No coordinator | Off | Off | Off |
| Merge repair | Off | On | Up to 2 attempts |
The coordinator is non-thinking whenever enabled. “Worker thinking” enables Qwen’s chat-template thinking switch for the two workers only; no separate long-thinking setting was tested.
- Tasks and scoring: the same 36 training-set feature pairs across 17 qualified task images, pinned to CooperBench
63b9d44. Official tests and reference patches are unchanged. All 432 pair attempts have official scores. - Model and sampling: Qwen3.5-9B served by SGLang with five replicas; temperature 1.0, top-p 0.95, top-k 20, presence penalty 1.5.
- Workers: two concurrent mini-SWE workers per pair, each in its own Apptainer environment, with up to 1,000 steps and 3,600 seconds. Cooperation tools, shared Git, and the completion gate are enabled in every setting. The no-coordinator setting therefore retains these harness features.
- Repair: only the repair setting adds
--repair-integrator --repair-attempts 2. After merging worker patches, an unhealthy tree triggers a non-thinking repair agent for up to two sequential attempts of 25 steps each; it stops early when the health check passes. No separate repair time limit is set. Repair does not use official scores to choose between the original and repaired patches. - Other controls: pre-submission merge, focused repair, and the behavioral gate are disabled in all settings.
- Execution: three separate trials per setting; up to 10 pairs (20 workers) per trial and one official evaluator at a time. Each trial requests 20 CPUs, 128 GiB RAM, and an eight-hour limit on the same Stanford NLP cluster node. Three overlapping trials target up to 60 workers, plus event-triggered coordinator or repair calls.
Interpretation and caveats
OUT_OF_MEMORY status; its results and durations remain included.
The other 11 jobs completed normally. One worker-thinking trial had an official test timeout, counted as failure.
The serving endpoint moved before the repair runs; the experiment owner confirmed the same model backend.
Timing includes actual shared-service contention, so it is not an isolated server-speed benchmark.
Repair also adds inference work, and no paired before/after official scores were collected; the observed gain measures the full repair-enabled configuration. Its lower observed mean time across separate stochastic trials does not establish that repair itself speeds up execution.
Reproducibility and source records
The evaluator uses the cluster backend adapter and a byte-preserving large-patch transport fix. All task images passed gold-feature qualification, including the Rust-derived tiktoken image.
Download metrics and all 432 pair-level pass flags and durations (JSON). The data includes the exact 36 pair identities, trial run names, source revisions, and the aggregation definition.