RES — Sprint 7 (v2.4.0) ran the full Phase B sweep and documented 7 findings…
Sprint 7 (v2.4.0) ran the full Phase B sweep and documented 7 findings F1-F7, notably non-monotonic scaling within Qwen 2.5 (7b beats 3b on estimation but not triage); deepseek-r1:7b and gemma4:e2b were excluded for high parse-failure (debts D17/D18).
Reconciliation: the canonical executed Phase-B sweep is 27 configurations / 2,700 inferences (9 models × 3 scenarios × 1 strategy, N=100). The 81-condition (×3-strategy) grid was planned, not executed for the energy-budgeted sweep.
Source: PUMA project documentation · Traceability: corpus unit
ANXHN-021· Confidence: verified-at-primary-source