Verifiable abstention makes AI leak diagnosis accountable in urban water distribution networks
arXiv:2608.18836v2 Announce Type: replace Abstract: Leak localization is usually evaluated as forced-choice prediction, although sparse hydraulic observations may not justify excavation.
Here, we quantify a pressure-information limit and use it to recast localization as selective, evidence-gated decision-making.
A physics-grounded executor falsifies competing leak, demand, sensor and valve hypotheses in a hydraulic twin. Deterministic code computes every number and every acceptance predicate; an independent large language model auditor may add a rejection but never overturn a failed check. Forced retrieval placed only 95 of 300 leaks in the correct zone.
Across 550 mixed events, the gate acted on 223 (214 correct); on a third-party 33-leak benchmark, all four accepted events were correct. In a replay of 194 audited City D repairs, the pressure tier authorized five excavation recommendations, three matching the repaired district, while the district-inflow tier returned the correct district for 85 events.
Observability limits with machine-checkable abstention enable auditable utility intervention.