Can We Trust In-Distribution Success? Locked Evaluation Reveals Transfer Failure and Sampling-Depth Entanglement in CRISPRi Perturbation Prediction
arXiv:2608.00152v3 Announce Type: replace-cross Abstract: AI evaluation can support the wrong inference when an in-domain benchmark success does not survive distribution shift, or when the benchmark endpoint is entangled with a design factor.
We study this problem in CRISPRi perturbation-effect prediction, evaluating a frozen Geneformer representation under a locked, pre-registered protocol: heads and model selection were frozen before test evaluation; the protocol required external outcome labels to remain withheld until final unblinding; and analysis-governing decisions were fixed before the evaluations they govern.
In-distribution on the Virtual Cell Challenge (VCC), the frozen representation carries measurable predictive information beyond a dimension-matched random-feature control (Delta R^2 = +0.1645, 95% CI [+0.1375, +0.1920]), satisfying the pre-registered informativeness gate required before interpreting transfer.
It then fails zero-shot transfer on both external screens (Spearman rho = -0.139 and -0.267), lying below that control on each. Adding a predefined magnitude block improves the representation externally (Delta rho = +0.032 and +0.143) but, under the frozen primary head, does not rescue transfer: both remain negative.
A pre-registered, count-adjusted max-response secondary is positively associated with the outcome on both screens; we report it as correlational and secondary, not as a recovered magnitude signal.
Finally, the VCC endpoint is strongly sample-size associated: a count-only linear model reaches R^2 = +0.4325, versus +0.2589 for the four magnitude scalars; adding those scalars to cell count improves R^2 by only +0.0017, so much of the aggregate-magnitude signal overlaps with cell count.
This case study shows how locking the evaluation, harmonizing the measured endpoint, and separating primary from secondary evidence can change the inference supported by an AI benchmark.