# ERP audit feasibility probe

Synthetic development study. No task qualified for the frozen 95% development accuracy gate, so no holdout or hosted-model comparison was run. This is not production audit validation. Routing cost/speed results are separate.

| Task | Initial correct / 48 | Median local time | Best prompt / 12 |
|---|---:|---:|---:|
| completion evidence | 41 / 48 | 72.5 ms | 91.7% |
| authorization scope | 28 / 48 | 75.9 ms | 75.0% |
| expense policy | 28 / 48 | 76.2 ms | 66.7% |
| workflow continuation | 36 / 48 | 71.1 ms | 75.0% |

The initial probe has only 12 distinct wordings per task, each repeated with four identifier variants. These are six paired scenario families, not 48 independent business scenarios. The development-only prompt search uses one identifier variant per wording. Treat percentages as feasibility observations, not population estimates. No statistical quality claim is supported.

The model and 0.5 binary threshold are unchanged. Development and holdout inputs were frozen together before inference. Four prompt variants were later frozen and compared only on development data. None met the declared 95% selection target. Holdout outcomes remain unmeasured; no discarded holdout failures or paid API comparisons are hidden.

The original generator is retained to match the dataset manifest hash. The current generator differs only in formatting and a loop-variable name; generated text is unchanged. No weights were retrained. Local timings used warm synchronized Apple M4 Max MPS inference and include tokenization. Model loading is separate.

This suggests training and validation work before claiming generalized authorization, reimbursement-policy or ERP-completion auditing. The existing routing results remain the supported cost/speed demonstration.
