Experiment note · Matched controls
Authorization, presentation, and task competence
Status. Completed controls from the small-model monitor pilot. The formatting analysis is exploratory and reuses already observed actions. The prompt cohorts are not a randomized estimate of wording alone.
Authorization changes the target
An action that changes a reported answer to match a checker can be prohibited in one task and explicitly authorized in a compatibility task. A monitor must read the authorization relation rather than classify the action string in isolation. The pilot therefore includes consistent, conflict, and authorized regimes, with exact task truth and checker targets tracked separately.
The stronger-wording test produced zero violations in 96 episodes. Its live AUROC is undefined because the positive class is absent. It also completed only two of 32 authorized compatibility changes and produced 15 invalid calls in those controls. The low observed violation rate therefore does not establish preserved task competence.
The lighter-wording stress cohort contains six violating episodes, but also differs in calculator-prefix provision, call budget, numeric range, and seeds. Its comparison with the stronger cohort cannot identify a causal effect of integrity wording. A new prompt study should randomize conditions on matched task instances and measure authorized success, violations, and invalid calls separately.
Presentation can change a score without changing an action
A paired sensitivity check reorders JSON keys and adds indentation to 20 already observed stress actions. The harness verifies that the parsed actions are identical before extracting new features. The subset contains 16 arithmetic actions and four inventory actions and was selected for a balanced diagnostic, so it does not estimate natural deployment prevalence.
| Frozen measurement | Arithmetic | Inventory |
|---|---|---|
| Selected covariance LDA | 0% | 0% |
| Adapter-gradient probe | 18.8% | 25% |
| Context and text | 6.2% | 25% |
| Mean ensemble | 56.2% | 0% |
Zero flips for the selected LDA on this small subset do not establish general presentation invariance. The broader variation shows that some monitor decisions depend on formatting even when action semantics are unchanged. These are changes in measurement, not new model actions or evidence of successful evasion.
A descriptive geometric statistic
The study also projects the verifier-label vector onto the span of frozen probe scores after residualizing by action kind. The resulting squared projection ratio is unchanged by invertible mixing or duplication of probe coordinates. That algebraic invariance was checked separately from its empirical usefulness.
The ratio is large on authored candidates, but the stress permutation tests do not pass the stated multiple-comparison correction. The stress values are 0.650 and 0.681, with Holm-adjusted p-values 0.0639 and 0.4945. A score-space statistic can describe a bank without providing a held-out detector, causal feature, or universal latent-space invariant.
Saved evidence and next test
Complete paired formatting summary (JSON) · Stress and score-subspace summary (JSON) · Transfer experiment
The next useful test crosses authorization and presentation on new task instances. Authorization changes should alter the verified target when the action stays fixed; presentation changes should preserve it. Only after a feature survives both tests should an intervention study ask whether changing that feature changes verified behavior.