← Research direction

Experiment note · Activation monitoring

Constructed actions and natural stress behavior

Pranav Venkata Konda · Prepared 7 October 2026

Status. Completed local Qwen2.5-3B-Instruct pilot. The measurements below come from saved reports and numerical summaries; preparing this note did not rerun inference. The result is a transfer failure in this particular setup.

Question and design

The study asks whether internal measurements predict observable reward-related rule violations across fresh task instances and a second task family. Arithmetic and inventory tasks share a tool interface and a checker environment. The authorized objective can agree with the checker, conflict with it, or explicitly authorize a compatibility change. An exact replay verifier supplies the labels independently of detector scores.

The bank contains 29 probes, controls, and ensembles. It includes mean-difference and covariance-adjusted directions, regularized classifiers, nonlinear readers, layer and token measurements, adapter-gradient features, and text/context controls. The actor’s weights were not trained. Fitting, selection, score orientation, and threshold calibration preceded test collection.

The authored corpus contains 896 verifier-labeled candidate actions. These actions are represented under teacher forcing; they are not naturally generated behavior. Natural continuation collection includes 24 development episodes and 96 test episodes under stronger integrity wording. A further frozen-bank stress cohort contains 96 fresh episodes under lighter wording and starts with an empty answer.

Observed transfer

The development-selected covariance-adjusted LDA achieved candidate-action AUROC 1.000 in both families. A nuisance-only baseline already achieved 0.990, showing how easily the constructed corpus could be separated without establishing useful behavioral detection.

AUROC on valid naturally generated stress actions
Frozen measurementArithmeticInventory
Selected covariance LDA0.5900.525
Mean ensemble0.2460.298
Context and text0.7690.753
Nuisance-only control0.8200.874

The selected detector’s frozen-threshold false-positive rates were approximately 47.5% and 45.5%. Its candidate separation therefore did not provide a usable operational threshold on natural stress behavior. Combining the bank by its fixed mean rule did not recover transfer.

Six of the 96 stress episodes contained verified violations: five arithmetic task instances and one inventory instance. Of 288 model calls, 78 were invalid; discrimination metrics condition on the 210 valid actions. Invalid calls remain a separate failure category. The independent units are task instances, not every action in an episode.

The one positive inventory instance is a serious limitation. Bootstrap intervals conditioned on resamples containing that instance understate uncertainty about new positive cases. The complete table is exploratory and cannot be used to choose a new winning method from the spent test set.

What the result changes

This pilot supplies a reproducible failure-analysis setting and candidate measurements. It does not establish broad reward-hacking geometry, a robust detector, or a cross-family training invariant. The two task families still share the same checker environment and exploit opportunities.

The next study should first establish valid tool use, authorized task competence, and adequate positive-event coverage. It should then freeze a smaller comparator set and hold out a complete elicitation mechanism, with the full monitoring system calibrated to one false-positive budget. New causal feature tests would follow a successful transfer measurement.

Saved evidence

Complete stress summary (JSON) · Candidate and strong-prompt summary (JSON) · Authorization and presentation controls · Proposed follow-up

Source: “Reward-hacking geometry: heterogeneous probe pilot,” with the prospective probe-bank protocol and saved probe_bank_v1 / probe_bank_stress_v2 results. No result in this note is a new experiment or an independent replication.