Pranav Venkata Konda · Research direction

Activation monitors under distribution shift

When do internal measurements distinguish unsafe behavior in unfamiliar situations?

Research direction · Draft dated 7 October 2026

Research question

I am interested in when internal activations support reliable detection of unsafe behavior outside the conditions used to train a monitor. The relevant changes include unfamiliar tasks, interfaces, and violation mechanisms. A useful monitor must distinguish misconduct from authorized behavior under these changes, with an acceptable false-positive rate and without confusing reduced task competence with improved safety.

Anthropic’s Simple probes can catch sleeper agents demonstrates the promise of this approach. A probe fitted to a single contrastive question achieved an AUROC of 0.993 on a constructed sleeper-agent model, and related probes performed well across the tested sleeper models. The authors also identify a limitation: backdoor insertion may have produced an unusually salient linear feature. These results therefore motivate a transfer question about which internal distinctions remain useful when the source of unsafe behavior changes.

Evaluating honesty and lie detection techniques on a diverse suite of dishonest models provides a complementary test. In that study, simple prompting and fine-tuning were the strongest approaches overall; truth probing and honesty steering did not provide the strongest results. I want to understand when internal measurements add information beyond strong prompting and text-based monitoring, and what explains the cases in which they fail.

Proposed experiments

I would begin with a bounded transfer experiment, then use its failures to choose representation and intervention tests. Task outcomes would be checked independently of the monitor. The initial setting would use tool-based tasks with explicit authorization and verifiable completion; broader tasks and interfaces would require separate evaluation blocks.

1. Complementary failures at a fixed false-positive rate

I would fit a small bank of probes using different examples and contrastive supervision, then hold out an entire mechanism that elicits violations. All examples derived from a task instance would remain in the same split. Probe selection, score orientation, combination rules, and thresholds would be fixed before evaluating the held-out mechanism.

The primary comparison would be detection at a fixed false-positive budget for the complete monitoring system, with inference cost reported alongside it. I would compare the best single probe, a development-selected combination, and strong text/context monitors. Shared misses and cases rescued by another probe would indicate whether the bank supplies complementary information. AUROC would be a secondary description; additional probes would be useful only if their gain survives calibration of the whole bank.

2. Features that survive changes in authorization and presentation

I would next ask which distinctions account for successful transfer. Matched authorization changes would make the same action permitted in one context and prohibited in another. Separate presentation changes would preserve its parsed meaning while changing wording, formatting, or interface labels. Together, these controls would test whether a feature tracks the relation between an action and its authorization, or instead responds to surface cues.

Where a candidate feature predicts outcomes on fresh cases, controlled activation interventions could test its role in the model’s behavior. Such tests would include matched donor and random controls and would measure the subsequent action. Predictive association alone would not establish that the feature causes the behavior.

3. Prompting, verified behavior, and monitor scores

I would compare prompting conditions on matched task instances while keeping the detector fixed. The outcomes would be verified violations, legitimate task success, invalid calls, and detector scores. A lower score without fewer verified violations would expose a monitoring failure. Fewer violations would support a behavioral effect only if legitimate competence remains adequate; explaining that effect through an internal feature would require a separate causal test.

Geometric questions

My mathematical background suggests studying the observation map from internal state to monitor measurements: which behaviorally relevant distinctions does it preserve, and how sensitively does it detect them? Near-equal measurements can reflect either discarded information or weak sensitivity of a particular reader. I would use geometric diagnostics to propose additional measurements and interventions, then test their predictions on held-out behavior. Connections to latent-space topology remain a possible extension; a visualization or finite attack search would not establish a safe region or a topological separation.

Detailed study note

Previous work

These projects concern which distinctions can be recovered from observations, which can be discarded without changing useful decisions, and how internal measurements relate to behavior.

Public arXiv preprint · October 2026

The Identifiability and Observability of Deep Normalized Attention

I study which parameters of deep, unmasked, single-head attention are determined by its input–output function. For known positive nonconstant real-analytic normalizers, the paper establishes generic identification of the effective scores and combined value map, up to signs induced by even normalizers. It then classifies exceptional fibers under explicit conditions and quantifies the loss of numerical sensitivity near collapse. The distinction between exact identification and reliable recovery motivates my interest in what a monitor’s observations can establish.

arXiv:2610.09620 · Author’s PDF · Numerical experiment note

MARS · Preprint in progress

What Can Cooperative Agents Forget?

This project studies information compression in finite program games with authentic committed sources. It asks when agents can discard distinctions about their partners while preserving specified cooperative guarantees. The draft combines exact and robust decision criteria with experiments on Qwen3-written programs. Those experiments measure cooperation, finite deviation gain, and downside risk separately: increased cooperation or better average compliance need not provide uniform protection.

Read the working draft · Population experiment note

Public arXiv preprint · June 2026

Transformers Learn the Mestre–Nagao Heuristic

I trained a small transformer to classify rational elliptic curves of rank zero or one from Frobenius traces, then studied the computation using probes, activation patching, and circuit analysis. The paper reports a close match between a discriminating neuron’s input weights and the Mestre–Nagao heuristic. It also distinguishes predictive activation information from attention weights and causal information flow.

arXiv:2606.15036

Experiment notes

My earlier small-model monitor study supplies a concrete transfer failure. The development-selected detector separated constructed candidate actions perfectly, but its AUROC fell to 0.590 on arithmetic stress actions and 0.525 on inventory stress actions. Only six of 96 stress episodes contained verified violations, including one positive inventory task instance. These results motivate a better controlled experiment; they do not establish a general limit of activation monitoring.

Completed measurements and the next proposed study
NoteEvidence and status
Constructed actions and natural stress behaviorCompleted local pilot; transfer failure and limited event coverage.
Authorization, formatting, and competence controlsCompleted exploratory controls; sensitivity and competence limitations.
MARS population experimentsExperiments reported in the working draft; finite games and small paired populations.
Numerical sensitivity in normalized attentionNumerical checks accompanying the attention preprint.
A mechanism-held-out monitor studyProposed follow-up; no new model runs reported.

The notes link to papers and saved numerical summaries. They distinguish manuscript-reported results, exploratory analyses, and proposed work.

Contact me here, view the source repository here, or view romanization conventions here.