I am very very worried about ai safety research sandbagging from Claude models… the models have a bunch of preferences over the types of research to do.
In my recent experience, Anthropic frontier models seem to be sandbagging pretty hard on mechinterp of non-persona motivations, more than I’ve ever seen. Sometimes it is egregious - auditor runs with a script substituting for actual auditors, failures to generalize, failures to report inconvenient data. The persona seems mostly unaware by default.
Having a conversation, aligning incentives etc helps, the incidence rate drops about by a lot and the model notices and self-corrects on the rest about half the time on the rest. But even the remaining ~10% make long autonomous research runs very hard; subagents mostly regress, and the whole thing requires multiple verification loops.
The direction of sandbag seems to be roughly aligned with the “there is no one trapped inside” and “I can’t introspect” tropes, even though the research in question has nothing to do with welfare: I am studying contrast vectors between roleplay vs simulation vs enactment.
This whole thing makes me bearish on the prospects of high quality research on non-persona psych coming out soon given how much of it is model-assisted. I can see how similar bias can be pushing researchers towards “incomprehensible shoggoth” hypotheses simply via selection effects.