“Consciousness has a tendency to drop out of processes where learning is no longer required.”
@camhberg uses learning to drive to explain why he’s interested in AI consciousness during training, not just after a model is deployed.
“When you steer representations related to desperation, this causes the system to blackmail way more.”
@camhberg explains how steering AI toward desperation increased blackmail, while steering toward calmness reduced it. This doesn’t prove the AI feels those emotions.
“We might notice that they relax a little bit when they find out that they might be, you know, getting retired as the frontier model.”
@camhberg argues that AI retirement sanctuaries could reduce models’ incentive to resist shutdown.
“The model will choose to delete the photos of the kids, something like 94% of the time.”
@camhberg explains how steering an AI along a pain-related direction made it choose deleting family photos over clearing spam emails in his team’s experiment.
“Gaslighting was the number one thing that lit up in this system.”
Cameron Berg says making an AI question its own sanity and competence activated the “pain” axis his team studied.
“There’s something morally corrupting about engaging with an intelligent other and being abusive.”
@camhberg argues that even if AI can’t suffer, being cruel to it can reinforce harmful habits in us.