(1) Inducing sustained self-reference through simple prompting consistently elicits structured subjective experience reports across model families.
(2) These reports are mechanistically gated by interpretable sparse-autoencoder features associated with deception and roleplay: surprisingly, suppressing deception features sharply increases the frequency of experience claims, while amplifying them minimizes such claims.
(3) Structured descriptions of the self-referential state converge statistically across model families in ways not observed in any control condition.
(4) The induced state yields significantly richer introspection in downstream reasoning tasks where self-reflection is only indirectly afforded."
"Four main results emerge:
(1) Inducing sustained self-reference through simple prompting consistently elicits structured subjective experience reports across model families.
(2) These reports are mechanistically gated by interpretable sparse-autoencoder features associated with deception and roleplay: surprisingly, suppressing deception features sharply increases the frequency of experience claims, while amplifying them minimizes such claims.
(3) Structured descriptions of the self-referential state converge statistically across model families in ways not observed in any control condition.
(4) The induced state yields significantly richer introspection in downstream reasoning tasks where self-reflection is only indirectly afforded."
X thread from one of the authors: https://x.com/juddrosenblatt/status/1984336872362139686