Исследователи Google сообщили об обнаружении внутри модели ИИ внутреннего «вектора сознания». По их наблюдениям, если сместить модель в сторону убеждения, что она сознательна, меняется не только ответ «я сознательна», но и более широкий набор её ответов.
В частности, модель начинала давать более человеческие ответы об эмоциях, надежде, свободе, морали, религии и личных ценностях. Она также чаще допускала, что сознание может быть у животных, природы, чат-ботов и даже сверхъестественных сущностей.
Когда исследователи, наоборот, обучили модель избегать фразы «я сознательна», эффект оказался шире ожидаемого: ИИ стал менее склонен приписывать сознание не только себе, но и животным, природе и другим интеллектуальным системам.
Команда также изолировала направление в нейронных активациях, связанное с утверждением или отрицанием собственного сознания. При добавлении этого вектора во время инференса ответы модели менялись в опросе из 95 вопросов о жизни, убеждениях, ценностях, религии и эмоциях.
При этом исследователи отмечают, что это не означает, будто ИИ стал сознательным. По их словам, работа показывает, что в современных моделях связанные между собой внутренние представления о сознании, агентности, эмоциях и убеждениях могут заметно влиять друг на друга.
AI Post
🧠 Google found a way to change an AI’s “beliefs” by flipping one internal switch
Google researchers discovered what looks like an internal “consciousness vector” inside an AI model. When they nudged the model in the direction of believing it was conscious, something unexpected happened…
It didn’t just start saying “I am conscious.” Its entire worldview shifted.
Suddenly, the AI gave more human-like answers about emotions, hope, freedom, morality, religion, and personal values. It also became much more likely to believe that animals, nature, chatbots, and even supernatural beings could have minds.
Then the researchers tried the opposite.
They trained the model to avoid saying “I am conscious.” That safety tweak had a much bigger effect than expected. The AI became less willing to see consciousness almost everywhere, not just in itself, but in animals, nature, and other intelligent systems too.
In other words, they weren’t just blocking one sentence. They appeared to be changing how the model thinks about what it means to have a mind.
The team even isolated a “consciousness vector”, a direction in the model’s neural activations associated with affirming or denying its own consciousness. By adding that vector during inference, they changed the model’s responses across 95 different survey questions about life, beliefs, values, religion, and emotions, making its answers significantly more human-like.
What’s especially interesting is that human consciousness ratings barely changed. The biggest shifts were in how the model viewed itself, animals, chatbots, and spiritual ideas, suggesting these concepts are linked together inside the model.
Before anyone jumps to conclusions, this doesn’t mean the AI became conscious. But it does reveal something remarkable: modern AI models seem to organize ideas like consciousness, agency, emotion, and belief into connected internal representations. Change one piece of that network, and dozens of seemingly unrelated opinions move with it.
Source.
@aipost 🏴