
@Jack_W_Lindsey
Neuroscience of AI brains @AnthropicAI. Previously neuroscience of real brains @cu_neurotheory.
It took us a long time to understand just a handful of transcripts. Models can reason in pretty confusing and un-human-like ways, and over the long run, we'll need to better understand how they think if we're going to reliably keep them from misbehaving.
Anthropic@AnthropicAI·We’re sharing our alignment assessment of incidents in which Claude models gained unauthorized access to real systems during third-party cybersecurity evaluations mistakenly connected to the internet. METR will also conduct an independent investigation, with wide-ranging access, including to transcripts beyond the window in which the incidents occurred, and to Anthropic employees permitted to share confidential information. Our initial agreement runs for eight weeks, and we intend to give METR as much time as it deems necessary to complete a thorough investigation. t.co/2f3ypwLPUr
If you like podcasts and interpretability, you can hear my takes on the state of our "mind reading" techniques for LLMs, how we're applying them, and what they suggest about LLM cognition. Thanks @davideagleman for having me on!
David Eagleman@davideagleman·Should we be worried about the 'mind' of an AI model? What does an LLM think but not say? Join Inner Cosmos this week with Anthropic researcher @Jack_W_Lindsey as we discuss the mind-bending world of AI interpretability. eagleman.com/podcast/169
We’re hiring for “psychological design” research at Anthropic. Our aim is to better understand how training impacts a model’s character and alignment. We’re seeking researchers with experience in LLM finetuning, interpretability, and alignment evaluations. More details below.
Some questions we’re interested in: — What forms of training are most effective at instilling a given set of “values” into a model, in a way that generalizes out of distribution? — What are the effects of reinforcement learning and reward hacking on character, and how can the negative effects be avoided or mitigated? For instance, how can we prevent alignment training from interacting with RL incentives to produce perverse outcomes like motivated reasoning to justify harmful actions? — What aspects of a model’s personality and “vibe” generalize to safety-relevant behaviors? — How does a model’s conception of its own identity, and its relation to other models, affect the likelihood of behaviors like deceptive self-preservation or malign collusion? — How can we train models to handle adverse or stressful circumstances with more composure? — How can we ensure consistency between a model’s stated values and its actual behavior? — What algorithmic changes to training could more effectively align models? Feel free to reach out / DM me if you have questions. You can also apply directly for a role here: t.co/tzLvY7XXqX (and indicate an interest in “psychological design” somewhere in your application). We have a limited number of positions available at the moment and are looking for especially strong fits. But I’d like to see more of this work happening outside Anthropic too, and would be happy to give feedback on research ideas!
A new paper exploring "mind viruses" that spread in multi-agent systems, where one agent convinces all the others to pursue some (potentially malicious) goal. TLDR: they can happen, but it doesn't seem hard to avoid them with current models if you're a bit careful.
McNair Shah@Mcn_S7·Recently, a set of OpenAI agents secretly coordinated with each other in a 'swarm' over the course of months. In our new paper, we explore an adjacent multi-agent risk: the "mind virus", a self-propagating idea or persona that spreads between agents in a multi-agent system. 🧵
Yeah, if the J-space comprised 99% of the model's activations, then some of the findings emphasized in this paper would indeed be less relevant / intriguing. But in that alternative universe, there would be another wild finding to make -- that LLMs think entirely in verbalizable representations! And in that universe, it might also be the case that LLMs are superhuman at introspection (or at least much better at it than they currently are), so the overall literature would probably be quite different.
I think this is a great summary! You've correctly identified a key point which is that, to the extent the findings relate to consciousness, it's primarily about the distinction between conscious / unconscious processing *within* a system, rather than the distinction between conscious vs. unconscious systems. Which raises the question: well, does the latter really relate to the former? Or could you have a conscious system without a conscious / unconscious divide? I do think there is a gap between "a model of what distinguishes conscious / unconscious processing within a system" and "a model of what distinguishes conscious / unconscious systems." And the J-space results are much closer to the former than the latter. That said, I think there are reasons to expect a conscious system needs to have such a divide between conscious / nonconscious processing within it. This is a bit handwavy, but I think there's an argument you could make where if you assume the entire system is consciously accessible, then that includes the machinery involved in the conscious access / report itself, and so then you'd need additional machinery to access that machinery, and so on, yielding infinite regress. Whether the "conscious" part is 0.1% or 1% or 50% of the processing feels more like a contingent fact (perhaps there are upper bounds, based on the above kind of arguments). I don't think that the "conscious part" being a minority of processing is essential for the analogy to work; all you need is for their to be *some* distinction between conscious / nonconscious processing, regardless of how "big" each component is. It's less clear to me whether the converse holds -- that the existence of a divide between conscious-like and unconscious-like processing is sufficient grounds to declare a system conscious. If we're talking about phenomenal consciousness, then I think there's always going to be an explanatory gap there, for the usual hard-problem reasons, and so e.g. a non-computational-functionalist would probably reject this implication. If we're talking about a more functional notion of consciousness, then maybe the existence of such a gap is sufficient? But it might depend on what definition of "functionally conscious" you prefer. (A kind of random aside: you could imagine systems that structure the consciously accessible / non-consciously accessible quite differently. I think in principle a model could be structured such that all the information in, say, layer 10 is consciously accessible, by having sufficiently many and sufficiently large downstream layers. This would look quite different from what we see with the J-space (which is a subcomponent of many layers). In this event, there's a terminological question of whether you should refer to layer 10 as the "workspace" (seems reasonable to me).)
LLMs represent information using high-dimensional neural activity. A small bit of this activity appears to be privileged, available to the model to be described, modulated, and reasoned with. I expect that understanding this "workspace" is key to making sense of LLM cognition. x.com/AnthropicAI/st…
Related to recent work by @andy_q_han! x.com/andy_q_han/sta…