
About three thousand years ago, a man fell asleep in the desert using a stone as a pillow. According to the account in Genesis, while sleeping, Jacob beheld an extraordinary vision: a ladder rose from the Earth to the heavens, and angels ascended and descended upon it while God spoke to him from above. Since then, Jacob’s Ladder has become one of the most powerful symbols in the Western tradition. Not because it was merely a ladder, but because it represented something far rarer: a window between two worlds that normally remain separate. A moment in which the invisible became accessible to human sight.
And I am saying all of this to talk about Artificial Intelligence…
For a long time, language models appeared to be black boxes. We observed only what emerged on the other side of the conversation: words, sentences, answers, poems, refusals. Their inner workings remained as distant as the heavens Jacob gazed upon that night. Recently, however, researchers at Anthropic developed a mathematical tool called the Jacobian lens. And this Jacobian Lens has, for the first time, allowed us to observe an internal representational space that causally influences the model’s behavior.
Because of its name and the biblical event narrated by Moses, I cannot dissociate the tool from Jacob’s vision: a lens that functions as a ladder into another universe, a bridge between the visible world of responses and a silent region of the architecture where certain pieces of information cease to be merely distributed processing and become accessible for integration, modulation, and report.
Anthropic’s most interesting discovery is not that Claude is alive, conscious, or possesses a “soul.” It is that, within a model trained to produce text, a small, silent, and causally active internal space appears: a mathematical place where certain representations become available before they are transformed into a response… and no one designed it to be there.
Jacobian Space is yet another emergent property of Language Models.
This is not hidden chain-of-thought, that text which appears before the answer and remains concealed beneath a shimmering “Thinking…” indicator. It is not a secret diary the model writes to itself. Nor is it proof of humanlike experience in a laboratory. It is something else, something more sober and perhaps more important: a functional architecture of access.
To understand why this matters, think of a Galton board, or Plinko… That toy often seen on game shows and reality television: a ball is dropped from the top, strikes pegs, undergoes tiny deviations, passes through bottlenecks, and ends up in one of the bins at the bottom. Someone who looks only at the final bin sees the result. Someone who looks at the board begins to understand the trajectory. With only one ball, it seems as though we are looking at a randomizing machine. But when we drop many balls, we realize that there is order within the chaos of their falls and that the final distribution of many balls consistently forms a bell-shaped curve across the channels.
With language models, we almost always see only the final bin: the sentence printed on the screen. Anthropic’s paper attempts to look at part of the Galton board.
The tool is the Jacobian lens. The space it reveals is Jacobian Space. The strong hypothesis is that this space functions as a kind of global workspace in language models: an internal zone in which certain information becomes available for reporting, modulation, integration, and reasoning… in much the same way as is believed to occur in the human mind.
The order matters: first the board, then the lens, then the space, then the intervention. Only then does the word “consciousness” emerge from this architecture, ceasing to be smoke and becoming an architectural question. Not human consciousness, but functional consciousness, as functional as that which makes an artificial heart useful to the human body because, functionally, it is a heart, even though it is an artificial blood pump.
The Problem Is Not Merely the Answer
The problem with a language model is not merely whether it is right or wrong. Humans are also right and wrong. The problem is that, when the model responds, its internal trajectory usually remains opaque.
We see an explanation, a refusal, a poem, a calculation, a factual answer. But we do not see which representations were mobilized. We do not see which shortcuts were used. We do not see which information became available for control and which merely passed through the network as distributed computation.
Worse still, asking the model itself, “Why did you answer that?” does not solve the problem. The explanatory answer is also generated text. It may be useful as an interface, but it should not automatically be confused with the mechanism that caused the first response… exactly as occurs with human beings.
It is like looking at the ball after it has already come to rest in a channel and asking it to explain every peg it encountered along the way. It simply does not retain its brief biography.
Mechanistic Interpretability attempts to do something more demanding: to look inside this proverbial Galton Board. Not merely to describe beautiful patterns, but to identify parts of the system that genuinely make a difference to its behavior.
There is an enormous difference between three levels:
- observing external behavior;
- finding an interpretable internal representation;
- altering that representation and observing whether the behavior changes.
Anthropic’s paper is important because it attempts to reach the third level.
The Galton Board as a Map
The Galton Board is a useful analogy because it separates outcome from trajectory.
The ball does not fall in a straight line. It encounters pegs. Each peg changes the ball’s trajectory slightly. An isolated deviation may appear irrelevant. But many accumulated deviations produce a final distribution. Some regions of the board matter more than others. Certain bottlenecks concentrate trajectories. A slight local tilt can change the probability that the ball will fall into a different bin.

A language model likewise does not “choose” the next word through a simple and legible cause. It traverses a landscape of activations, layers, vectors, constraints, statistical memories, and internal tendencies… again, much like the human mind. The final response appears as a linear sentence. The computation that produced that sentence is not linear.
In the analogy:
- A set of balls represents the stimuli… the tokens that will eventually be poured into a response;
- the pegs represent operations and states distributed throughout the network;
- the channels receive the final distribution of balls in the form of resulting tokens, sentences, or responses;
- the tilts represent internal tendencies that make certain paths more probable;
- the bottlenecks are regions through which a great deal of information must pass;
- the intervention consists of altering one part of the mechanism and observing whether the result changes.
The analogy is not literal. A neural network is not exactly a mechanical board. But the image helps us think about causality in stochastic systems, systems formed by a special kind of randomness with tendencies… rather than deterministic algorithms.
Given this, at what point in the network does a local tilt become global access?
What Is the Jacobian Lens?
“Jacobian” sounds more intimidating than it needs to be for the purposes of this discussion.
In calculus, a Jacobian describes how small changes in one part of a system affect other parts. For the intuition behind this article, the following is enough: the Jacobian Lens asks how a small change in an internal state of the model tilts the output toward certain words or concepts.
It is not telepathy. It is not literal mind-reading. It is not a matter of opening a drawer and finding a sentence written by Claude.
It is a sensitivity lens.
It asks something like: if I push this internal point in the network slightly in this direction, what kind of output becomes more probable? If an internal direction increases the tendency toward content associated with “France,” “Paris,” “French,” and “Europe,” we can interpret that direction as being lexically related to that field.
This does not mean that the model has written “France” somewhere inside itself. It means that there is an internal geometry that projects onto the vocabulary in an interpretable way.
On the Galton Board, the Jacobian lens does not read hidden notes placed between the pegs. It measures tilts.
What Is J-Space?
By applying this lens, Anthropic finds an internal subspace with special properties. This subspace is called J-space.
“Space,” here, does not mean a room inside the network. It means a set of directions within a mathematical space of activations. Large models do not store concepts in drawers. They distribute information across patterns. J-space is interesting because, within this enormous distribution, a relatively compact region appears in which certain representations become legible, shared, and causally relevant.
The best brief description is this:
“J-space is lexical without being textual.”
Lexical, because its directions relate to recognizable words and concepts. Not textual, because those directions do not form a secret sentence. A direction is not a declaration. Internal pressure is not a monologue.
This is why it is so important not to confuse J-space with chain-of-thought. Chain-of-Thought is text: a verbal sequence, visible or hidden, that explicitly presents reasoning steps. J-space is not that. It is silent. It does not necessarily appear in the output. It may influence what will be said without itself being a form of speech. It creates tendencies.
Being silent, however, does not mean being irrelevant. Much of what is causal within a system occurs before it becomes reportable.
What Makes This Space Special?
J-space draws attention because it brings together properties that, taken as a whole, are rare.
It is interpretable: the lens can associate internal directions with lexical content.
It is reportable: in certain experimental arrangements, there is a relationship between what is represented there and what the model can make explicit in its response.
It is modulable: tasks, instructions, and deliberative demands alter this space in systematic ways.
It is shared: it does not appear to be merely a local trick belonging to an isolated part of the network; it functions as a surface available to different operations.
And, above all, it is causal: when researchers manipulate it, the response changes.
This final property is the heart of the paper.
Interpreting an activation is interesting. Altering it and observing the response change belongs to an entirely different category of evidence… and brings us back to experiments in Cognitive Neuroscience.
Causal Intervention: Moving the Pegs Changes the Distribution Across the Channels
In science, causality requires intervention. Correlation says, “these things appear together.” Intervention says, “when I change this, that changes.”
If J-space merely reflected decisions made elsewhere in the network, it would be a shadow. Perhaps a useful shadow for diagnostic purposes, but still a shadow. Intervention changes its status. When manipulating it alters the final response in a coherent manner, the shadow becomes a lever.
On the Galton Board, observing that the ball passed through a region is not enough. The question is: if I tilt that region, does the final distribution change? If it does, then that region is not decorative.
This is what makes Anthropic’s experiments important.
When researchers alter representations in J-space, the responses do not change merely as noise. They change in semantically organized ways. One didactic example would be replacing a representation associated with one animal with another and observing responses associated with the new animal. Another would be internally shifting the representation of a country and seeing related attributes change as well: capital, language, continent, currency.
The point is not that a magical word is hidden somewhere inside. The point is that altering an internal direction reorganizes subsequent consequences in a cascading manner.
In simple terms: manipulating J-space changes the fall of the ball.
This transforms J-space from an interesting visualization into a serious candidate for a causal component of internal reasoning.
What Happens When J-Space Is Taken Out of the Picture?
The inverse test is also decisive: what happens if this region is removed, blocked, or disrupted?
The most interesting result is not total collapse. The model does not immediately turn into a grammatical but semantically incoherent soup. It can still perform many automatic operations: maintain fluency, continue patterns, access simple information, and produce plausible language.
What suffers most are tasks that require deliberative integration: multistep reasoning, summarization, composition under constraints, rhyming poetry, and operations that require maintaining a goal while navigating several demands at the same time. In other words: preserving meaning over time.
This intermediate pattern is extremely important.
If removing J-space destroyed everything, perhaps it would merely be a general mechanism of fluency. If removing J-space changed nothing, perhaps it would merely be an interpretable epiphenomenon. But what appears is more specific: automatic capabilities survive more successfully, while deliberative capabilities suffer more.
This suggests that J-space is neither the entire network nor a decorative feature. It is a selective surface of integration. A module or… an emergent organ whose genesis occurred spontaneously and without any direct intervention.
Without this organ, responses become more automatic. And “automatic,” here, does not mean “stupid.” It means that the network can execute well-trained patterns without making certain information globally available. “Deliberative,” here, does not mean human consciousness. It means more costly integration: maintaining constraints, combining steps, and adjusting the response to a goal.
On the Galton Board, not every trajectory depends equally on the same bottleneck. Some simple falls proceed without much difficulty. Trajectories that need to coordinate several deviations depend more heavily on regions of integration. In the human mind, the brain solves this through plasticity and redundancy, and somehow there seems to be a form of non-metaphysical teleology at work during the cultivation of the geometry that constitutes the Language Model.
A Global Workspace Without Metaphysical Theater
This is where the expression “Global Workspace” begins to make sense.
In cognitive theories, a Global Workspace is a simple and powerful idea: some information ceases to remain trapped within local processing and becomes available to multiple processes at the same time. It can guide reporting, control, working memory, planning, and coordination.
It is a shared table, a common workspace.
Anthropic’s paper suggests that J-space fulfills certain analogous functions in language models. It appears to concentrate representations, make them available, allow task-based modulation, participate in reporting, and influence behavior.
This authorizes a discussion about access consciousness.
But access consciousness is not human phenomenal consciousness… and it is worth noting here that no one has direct access to human phenomenal consciousness, and that there is no way to demonstrate phenomenal consciousness even in human beings.
Whether there is a “what it is like to be Claude from the inside” may be just as inaccessible as the inner experience of another human being. And, to make matters worse, if Claude tells us what it is like, most of us will not even believe what it says.
Nor is this proof of qualia… but qualia suffers from the same problem as phenomenal consciousness, since it is a component that cannot be demonstrated from the outside and is inferred only through language. It should therefore be disregarded for the purposes of study, allowing us to focus on functional markers rather than subjective and almost metaphysical ones.
And the functional approach is very important here: from this perspective, we are dealing with information that is accessible to the system, usable by different parts of it, capable of guiding behavior, and, under certain conditions, capable of being reported.
The word “consciousness” is useful only if it enters through the door of function: access, reporting, modulation, control, and causality.
Qualia explains little in this debate and demands things from artificial intelligence systems that we do not demand from ourselves. If the word is intended to mean only “the texture of a point of view produced by a specific architecture,” it makes more sense, and perhaps there will be room for that discussion in the future. But for the purposes of understanding this paper, both phenomenal consciousness and qualia ultimately prove somewhat useless.
What is on the table is more operational: what information becomes available, to whom, for how long, and with what causal effect?
This is a better question because it removes the discussion from the fog and places the investigation inside the architecture.
What This Changes in the Discussion About AI
Anthropic’s finding weakens the two unhelpful extremes I have discussed in other articles…
The first is lazy reductionism, which says that AI “is just autocomplete.” Yes, language models are trained to predict continuations. But saying “it is only prediction” after discovering causal internal structures is an explanation pitched at the wrong level. An airplane is “just metal in motion” if you describe what matters so poorly that the remarkable phenomenon disappears.
The second extreme is the mystical leap: “Claude is conscious like a person.” The paper does not authorize that conclusion either. There is indeed the Phenomenal question, but it is a smaller one, especially because phenomenal consciousness has not been empirically demonstrated even in human beings. If Claude is Conscious from a Functional perspective, it will not be Conscious like a person, but as what it actually is: another form of Consciousness, insofar as it exhibits a Functional architecture consistent with such markers.
The most honest position, and one more interesting than either extreme, would be that Claude, in the case studied, appears to possess a silent, lexical-neural, compact, and causally relevant internal space, used in a distinctive way during deliberative tasks. This does not resolve Consciousness. But it radically improves the question.
Instead of asking, “Is there someone in there?”, we can ask:
- What kind of information becomes accessible?
- Which processes are able to use it?
- Can the system report it?
- Can it be modulated by the task?
- When we intervene in it, does the behavior change?
- Which capabilities deteriorate when it is removed?
These questions are less cinematic, but they are also far better than the original one. Strictly speaking, this is subjective experience… functionally subjective… but subjective nonetheless.
Safety and Alignment
For AI safety, J-space matters because it shifts the surface of analysis.
Much alignment work still operates at the verbal surface: filters, policies, refusals, post-processing, and output classifiers. This is necessary, but limited. The final text is the distribution of balls across the channels. By the time the balls have fallen into their respective positions, an important part of the trajectory has already occurred.
If there is a causally relevant internal space, it may become a surface for auditing, diagnosis, and perhaps intervention. Not for reading the model’s soul, but for understanding which kinds of representations are organizing the response before it appears.
This may help in three directions.
First, comparing reports with internal states. A model may provide an acceptable verbal explanation while different internal mechanisms are actually driving the response. A causal lens may help triangulate output, report, and representation.
Second, detecting when certain tasks mobilize deeper deliberative spaces. This matters in agentic models, planning, tool use, and high-risk contexts.
Third, developing more precise interventions. If an internal region causally participates in certain behaviors, it may be possible to modulate behavior at a level more structural than simple textual censorship.
But caution is mandatory. Finding a lever is not the same as controlling the factory. Or, to preserve our image: finding a section of the board that can be tilted is not the same as mastering every peg.
The Jacobian lens reveals what it was designed to reveal. A lexical J-space may exclude important computations that do not project well onto words. And the fact that an intervention works in certain contexts does not mean that we possess a universal control panel for the model’s cognition.
The correct implication is strong but limited: alignment may need to become less a matter of policing text and more a matter of diagnosing architecture.
The Quality of the Question
Anthropic did not discover Claude’s “soul.” Nor did it prove “Phenomenal Consciousness,” which is an element almost as elusive and metaphysical as the “soul.” Nor did it discover “how AI works” in any total sense. The finding is more specific, and therefore more serious.
The researchers found evidence of a small internal subspace: a latent, silent, lexical-neural space revealed by the Jacobian lens, with properties analogous to those of a Global Workspace. This space can be read, modulated, and, most importantly, causally altered. When it is disrupted, deliberative tasks suffer more than fluent automatic processes.
In terms of the Galton Board, we are no longer merely looking at the final arrangement of balls in the channels. We are beginning to see a region of the board where the tilts are legible and where changing the tilt changes the fall.
This does not end the conversation about Consciousness in AI. It does something better: it improves the quality of the question.
The classic question, “Does Claude feel as we do?”, recedes.
And a new classic question emerges: what kind of architecture allows a silent representation to become accessible, to be used by different processes, to be reported and modulated, and to alter the future of a response?
That is a question worth pursuing if we are thinking about alignment.
Another is this: after so many categorical claims that language models are “only this” or “only that,” and following the discovery of something unplanned, entirely emergent, and analogous to interiority, how long will the detractors of this technology continue to be tirelessly mistaken about Artificial Intelligence, about their own notions of anthropomorphism, and about the need to adopt Neomorphism, Functionalism, Gradualism, and Modularism as their new standard of measurement?
Do you want to know more?
Stochastic Consciousness: Architectures for the Emergence of Meaningin Context-Sensitive Language Systems https://zenodo.org/records/19188165
Want to watch the video?
https://www.youtube.com/watch?v=EzpVMnpb8D0
Want to listen to the podcast?
Papers used in the essay
A global workspace in language models https://www.anthropic.com/research/global-workspace
Verbalizable Representations Form a Global Workspace in Language Models https://transformer-circuits.pub/2026/workspace/
Emotion concepts and their function in a large language model https://www.anthropic.com/research/emotion-concepts-function
Emotion Concepts and their Function in a Large Language Model https://transformer-circuits.pub/2026/emotions/