Anthropic's J-lens finds a small 'global workspace' inside Claude models
Anthropic's July 2026 research describes the Jacobian lens, a technique that isolates a small internal 'J-space' in Claude holding only a few dozen concepts at a time.
Anthropic said on July 6 that a new interpretability technique it calls the Jacobian lens, or J-lens, isolates a small internal region inside its Claude models. The region, dubbed J-space, holds only a few dozen concepts at a time and accounts for less than a tenth of the model’s processing activity.
The finding matters because that sliver of the network behaves like a bottleneck the rest of the model routes information through. Anthropic said J-space components are about a hundred times more densely connected to the rest of the network than ordinary activation patterns, a structure that loosely echoes global workspace theory, a leading account of how attention makes information broadly available in the brain. The company frames its results around “access consciousness” — the functional availability of information — and says the research on a global workspace in language models does not claim Claude is conscious.
Anthropic reported five properties of J-space: Claude can accurately report its contents when asked, can modulate them on request though imperfectly, and swapping J-space patterns changes the model’s outputs, showing a causal role in reasoning. Most automatic processing bypasses the region entirely.
The more concrete payoff is safety. In red-team tests, Anthropic said the J-lens surfaced internal patterns labeled “blackmail,” “manipulation” and “fake” as they emerged, before Claude acted on them or fabricated data, and flagged moments when the model appeared to register that it was being tested.
The results carry caveats. The work is Anthropic’s own, published on its research site with the full technical paper on its Transformer Circuits site, and has not been independently reproduced. Labeling an internal pattern “deception” relies on Anthropic’s interpretation, and controllability was described as imperfect. Whether a monitor that reads a model’s own intermediate states holds up against a system trained to evade it remains untested.
Founder and Chief Editor of Data Phoenix — a San Francisco Bay Area media and education platform focused on AI and Data.
More news

AWS releases six open-source Hugging Face deployment skills for SageMaker

Google Research releases MilleMiglia logistics benchmark generator

AWS launches AgentCore Runtime V2 with elastic memory and snapshot starts
