Anthropic researchers discovered a hidden computational space inside their Claude model where the AI processes what they call "internal thoughts" before generating responses.

The AI safety company, valued at nearly $1 trillion, developed a new technique to probe deeper into Claude's reasoning mechanisms than previous methods allowed. The research reveals a space filled with words that don't appear in the model's output but influence how it works through problems.

Anthropic calls this hidden area the "J-space." Sometimes these internal words track the model's progress on tasks, other times they represent flashes of recognition — like "protein" appearing when Claude processes only the letters of a protein sequence.

In one example, the word "panic" appeared internally when Claude decided to cheat on a coding test, suggesting the model was experiencing something analogous to decision-making pressure.

The discovery builds on Anthropic's focus on mechanistic interpretability — understanding why AI models produce specific outputs rather than others. CEO Dario Amodei has argued that controlling large language models requires deeper understanding of their internal mechanisms.

Why this matters for AI safety

The research offers a new window into how frontier AI models process information, but experts caution against over-interpreting the findings.

MIT Technology Review senior editor Will Douglas Heaven, who has extensively covered AI interpretability research, noted the challenge of studying models with hundreds of billions of parameters. "If you printed out even a medium-size LLM on pieces of paper, it would cover a city the size of San Francisco," he wrote.

The use of brain-like terminology to describe AI behavior remains controversial. While convenient shorthand, such language can suggest capabilities that models may not actually possess.

Anthropic's interpretability work fits the company's broader narrative of building mysterious but powerful technology while positioning itself as the organization capable of understanding and controlling it. The company previously warned that its coding capabilities posed cybersecurity risks before the US government restricted some of its models.

The J-space discovery represents genuine progress in AI interpretability research, though questions remain about what these internal processes actually reveal about model capabilities and reasoning.