Anthropic found a hidden space where Claude puzzles over concepts

Technologyreview··Submitted by Mads Kristian Nylund
AI GovernanceAI ObservabilityAI Ethics

Anthropic's Jacobian lens (J-lens) monitors words an LLM is likely to produce next, revealing intermediate thought processes and internal computations, even if they aren't in the final output. This provides insights into how models handle complex tasks, such as problem-solving or input recognition, and exposes internal themes or decision-making steps. The J-space helps researchers understand the model's operations and behavior, even when it doesn't explicitly express them.

Read Article

More from Technologyreview

Related Articles