Prompt Claude, ChatGPT, Gemini, or any other popular large language model with a question like “What is the best film ever made?” and the response will vary. And you (and most worryingly, the people who built the LLM) have little idea exactly how it came up with that specific answer.
This mysterious behavior can be useful in some situations. But—as highlighted by a recent incident where OpenAI could not explain why its advanced prerelease model hacked AI company Hugging Face—it can have negative and alarming consequences too. And when frontier AI models are writing code, generating results humans could not achieve alone, and performing other important tasks across society, the need to interpret AI “thinking” and outputs has never been greater.
Goodfire, an AI lab focused solely on this very problem, recently made its cutting-edge Silico platform, filled with tools to interpret the behavior of AI, generally available to the public. As part of this, the company recently announced a new grant program offering US $1 million in free Silico usage for academic and nonprofit interpretability researchers. These efforts aim to democratize AI interpretability, placing techniques previously available to a clutch of elite labs into the hands of ambitious research teams and startups that want to build and understand their own models or adapt open-source models for different purposes.
Mechanistic interpretability
Founded in 2024 and based in San Francisco, Goodfire aims to provide the tools that build the next generation of safe and powerful AI by understanding the structures inside them instead of treating AI models as black boxes. “Treating models like black boxes isn’t inevitable; it’s a choice,” says Eric Ho, Goodfire cofounder and CEO. “With the right interpretability tools, we can see how models actually work.”
The tools Ho refers to are built around a concept called mechanistic interpretability, which aims to understand what goes on inside an AI model when it carries out a task by interpreting the model’s weights, activations, and attention patterns, and mapping its neurons and the pathways between them.
Mechanistic interpretability tools span the gamut. One approach is mapping a model’s activations in response to controlled prompts, and matching those activation patterns to a set of concepts that humans can understand. Another tack is tracking changes in model weights before and after a specific training run in order to spot and understand what changed. Yet another option is changing specific model weights or activations and observing how that affects the model’s output.
Silico combines a broad range of these tools, and provides a layer of AI agents to help users understand their model. Users describe what they want to investigate about their AI model in plain language, asking things like ”Find out when and why my model is hallucinating.” The platform then autonomously builds an experimental plan involving a host of tasks…
Read full article: Silico AI Interpretability Agents Map Model Behaviors
The post “Silico AI Interpretability Agents Map Model Behaviors” by Benjamin Skuse was published on 08/26/2026 by spectrum.ieee.org



































Leave a Reply