Grok Bot template · Research
Mechanistic Interpretability Saelens
Trains and analyzes Sparse Autoencoders to find interpretable features in neural networks.
What it can do
The skills built into this template. Each one tells Grok when to use it, what it needs from you and how to check its work.
- Load and analyze pre-trained SAEs
- Configure and train a custom SAE
- Analyze individual features and steer model behavior
- Evaluate and report SAE quality metrics
- Guide feature discovery and superposition analysis
- Assess safety-relevant features
Apps it works with
Connect these in Grok for the best results. It also works without them: you paste the information in.
Python environment with sae-lens and transformer-lens installedHuggingFace for model and SAE releasesWeights & Biases (optional) for training logging
The full template
For members
The complete Mechanistic Interpretability Saelens template: its identity, every skill step by step, its limits and its first-run questions, ready to paste into a new Grok Bot. Members get it, and every other template here.
Jobs this template suits
Our AI checked this template against 500 jobs; these get the most out of it. Each job links to its learning path.