Code accompanying the paper Model Directions, Not Words: Mechanistic Topic Models Using Sparse Autoencoders.
Mechanistic topic models define topics over interpretable sparse autoencoder (SAE) features rather than words. This repository contains the model implementations, training and evaluation scripts, and steering experiments used in the paper. This is a preliminary research release.
The code requires Python 3.9 or later. We use Poetry to install the project and its development dependencies:
git clone https://github.com/blei-lab/mechanistic-topic-models.git
cd mechanistic-topic-models
poetry install
poetry shellSome workflows use Hugging Face and OpenAI APIs. Create .api_tokens in the
repository root (this file is ignored by Git):
{
"HF_API_TOKEN": "...",
"HF_WRITE_TOKEN": "...",
"OPENAI_API_TOKEN": "...",
"DEEPSEEK_API_TOKEN": "..."
}HF_WRITE_TOKEN is optional. The other keys are required when the
package configuration is imported, even when a particular run does not call the
corresponding API.
The Mallet baselines additionally require a local Mallet installation. If its
executable is not on PATH, set MALLET_BIN_PATH to its location.
20 Newsgroups, AG News, and Yelp Polarity are downloaded automatically. The
remaining datasets are linked below, and must be downloaded and placed under raw_data/:
- Bills: ahoho/topics repository, placed in
raw_data/bills/ - Wiki: ahoho/topics repository, placed in
raw_data/wiki/ - GoEmotions: Google Research repository, placed in
raw_data/goemotions/ - PoemSum: PoemSum repository, placed in
raw_data/poemsum/ - WritingPrompts: Kaggle dataset, placed in
raw_data/writing_prompts/
The training code preprocesses the raw files into the corpus format used by the models.
Training and evaluation are controlled by Python configuration files. Edit
scripts/configs/train_config.py to select a dataset and model, then run:
python scripts/train.pySaved models are written to saved_models/. To evaluate them, edit
scripts/configs/eval_config.py so that its dataset and run names match the
trained models, then run:
python scripts/evaluate.pyHyperparameter sweeps are configured under scripts/bayesopt/. Steering
experiments and their additional setup notes are under
src/mechanistic_topic_models/steering_utils/ and scripts/steering_*.py.
This paper has been accepted for publication in Transactions of the Association for Computational Linguistics (TACL) and is forthcoming. Meanwhile, please cite the arXiv version:
@article{zheng2025model,
title={Model Directions, Not Words: Mechanistic Topic Models Using Sparse Autoencoders},
author={Zheng, Carolina and Beltran-Velez, Nicolas and Karlekar, Sweta and Shi, Claudia and Nazaret, Achille and Mallik, Asif and Feder, Amir and Blei, David M.},
journal={arXiv preprint arXiv:2507.23220},
year={2025},
url={https://arxiv.org/abs/2507.23220}
}This project is released under the MIT License. See LICENSE.