Skip to content

Repository files navigation

Mechanistic Topic Models

Code accompanying the paper Model Directions, Not Words: Mechanistic Topic Models Using Sparse Autoencoders.

Mechanistic topic models define topics over interpretable sparse autoencoder (SAE) features rather than words. This repository contains the model implementations, training and evaluation scripts, and steering experiments used in the paper. This is a preliminary research release.

Setup

The code requires Python 3.9 or later. We use Poetry to install the project and its development dependencies:

git clone https://github.com/blei-lab/mechanistic-topic-models.git
cd mechanistic-topic-models
poetry install
poetry shell

Some workflows use Hugging Face and OpenAI APIs. Create .api_tokens in the repository root (this file is ignored by Git):

{
  "HF_API_TOKEN": "...",
  "HF_WRITE_TOKEN": "...",
  "OPENAI_API_TOKEN": "...",
  "DEEPSEEK_API_TOKEN": "..."
}

HF_WRITE_TOKEN is optional. The other keys are required when the package configuration is imported, even when a particular run does not call the corresponding API.

The Mallet baselines additionally require a local Mallet installation. If its executable is not on PATH, set MALLET_BIN_PATH to its location.

Datasets

20 Newsgroups, AG News, and Yelp Polarity are downloaded automatically. The remaining datasets are linked below, and must be downloaded and placed under raw_data/:

The training code preprocesses the raw files into the corpus format used by the models.

Running

Training and evaluation are controlled by Python configuration files. Edit scripts/configs/train_config.py to select a dataset and model, then run:

python scripts/train.py

Saved models are written to saved_models/. To evaluate them, edit scripts/configs/eval_config.py so that its dataset and run names match the trained models, then run:

python scripts/evaluate.py

Hyperparameter sweeps are configured under scripts/bayesopt/. Steering experiments and their additional setup notes are under src/mechanistic_topic_models/steering_utils/ and scripts/steering_*.py.

Citation

This paper has been accepted for publication in Transactions of the Association for Computational Linguistics (TACL) and is forthcoming. Meanwhile, please cite the arXiv version:

@article{zheng2025model,
  title={Model Directions, Not Words: Mechanistic Topic Models Using Sparse Autoencoders},
  author={Zheng, Carolina and Beltran-Velez, Nicolas and Karlekar, Sweta and Shi, Claudia and Nazaret, Achille and Mallik, Asif and Feder, Amir and Blei, David M.},
  journal={arXiv preprint arXiv:2507.23220},
  year={2025},
  url={https://arxiv.org/abs/2507.23220}
}

License

This project is released under the MIT License. See LICENSE.

About

No description, website, or topics provided.

Resources

Stars

6 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages