gpt2-gsm8k-cot / README.md
connordilgren's picture
Add model card
19adef4 verified
|
Raw History Blame Contribute Delete
1.95 kB
metadata
license: mit
base_model: openai-community/gpt2
language:
  - en
tags:
  - latent-reasoning
  - interpretability
  - reasoning
  - cot
  - gsm8k

CoT · gpt2 · GSM8k-Aug

This is the CoT checkpoint trained on GSM8k-Aug with base model openai-community/gpt2, from the paper Are Latent Reasoning Models Easily Interpretable? (Dilgren & Wiegreffe, 2026).

Files

This repository contains a single raw PyTorch checkpoint, checkpoint_25 — the state dict as saved by the training framework. It is not a from_pretrained-style model; it is loaded by the paper's evaluation code, which builds the base model and applies this checkpoint.

Usage

The evaluation code in the repository loads this checkpoint from the local path configured in model_paths.yaml. Download it to the expected location with:

hf download connordilgren/gpt2-gsm8k-cot checkpoint_25 --local-dir checkpoints/gsm-cot

This places the file at checkpoints/gsm-cot/checkpoint_25, which is the path referenced for this model (gpt2 → gsm8k → cot) in model_paths.yaml. See the repository README for full setup and evaluation instructions.

Citation

@misc{dilgren2026latentreasoningmodelseasily,
      title={Are Latent Reasoning Models Easily Interpretable?},
      author={Connor Dilgren and Sarah Wiegreffe},
      year={2026},
      eprint={2604.04902},
      archivePrefix={arXiv},
      primaryClass={cs.LG},
      url={https://arxiv.org/abs/2604.04902},
}