Can Rager

According to our database¹, Can Rager authored at least 16 papers between 2023 and 2025.

Collaborative distances:

Dijkstra number² of four.
Erdős number³ of four.

Timeline

Legend:

Book

In proceedings

Article

PhD thesis

Dataset

Other

Links

On csauthors.net:

Bibliography

2025

Priors in Time: Missing Inductive Biases for Language Model Interpretability.

[BibT_eX]

[DOI]

Ekdeep Singh Lubana

Can Rager

Sai Sumedh R. Hindupur

CoRR, November, 2025

Automatically Finding Rule-Based Neurons in OthelloGPT.

[BibT_eX]

[DOI]

Aditya Singh

Zihang Wen

Srujananjali Medicherla

Adam Karvonen

Can Rager

CoRR, November, 2025

Discovering Forbidden Topics in Language Models.

[BibT_eX]

[DOI]

CoRR, May, 2025

SAEBench: A Comprehensive Benchmark for Sparse Autoencoders in Language Model Interpretability.

[BibT_eX]

[DOI]

CoRR, March, 2025

Sparse Feature Circuits: Discovering and Editing Interpretable Causal Graphs in Language Models.

[BibT_eX]

[DOI]

Proceedings of the Thirteenth International Conference on Learning Representations, 2025

NNsight and NDIF: Democratizing Access to Open-Weight Foundation Model Internals.

[BibT_eX]

[DOI]

Proceedings of the Thirteenth International Conference on Learning Representations, 2025

2024

Evaluating Sparse Autoencoders on Targeted Concept Erasure Tasks.

[BibT_eX]

[DOI]

CoRR, 2024

The Quest for the Right Mediator: A History, Survey, and Theoretical Grounding of Causal Interpretability.

[BibT_eX]

[DOI]

Aruna Sankaranarayanan

CoRR, 2024

NNsight and NDIF: Democratizing Access to Foundation Model Internals.

[BibT_eX]

[DOI]

CoRR, 2024

Measuring Progress in Dictionary Learning for Language Model Interpretability with Board Game Models.

[BibT_eX]

[DOI]

Claudio Mayrink Verdun

David Bau

Samuel Marks

Proceedings of the Advances in Neural Information Processing Systems 38: Annual Conference on Neural Information Processing Systems 2024, 2024

2023

Structured World Representations in Maze-Solving Transformers.

[BibT_eX]

[DOI]

Michael I. Ivanitskiy

Cecilia G. Diniz Behn

Katsumi Inoue

Samy Wu Fung

CoRR, 2023

Attribution Patching Outperforms Automated Circuit Discovery.

[BibT_eX]

[DOI]

Aaquib Syed

Can Rager

Arthur Conmy

CoRR, 2023

An Adversarial Example for Direct Logit Attribution: Memory Management in gelu-4l.

[BibT_eX]

[DOI]

CoRR, 2023

A Configurable Library for Generating and Manipulating Maze Datasets.

[BibT_eX]

[DOI]

Michael Igorevich Ivanitskiy

Cecilia G. Diniz Behn

Samy Wu Fung

CoRR, 2023

Safety of self-assembled neuromorphic hardware.

[BibT_eX]

[DOI]

Can Rager

Kyle Webster

CoRR, 2023

Linearly Structured World Representations in Maze-Solving Transformers.

[BibT_eX]

[DOI]

Michael I. Ivanitskiy

Cecilia G. Diniz Behn

Katsumi Inoue

Samy Wu Fung

Proceedings of UniReps: the First Workshop on Unifying Representations in Neural Models, 2023

Can Rager

Timeline

Legend:

Links

On csauthors.net:

Bibliography

Loading...