Chris Olah

According to our database¹, Chris Olah authored at least 24 papers between 2015 and 2026.

Collaborative distances:

Dijkstra number² of three.
Erdős number³ of four.

Timeline

Legend:

Book In proceedings Article PhD thesis Dataset Other

Links

On csauthors.net:

Bibliography

2026

Scaling Monosemanticity: Extracting Interpretable Features from Claude 3 Sonnet.

[BibT_eX]

[DOI]

CoRR, May, 2026

Emotion Concepts and their Function in a Large Language Model.

[BibT_eX]

[DOI]

CoRR, April, 2026

When Models Manipulate Manifolds: The Geometry of a Counting Task.

[BibT_eX]

[DOI]

CoRR, January, 2026

2025

Auditing language models for hidden objectives.

[BibT_eX]

[DOI]

CoRR, March, 2025

2023

The Capacity for Moral Self-Correction in Large Language Models.

[BibT_eX]

[DOI]

CoRR, 2023

Discovering Language Model Behaviors with Model-Written Evaluations.

[BibT_eX]

[DOI]

Timothy Telleen-Lawton

Proceedings of the Findings of the Association for Computational Linguistics: ACL 2023, 2023

2022

Discovering Language Model Behaviors with Model-Written Evaluations.

[BibT_eX]

[DOI]

Timothy Telleen-Lawton

CoRR, 2022

Constitutional AI: Harmlessness from AI Feedback.

[BibT_eX]

[DOI]

Timothy Telleen-Lawton

CoRR, 2022

Measuring Progress on Scalable Oversight for Large Language Models.

[BibT_eX]

[DOI]

CoRR, 2022

In-context Learning and Induction Heads.

[BibT_eX]

[DOI]

CoRR, 2022

Toy Models of Superposition.

[BibT_eX]

[DOI]

CoRR, 2022

Red Teaming Language Models to Reduce Harms: Methods, Scaling Behaviors, and Lessons Learned.

[BibT_eX]

[DOI]

CoRR, 2022

Language Models (Mostly) Know What They Know.

[BibT_eX]

[DOI]

CoRR, 2022

Scaling Laws and Interpretability of Learning from Repeated Data.

[BibT_eX]

[DOI]

CoRR, 2022

Training a Helpful and Harmless Assistant with Reinforcement Learning from Human Feedback.

[BibT_eX]

[DOI]

CoRR, 2022

Predictability and Surprise in Large Generative Models.

[BibT_eX]

[DOI]

CoRR, 2022

Predictability and Surprise in Large Generative Models.

[BibT_eX]

[DOI]

Proceedings of the FAccT '22: 2022 ACM Conference on Fairness, Accountability, and Transparency, Seoul, Republic of Korea, June 21, 2022

2021

A General Language Assistant as a Laboratory for Alignment.

[BibT_eX]

[DOI]

CoRR, 2021

2018

Is Generator Conditioning Causally Related to GAN Performance?

[BibT_eX]

[DOI]

Proceedings of the 35th International Conference on Machine Learning, 2018

2017

Conditional Image Synthesis with Auxiliary Classifier GANs.

[BibT_eX]

[DOI]

Augustus Odena

Christopher Olah

Jonathon Shlens

Proceedings of the 34th International Conference on Machine Learning, 2017

Changing Model Behavior at Test-time Using Reinforcement Learning.

[BibT_eX]

[DOI]

Augustus Odena

Dieterich Lawson

Christopher Olah

Proceedings of the 5th International Conference on Learning Representations, 2017

2016

Concrete Problems in AI Safety.

[BibT_eX]

[DOI]

CoRR, 2016

TensorFlow: Large-Scale Machine Learning on Heterogeneous Distributed Systems.

[BibT_eX]

[DOI]

CoRR, 2016

2015

Document Embedding with Paragraph Vectors.

[BibT_eX]

[DOI]

Andrew M. Dai

Christopher Olah

Quoc V. Le

CoRR, 2015

Chris Olah

Timeline

Legend:

Links

On csauthors.net:

Bibliography

Loading...