Tristan Hume

According to our database¹, Tristan Hume authored at least 13 papers between 2022 and 2026.

Collaborative distances:

Dijkstra number² of four.
Erdős number³ of four.

Timeline

Legend:

Book In proceedings Article PhD thesis Dataset Other

Links

On csauthors.net:

Bibliography

2026

Scaling Monosemanticity: Extracting Interpretable Features from Claude 3 Sonnet.

[BibT_eX]

[DOI]

CoRR, May, 2026

2023

Specific versus General Principles for Constitutional AI.

[BibT_eX]

[DOI]

CoRR, 2023

Measuring Faithfulness in Chain-of-Thought Reasoning.

[BibT_eX]

[DOI]

CoRR, 2023

The Capacity for Moral Self-Correction in Large Language Models.

[BibT_eX]

[DOI]

CoRR, 2023

Discovering Language Model Behaviors with Model-Written Evaluations.

[BibT_eX]

[DOI]

Timothy Telleen-Lawton

Proceedings of the Findings of the Association for Computational Linguistics: ACL 2023, 2023

2022

Discovering Language Model Behaviors with Model-Written Evaluations.

[BibT_eX]

[DOI]

Timothy Telleen-Lawton

CoRR, 2022

Constitutional AI: Harmlessness from AI Feedback.

[BibT_eX]

[DOI]

Timothy Telleen-Lawton

CoRR, 2022

Measuring Progress on Scalable Oversight for Large Language Models.

[BibT_eX]

[DOI]

CoRR, 2022

Toy Models of Superposition.

[BibT_eX]

[DOI]

CoRR, 2022

Red Teaming Language Models to Reduce Harms: Methods, Scaling Behaviors, and Lessons Learned.

[BibT_eX]

[DOI]

CoRR, 2022

Language Models (Mostly) Know What They Know.

[BibT_eX]

[DOI]

CoRR, 2022

Scaling Laws and Interpretability of Learning from Repeated Data.

[BibT_eX]

[DOI]

CoRR, 2022

Training a Helpful and Harmless Assistant with Reinforcement Learning from Human Feedback.

[BibT_eX]

[DOI]

CoRR, 2022

Tristan Hume

Timeline

Legend:

Links

On csauthors.net:

Bibliography

Loading...