Introduction to Mechanistic Interpretability
Fetch error
Hmmm there seems to be a problem fetching this series right now. Last successful fetch was on January 02, 2025 12:05 ()
What now? This series will be checked again in the next hour. If you believe it should be working, please verify the publisher's feed link below is valid and includes actual episode links. You can contact support to request the feed be immediately fetched.
Manage episode 458945499 series 3498845
Our introduction introduces common mech interp concepts, to prepare you for the rest of this session's resources.
Original text: https://aisafetyfundamentals.com/blog/introduction-to-mechanistic-interpretability/
Author(s): Sarah Hastings-Woodhouse
A podcast by BlueDot Impact.
Learn more on the AI Safety Fundamentals website.
فصول
1. Introduction to Mechanistic Interpretability (00:00:00)
2. Why might mechanistic interpretability be useful? (00:01:16)
3. Looking inside neural networks (00:03:34)
4. What makes mechanistic interpretability hard? (00:06:33)
5. Addressing polysemanticity (00:08:34)
85 حلقات