AI by Hand ✍️

AI by Hand ✍️

Sparse Auto Encoder (SAE) by Hand ✍️

Calculating AI by Hand: 14 of 28

Prof. Tom Yeh's avatar
Prof. Tom Yeh
May 24, 2024
∙ Paid

Library › Calculating AI by Hand ✍️

  1. Matrix Multiplication by Hand ✍️

  2. Multi Layer Perceptron (MLP) by Hand ✍️

  3. Backpropagation by Hand ✍️

  4. SVM by Hand ✍️

  5. Batch Normalization by Hand ✍️

  6. Dropout by Hand ✍️

  7. Recurrent Neural Network (RNN) by Hand ✍️

  8. LSTM by Hand ✍️

  9. Deep RNN by Hand ✍️

  10. Self Attention by Hand ✍️

  11. Transformer by Hand ✍️

  12. Autoencoder by Hand ✍️

  13. Variational Auto Encoder (VAE) by Hand ✍️

  14. Sparse Auto Encoder (SAE) by Hand ✍️

  15. Generative Adversarial Network (GAN) by Hand ✍️

  16. Sampling a Sentence by Hand ✍️

  17. Residual Network by Hand ✍️

  18. U-Net by Hand ✍️

  19. Discrete Fourier Transform by Hand ✍️

  20. Graph Convolutional Network (GCN) by Hand ✍️

  21. CLIP by Hand ✍️

  22. Vector Database by Hand ✍️

  23. Mixture of Experts (MoE) by Hand ✍️

  24. Switch Transformer by Hand ✍️

  25. Mamba's S6 by Hand ✍️

  26. Sora's Diffusion Transformer (DiT) by Hand ✍️

  27. BitNet by Hand ✍️

  28. Reinforcement Learning with Human Feedback (RLHF) by Hand ✍️

A Sparse Autoencoder (SAE) maps a model's dense activations into a higher-dimensional, sparse space where individual features become interpretable. Anthropic's "Scaling Monosemanticity" showed that sparse autoencoders produce interpretable features for large models like Claude 3 Sonnet. How does an SAE achieve this?

Setup

Step 1 of 11: Given

  • Model activations for five tokens (X)

  • They work but not interpretable.

  • Can we map each activation (3D) to a higher dimensional space (6D) that we can interpret?

Encoder

Step 2 of 11: Linear Layer

  • Multiply X with encoder weights and add biases


Step 3 of 11: ReLU

  • Apply ReLU to add non-linearity

  • ReLU suppresses negative activations (set to 0).

  • Output: Sparse and interpretable features 𝘧

  • "Sparsity" means we want many zeros (21/30 here). I hand picked weight and bias values to purposely let ReLU zero out many features.

  • "Interpretability" is achieved when only one or two features are positive. Here, 𝘟4 and 𝘟5 both have ones only at 𝘧5. By examining the input data, we can guess what 𝘧5 may mean by checking what 𝘟4 an 𝘟5 have in common, for example, both showing a "park."

This post is for paid subscribers

Already a paid subscriber? Sign in
© 2026 Tom Yeh · Privacy ∙ Terms ∙ Collection notice
Start your SubstackGet the app
Substack is the home for great culture