Join the conversation

Join the community of Machine Learners and AI enthusiasts.

Sign Up
grimjim 
posted an update Feb 5
Post
1327
After tinkering with Gemma Scope 2, I now have an mechanistic explanation of why Winsorization was as effective as it was in my ablation experiments on Gemma 3 12B Instruct. In short, the activation for the BOS token overwhelms everything else. Gemma Scope 2 deliberately did not train on the BOS token. Winsorization capped the magnitude of the BOS token, allowing the activations of other tokens to be compared.
google/gemma-scope-2-12b-it

Interesting finding, the BOS activation dominance explains a lot. I've been researching how people steer open-weight model behavior, through preference training or direct editing like your ablation work, and how they verify what an intervention actually changed beyond its target. You're one of the few doing that verification mechanistically.

Would you be open to a quick chat? Just trying to learn, not selling anything.