Prabhat TiwariJournal

Reading attention maps like margin notes

A way of looking at model internals that borrows from how we annotate books: lightly, and with questions.

Prabhat Tiwari2 min read

Attention maps are one of the more approachable ways to look inside a language model. They show, roughly, which parts of the input a model weighed while producing an output.

They are easy to over-read. A gentler approach is to treat them the way we treat notes in the margins of a book.

Traces of attention

Margin notes are not the meaning of a text. They are traces of a reader’s attention: what caught the eye, what raised a question. Attention maps can be read in the same spirit.

“Read attention maps lightly, and with questions.”

Prompts, not explanations

Used this way, they become prompts for further investigation rather than explanations in themselves. A surprising pattern is a reason to test something, not a conclusion to report.

Reading with curiosity

Interpretability is still a young field. Approaching its tools with curiosity and a little humility keeps the questions open long enough to learn something from them.

Latest writing