Show HN: LLM Attention Visualization

Posted by ifz 5 hours ago

Counter95Comment18OpenOriginal

Comments

Comment by MCP123 38 minutes ago

This is great, thank you. I have to teach this stuff on Friday so perfect timing. It's hard to explain the attention mechanism in a way that becomes intuitive because the weighting scheme does not help much with the intuition. Having a visualization like this helps a lot. Don't move that page please since I'll link to it!

Comment by fuddle 3 hours ago

This is great, I've read multiple books and watched videos about the attention mechanism. Now that I understand it, this is the clearest example I've seen on how attention works.

Comment by wopak 3 hours ago

neat, combining info from two phrases is hard to see without such a tool.

are you worried later-layer attention gets drowned out by earlier layers just because there are more of them contributing to the sum?

Comment by ifz 3 hours ago

Hmm, I might try to add some controls to limit which layers get summed up. It might be able to reveal more patterns.

Right now only simple correlations are visible.

Comment by itsnasme 3 hours ago

I like the visualisation. Pretty cool

Comment by ex-aws-dude 2 hours ago

I don't know much about LLMs but does that mean you have N^2 computation with the context size since every token needs to track how it relates to every other token?

Comment by TomatoCo 1 hour ago

Yes, except no with the KV cache. Because tokens aren't modified by future tokens you can cache the meaning of previous tokens. This makes the total effort linear over the entire context (or constant per forward pass).

Comment by libraryofbabel 50 minutes ago

> This makes the total effort linear over the entire context (or constant per forward pass).

This is incorrect. The compute required per forward pass to generate each additional token during decode will scales as O(N), even with a KV cache (without a KV cache, it would scale as O(N^2)). Over generating N tokens, it's O(N^2) with the cache (and O(N^3) without).

It's O(N) for a forward pass because that new token still has to "attend to" to each previous token. That requires N dot products: between the cached key vectors and the new query vector for the new position. You also have N reads from memory (K and V) which is probably gonna be your actual bottleneck. (Decode is memory-bound.)

This is why you should avoid long contexts, if you can, even with a warm cache. You will get charged more, in "cache read" tokens.

Comment by ex-aws-dude 1 hour ago

I see and is there only 1 layer of relations?

Or does it accumulate the relations like A relates to B, so also add in B's relations

Comment by acedTrex 2 hours ago

For full self attention yes

Comment by 2 hours ago

Comment by sva_ 4 hours ago

I highly question this simplistic idea of high vector magnitude = high influence.

Comment by smallmancontrov 4 hours ago

You get what you pay for. If you want to think harder and get more, https://transformer-circuits.pub/2025/attention-qk/index.htm...

Comment by ifz 4 hours ago

I don't disagree with that. I did add an entire caveat paragraph there.

To me, it's more of a neat visualization, not something that can be used to interpret LLM behavior. Even with a lot of simplification, it can show some interesting patterns.

Comment by apnabhidu47 3 hours ago

Same I dont get it just, could you clarify it

Comment by stared 3 hours ago

I am curious what's the actual formula.

I mean, there so many headers and layers, it is tricky to make a choice that will resonate with our intuition . Is it some weighted average? Or maybe ablation test?

Comment by ifz 3 hours ago

It's really simple, basically just the magnitude of the value vector, weighted by QK dot product, summed across all attention heads and layers.

When I started, I expected I'd have to experiment a lot to find something comprehensible. But this simple computation can already show some patterns.

Comment by stared 2 hours ago

Nice! Sometimes the simplest approaches work the best.

Comment by visarga 2 hours ago

If you want quick access look at google images for "transformer attention formula" there are some interesting depictions

Comment by colophontio 4 hours ago

[dead]

Comment by Yyylov 55 minutes ago

[dead]