Skip to main content
Softmax function

Softmax function

Search complete. 5 mentions across 5 episodes found for "Softmax function".

Sep 21, 2026

Type Three AudioNARRATOR
7:08
In Anthropic's mathematical framework for transformer circuits, they introduced QK and OV circuits.
Type Three AudioNARRATOR
7:15
They stop there because you can't compose other parts of the model together due to non-linearity such as Softmax or GELU.
Type Three AudioNARRATOR
7:22
With a tensor transformer, you can compose everything with everything, such as embedding.
Type Three AudioNARRATOR
7:27
Right arrow.
Arshavir BlackwellHOST
10:37
Attention has that structure, and not merely by analogy.
Arshavir BlackwellHOST
10:42
Every earlier word is scored, and the softmax forces the scores into a distribution summing to 1, so the candidates are in literal zero-sum competition for a fixed budget, exactly as q's are in Bates' model.
Arshavir BlackwellHOST
11:00
What the query and key lenses learn is the functional equivalent of q-strength.
Arshavir BlackwellHOST
11:06
How much a given kind of match counts, tuned by training toward whatever predicts well.
Bill DallyGUEST
7:02
You know, the application drives certain requirements, right? You have gems, you're gonna need some matrix multiply units.
Bill DallyGUEST
7:08
You have, you know, softmax and, and norms, you're gonna need some vector units and, and things that can do transcendental functions.
Bill DallyGUEST
7:14
You need a certain amount of memory capacity and a certain amount of memory bandwidth.
Bill DallyGUEST
7:17
You need a certain amount of communication bandwidth.
speaker_4HOST
15:54
Which is totally realistic.
speaker_5HOST
15:55
To get that, we cap the network off with a special function called softmax.
speaker_5HOST
15:59
Softmax.
speaker_5HOST
16:00
It looks at all the final scores and mathematically forces them to add up to exactly 100%.
speaker_5HOST
16:05
It gives us a clear, distinct distribution of probabilities we can actually read.
Artificial IntelligenceNARRATOR
20:58
Jürgen Schmidhuber published IDEN 1992 as a scheme for programming fast weights, a quarter of a century before attention existed, and a 2021 paper proved the two are formally the same idea in different notation.
Artificial IntelligenceNARRATOR
21:11
Softmax's attention is not the original design.
Artificial IntelligenceNARRATOR
21:14
It is the deviation, nor is it unbuilt, which is the part he finds funny.
Artificial IntelligenceNARRATOR
21:18
Go back to the model cards from the last chapter and read the architectures instead of the licenses.

We value your privacy

We use cookies to understand how you use our platform and to improve your experience. Click “Accept All” to consent, or “Decline non-essential” to opt out of non-essential cookies. Read our Privacy Policy.