17°

Reels · 9 of 40

46 s

Reel · 46 s

Transformer explained: it landed below three simpler models, and among the top candidates the grid ranked them backwards

In 46 seconds, narrated: why the Transformer landed below three simpler models; the 9,633 weights the rule picked and the 0.007 of AUC that separate it from boosting; how the ranking of eleven architectures flipped between validation and the test period; what happened when the attention was removed; and the layer that shields it from a unit change by leaving one column mute. Muted by default: turn the sound on in the controls.

Length
46 s
Published

Machine learningAlgoritmosPythonDatos

This reel sums up Transformer explained: it landed below three simpler models, and among the top candidates the grid ranked them backwards, where the method, the tables and what did not work are.

Read the articleOpen in the viewer