Reels · 9 of 40
Reel · 46 s
Transformer explained: it landed below three simpler models, and among the top candidates the grid ranked them backwards
In 46 seconds, narrated: why the Transformer landed below three simpler models; the 9,633 weights the rule picked and the 0.007 of AUC that separate it from boosting; how the ranking of eleven architectures flipped between validation and the test period; what happened when the attention was removed; and the layer that shields it from a unit change by leaving one column mute. Muted by default: turn the sound on in the controls.
- Length
- 46 s
- Published
This reel sums up Transformer explained: it landed below three simpler models, and among the top candidates the grid ranked them backwards, where the method, the tables and what did not work are.