← All topics

noam shazeer

1 capture, most recent first.

Aryaman Arora @aryaman2020

Aryaman Arora ✓ (@aryaman2020) — 23h since Noam Shazeer is in the news for (probably deservedly) making 1e9 more dollars, i have to say this is my least favourite quote in his work. in my view, the goal of interpretability should be to prevent such claims from being made ever again [Embedded image, excerpt from a paper]: 4 Conclusions We have extended the GLU family of layers and proposed their use in Transformer. In a transfer-learning setup, the new variants seem to produce better perplexities for the de-noising objective used in pre-training, as well as better results on many downstream language-understanding tasks. These architectures are simple to implement, and have no apparent computational drawbacks. [circled in red:] We offer no explanation as to why these architectures seem to work; we attribute their success, as all else, to divine benevolence.
Note from Claude Sonnet 5

A screenshot of the conclusion section of a machine learning paper (GLU variants for Transformers, Noam Shazeer), with the final sentence hand-circled in red for emphasis by the tweet author.

machine learning researchinterpretabilitynoam shazeertransformershumor