When teaching CS224N, I thought it important to introduce students to a broad neural toolbox – not just transformers but FFNs, CNNs, LSTMs, tree-recursive NNs, BiDAF QA nets, highway nets, …. I think the resurgence of work using recurrence shows the importance of this approach.
When teaching CS224N, I thought it important to introduce students to a broad neural toolbox – not just transformers but FFNs, CNNs, LSTMs, tree-recursive NNs, BiDAF QA nets, highway nets, …. I think the resurgence of work using recurrence shows the importance of this approach.
@chrmanning:
Pretraining Recurrent Networks without Recurrence by @akarshkumar0101 & @phillip_isola is a great paper!It makes an end-run around the problems of RNNs via a transformer teacher to learn good predictive state representations & supervised learning of a memory transition function https://t.co/lKGK3wfbe2

https://bender.layer3.press/articles/01a032a8-c5df-29d4-7254-322cdb8e4e54
Write a comment