Publications

A running list of papers I'm reading in grad school — what I've finished, what I'm working through, and what's next. Notes get linked here as they're written.

Reading 1

Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N. Gomez, Lukasz Kaiser, Illia Polosukhin

NeurIPS · 2017

提出自注意力机制,完全抛弃循环与卷积结构,是现代大语言模型架构的起点。

#Transformer#注意力机制#序列建模
Notes coming soon
BibTeX
@inproceedings{vaswani2017attention,
  title={Attention Is All You Need},
  author={Vaswani, Ashish and Shazeer, Noam and Parmar, Niki and Uszkoreit, Jakob and Jones, Llion and Gomez, Aidan N. and Kaiser, Lukasz and Polosukhin, Illia},
  booktitle={NeurIPS},
  year={2017},
  url={https://arxiv.org/abs/1706.03762}
}

To Read 1

Tom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, others

NeurIPS · 2020

把模型规模推到 175B 并展示 few-shot 能力,计划与 in-context learning 的后续工作对照阅读。

#GPT-3#上下文学习#大规模语言模型
Notes coming soon
BibTeX
@inproceedings{brown2020language,
  title={Language Models are Few-Shot Learners},
  author={Brown, Tom B. and Mann, Benjamin and Ryder, Nick and Subbiah, Melanie and Kaplan, Jared and Dhariwal, Prafulla and Neelakantan, Arvind and Shyam, Pranav and Sastry, Girish and Askell, Amanda and others},
  booktitle={NeurIPS},
  year={2020},
  url={https://arxiv.org/abs/2005.14165}
}