Transformer
loadingβ¦
loadingβ¦
Transformer is one of the most active areas in Awesome Large Language Models β 40 papers in this collection. A strong starting point is "Hybrid Linear Attention Done Right: Efficient Distillation and Effective Architectures for Extremely Long Contexts".