← all papers · overview

Mabvit -- Modified Attention Block Enhances Vision Transformers

Abstract

Recent studies have demonstrated the effectiveness of Gated Linear Units (GLU) in enhancing transformer models, particularly in Large Language Models (LLMs). Additionally, utilizing a parallel configuration within each Transformer block rather than the conventional serialized method has been revealed to accelerate the training of LLMs without significantly impacting performance. However, when the

Related papers

Ranked by semantic similarity — how closely each paper's abstract matches this one (100% = near-identical topic).