← all papers · overview

Theoretical Insights Into Fine-tuning Attention Mechanism: Generalization And Optimization

Abstract

Large Language Models (LLMs), built on Transformer architectures, exhibit remarkable generalization across a wide range of tasks. However, fine-tuning these models for specific tasks remains resource-intensive due to their extensive parameterization. In this paper, we explore two remarkable phenomena related to the attention mechanism during the fine-tuning of LLMs (where , \(\ma

Related papers

Ranked by semantic similarity — how closely each paper's abstract matches this one (100% = near-identical topic).