← all papers · overview

Building On Efficient Foundations: Effectively Training Llms With Structured Feedforward Layers

Abstract

State-of-the-art results in large language models (LLMs) often rely on scale, which becomes computationally expensive. This has sparked a research agenda to reduce these models' parameter counts and computational costs without significantly impacting their performance. Our study focuses on transformer-based LLMs, specifically targeting the computationally intensive feedforward networks (FFNs), whi

Related papers

Ranked by semantic similarity — how closely each paper's abstract matches this one (100% = near-identical topic).