← all papers · overview

How Large Language Models Get Stuck: Early Structure With Persistent Errors

Abstract

Linguistic insights may help make Large Language Model (LLM) training more efficient. We trained Meta's OPT model on the 100M word BabyLM dataset, and evaluated it on the BLiMP benchmark, which consists of 67 classes, each defined by sentence pairs that differ in a targeted syntactic or semantic rule violation. We tested the model's preference for grammatical over ungrammatical sentences across tr

Related papers

Ranked by semantic similarity — how closely each paper's abstract matches this one (100% = near-identical topic).