← all papers · overview

BTLM-3B-8K: 7B Parameter Performance In A 3B Parameter Model

Abstract

We introduce the Bittensor Language Model, called "BTLM-3B-8K", a new state-of-the-art 3 billion parameter open-source language model. BTLM-3B-8K was trained on 627B tokens from the SlimPajama dataset with a mixture of 2,048 and 8,192 context lengths. BTLM-3B-8K outperforms all existing 3B parameter models by 2-5.5% across downstream tasks. BTLM-3B-8K is even competitive with some 7B parameter mod

Related papers

Ranked by semantic similarity — how closely each paper's abstract matches this one (100% = near-identical topic).