Abstract
We introduce \ours{}, a recurrent Transformer architecture with fixed-size memory that generalizes sliding-window attention while remaining parallelizable during training. \ours{} consists of two coupled models: a prefiller , which leverages full attention\footnote{In practice, we use interleaved full and sliding-window attention for , as this yields stronger performance. The essential requirement is that be more expressive than , with access to the full history.} to produce memory targets , and a decoder , which uses only sliding-window attention and recurrent K/V injection to produce decoder memories for next-token prediction. We train \ours{} with a memory consistency loss that aligns with , allowing inference to use alone. Empirically, \ours{} improves validation loss and downstream pretraining benchmarks over sliding-window and latent recurrent transformer baselines. Moreover, sharing parameters between and reduces parameter memory while preserving most of the gains.