← all papers · overview

Fitv2: Scalable And Improved Flexible Vision Transformer For Diffusion Model

Abstract

\textit\{Nature is infinitely resolution-free\}. In the context of this reality, existing diffusion models, such as Diffusion Transformers, often face challenges when processing image resolutions outside of their trained domain. To address this limitation, we conceptualize images as sequences of tokens with dynamic sizes, rather than traditional methods that perceive images as fixed-resolution grids. This perspective enables a flexible training strategy that seamlessly accommodates various aspect ratios during both training and inference, thus promoting resolution generalization and eliminating biases introduced by image cropping. On this basis, we present the \textbf\{Flexible Vision Transformer\} (FiT), a transformer architecture specifically designed for generating images with \textit\{unrestricted resolutions and aspect ratios\}. We further upgrade the FiT to FiTv2 with several innovative designs, includingthe Query-Key vector normalization, the AdaLN-LoRA module, a rectified flow

Related papers

Ranked by semantic similarity — how closely each paper's abstract matches this one (100% = near-identical topic).