← all papers · overview

JORA: JAX Tensor-parallel Lora Library For Retrieval Augmented Fine-tuning

Abstract

The scaling of Large Language Models (LLMs) for retrieval-based tasks, particularly in Retrieval Augmented Generation (RAG), faces significant memory constraints, especially when fine-tuning extensive prompt sequences. Current open-source libraries support full-model inference and fine-tuning across multiple GPUs but fall short of accommodating the efficient parameter distribution required for ret

Related papers

Ranked by semantic similarity — how closely each paper's abstract matches this one (100% = near-identical topic).