← arXiv

arXiv·3 min read

New LoRA Skills Should Read but Never Write

The goal is to compose multiple LoRA adapters into a single model without interference or extra inference cost.

Low‑rank adapters let practitioners fine‑tune a large language model per task cheaply, but merging several independently trained adapters is problematic: naïve weight‑space addition causes interference, full retraining on all tasks is expensive, and routing between adapters defeats the purpose of a unified model. The authors set out to resolve this composition dilemma while preserving the cheap per‑task fine‑tuning advantage.

They introduce READ, which rewrites each adapter into a balanced canonical form that exactly preserves its delta and then links adapters with a one‑directional coupling matrix. The coupling allows a newly added skill to read the input subspaces of earlier skills but prevents it from writing into their output subspaces. Only the new skill’s row in the coupling matrix is trainable, and the resulting update is folded back into the base weights, incurring no runtime routing or task‑specific logic.

Evaluated on four benchmark suites and two model families, adding skills sequentially, READ outperforms the strongest baselines built from the same adapters on every suite average, gaining more than twenty points on SuperGLUE and over seven points on the domain suite. Nearly all complete addition sequences finish above every direct baseline, confirming that factor coordinates and coupling direction are the decisive factors for preserving composed skills.

TakeawayREAD composes LoRA adapters via a canonical read‑only coupling, achieving >20‑point SuperGLUE and >7‑point domain suite improvements.

Prodigy briefing — continue on the original for source material, discussion, and updates.

Read the paper ↗