Skip to main content
Cloud & AI Hub
Browse
Glossary AI Directory Playgrounds Models Prompts Explainers Strategy Matrix Benchmark Decoder

LoRA Adapter

A trainable low-rank decomposition matrix injected into transformer attention layers during model tuning.

Last reviewed: July 25, 2026

A LoRA (Low-Rank Adaptation) adapter is a small set of trainable parameters injected into a pretrained model’s layers to fine-tune its behavior, without updating the model’s original weights at all. It’s the most widely used technique for parameter-efficient fine-tuning (PEFT), and it made fine-tuning large models practical on far more modest hardware than full fine-tuning requires.

How It Works

Full fine-tuning updates every one of a model’s billions of parameters, requiring enough GPU memory to store gradients and optimizer states for the entire model — often several times the model’s base memory footprint. LoRA instead freezes all of the original weights and adds small trainable low-rank matrices alongside specific layers, typically the attention projection matrices. The mathematical insight is that the change needed to adapt a model to a new task can often be well approximated by a much lower-dimensional update than the full weight matrix, even though the weight matrix itself is high-dimensional — so training just this low-rank update captures most of the benefit of full fine-tuning at a small fraction of the parameter count and memory cost.

Practical Benefits

Because the base model weights never change, a single pretrained model can serve many different LoRA adapters for different tasks or customers, swapping in the relevant adapter at inference time rather than storing multiple full copies of the fine-tuned model. LoRA adapters are typically a few megabytes to a few hundred megabytes, compared to gigabytes for a full model checkpoint, making them cheap to store, version, and distribute.

QLoRA

A related and now very common variant, QLoRA, combines LoRA with 4-bit quantization of the frozen base model weights, further reducing the GPU memory needed to fine-tune large models — enabling fine-tuning of models with tens of billions of parameters on a single consumer or prosumer GPU.

Choosing the Rank Hyperparameter

LoRA’s most important hyperparameter is the rank of its low-rank matrices, typically set somewhere between 4 and 64 depending on the complexity of the task being fine-tuned for — lower ranks mean fewer trainable parameters and faster, cheaper training, but risk not capturing enough of the adaptation needed for more complex behavioral changes, while higher ranks capture more but start to erode LoRA’s efficiency advantage over full fine-tuning. In practice, many fine-tuning tasks (adapting tone, following a specific format, learning a narrow domain vocabulary) work well with quite low ranks, while more substantial behavioral or knowledge changes may require experimenting with higher rank values to find where quality plateaus relative to the added training cost.

Advertisement (In-Content)

Historical figures and technical concepts for informational purposes only. Not technical, professional, legal, or financial advice. Sources: Official Documentation.