← World of AI
Generative AI
QLoRA
Fine-tuning a quantised model through small adapter matrices, so it fits on one GPU.
Low-rank adaptation freezes the original weights and trains a pair of small matrices alongside them. QLoRA adds four-bit quantisation of the frozen base, cutting memory again.
The practical effect is that fine-tuning a substantial open model becomes a consumer-hardware task rather than a cluster task, and the resulting adapter is small enough to ship as a file.
JOIN NOW
Begin the first module
It is free, it is the real curriculum, and if it is not for you, you have lost nothing but an evening.
Join any time · Build AI skills at your pace