Skip to content
Learn AI by building
← World of AI

Generative AI

QLoRA

Fine-tuning a quantised model through small adapter matrices, so it fits on one GPU.

Low-rank adaptation freezes the original weights and trains a pair of small matrices alongside them. QLoRA adds four-bit quantisation of the frozen base, cutting memory again.

The practical effect is that fine-tuning a substantial open model becomes a consumer-hardware task rather than a cluster task, and the resulting adapter is small enough to ship as a file.

Also in Generative AI

JOIN NOW

Begin the first module

It is free, it is the real curriculum, and if it is not for you, you have lost nothing but an evening.

Join any time · Build AI skills at your pace