LoRA and QLoRA mechanics

LoRA (Low Rank Adaptation) fine tunes a model by adding small, trainable rank decomposition matrices to each attention and FFN layer instead of updating the full weight matrices. The number of trainable parameters drops from billions to millions, enabling fine tuning on consumer GPUs.