DOI: 10.54287/gujsa.1937530 ISSN: 2147-9542

QLoRA vs. LoRA: An Empirical Study of Performance–Efficiency Trade-offs in Fine-Tuning Mistral‑7B on Instruction‑Following Data

Abdullah Talha Kabakuş
Parameter‑Efficient Fine‑Tuning (PEFT) approaches have emerged as a key strategy for efficiently adapting Large Language Models (LLMs) in resource‑constrained environments. Within this family of methods, LoRA and its quantized extension QLoRA provide effective alternatives by substantially decreasing the number of trainable parameters required for adaptation. In this study, a systematic comparison of LoRA and QLoRA is presented using a controlled experimental setup with identical base models and training data, with all experiments repeated across three random seeds and statistical significance assessed via paired t‑tests. Both approaches are evaluated across multiple dimensions, including generation quality (ROUGE‑L, Exact Match, token‑level F1, and the semantic metric BERTScore), training efficiency, and GPU memory consumption. Baseline (untrained) model performance is reported to quantify the gain from fine‑tuning, and loss curves are provided to confirm convergence. The results show that QLoRA not only reduces memory usage by approximately 70% but also achieves superior performance compared to LoRA, with improvements in token‑level F1 and BERTScore F1 reaching statistical significance (p < 0.05), while other lexical metrics showed consistent but non‑significant improvements. Furthermore, the performance–efficiency trade‑offs are analyzed, and it is demonstrated that QLoRA offers a more favorable balance between computational cost and model quality. These findings highlight QLoRA as a highly effective and practical fine‑tuning strategy for resource‑constrained environments, supported by a reproducible, statistically robust evaluation pipeline.