Evaluating Prompt Engineering Techniques for LLaMA-3: A Study of Zero-Shot, Few-Shot, and Chain-of-Thought Prompts Across Reasoning and Classification Tasks
Darren Astle Travasso, Aboozar TaherkhaniPrompt engineering has emerged as a practical and resource-efficient alternative to fine-tuning large language models (LLMs), particularly as these methods have a lower computation cost than fine-tuning. In this paper, three widely adopted prompting techniques—Zero-Shot, Few-Shot, and Chain-of-Thought (CoT)—were assessed. While these prompting strategies are well established, practitioners still lack clear guidance on when each technique should be preferred, which types of tasks they fail to support reliably, and how performance trade-offs may affect the practical use of LLM-based systems. These techniques are tested across three benchmark tasks: sentiment classification (SST-2), multiple-choice questions (CommonsenseQA), and multi-step math problem solving (GSM8K) using Meta’s LLaMA-3 8B Instruct model. We provide a thorough performance comparison based on accuracy, F1 score, and solve rate. The solve rate is highlighted as a complementary metric for evaluating the usability of LLM outputs—a factor often overlooked in the existing literature. Experimental results showed that Few-Shot prompts are particularly effective in structured classification tasks, while CoT prompts excel in logic-heavy tasks that require multi-step reasoning. On the classification task, Few-Shot prompting improved the solve rate but achieved lower accuracy and F1 score than Zero-Shot prompting. On the multiple-choice questions, Zero-Shot, Few-Shot, and CoT prompting achieved a solve rate of 100%. On multi-step math problem solving, CoT improved the solve rate and interpretability compared to Zero-Shot but did not surpass Zero-Shot accuracy. Overall, the results demonstrate that the effectiveness of prompting strategies is task-dependent, with differences observed in both accuracy and output validity across the three benchmark tasks.