Assessing AI-Generated vs. Human-Authored Spear Phishing SMS Attacks: An Empirical Study
Jerson Francia, Derek Hansen, Benjamin Schooley, Matthew Taylor, Shydra Valynn Murray, Rebekah Cornelius, Greg SnowPersonalized phishing is difficult to defend against because messages can be tailored to a target’s work, interests, and social context. Large language models may make such tailoring faster and easier, but it remains unclear whether messages produced from simple prompts are more convincing than those written by people. This 25-target pilot study compared personalized smishing messages generated by GPT-4 with messages written by novice student authors working under time constraints. Using the proposed Threshold Ranking Approach for Personalized Deception (TRAPD), participants ranked 12 messages written for them, indicated the point at which they would intend to click, explained their reasoning, and judged whether each message was authored by GPT-4 or a human. GPT-4-generated messages elicited an intention to click more often than student-authored messages (28% versus 21%), although the difference was uncertain. More broadly, our findings suggest that a simple prompt can produce personalized messages that participants found comparably convincing within the uncertainty of this pilot study. Job-related messages were significantly more likely to elicit an intention to click than hobby- or social-media-related messages. When asked whether a message was written by a human or generated by AI, participants identified the source no more accurately than chance, although the two study-specific message sets remained computationally distinguishable based on their text. Together, these findings suggest that accessible AI-assisted personalization may increase the practical scale of social-engineering threats, while also demonstrating both the value and current limitations of TRAPD for controlled and ethical comparison.