DOI: 10.3390/arts15080179 ISSN: 2076-0752

From Alignment to Evocation: On the Capability Boundaries and Collaborative Paths of AI Art Creation—A Framework Based on the Neuroaesthetic “Ring Scale” and Prompt Engineering

Xianqun Yi, Hongsheng Li

Recent generative art outputs across music, literature, painting and moving-image media have attracted extensive scholarly and public interest, yet evaluations of their creative capacities are mostly limited to informal observational accounts. Drawing on neuroaesthetic reasoning, this paper puts forward a dual-layer analytical framework that differentiates two distinct modes of aesthetic reception: Alignment, defined as statistical template matching, and evocation, referring to the novel association of scattered embodied memory fragments. Building on this binary categorization, the study introduces the tentative Ring Scale taxonomy—a figurative target-shooting metaphor rather than quantitative metric—as a purely descriptive tool for stratifying relative aesthetic evocation intensity. This framework further unpacks the neurocognitive underpinnings of auditory, visual and textual aesthetic pathways, alongside their combined multimodal interactions within film and television works. It tentatively accounts for why generative systems tend to deliver more cohesive aesthetic outcomes within the auditory domain, and hypothesises a present functional limitation of current large models: these systems perform comparatively well within Alignment-driven aesthetic effects, while layered high-order evocation remains constrained by inherent structural limitations of statistical training architectures. From this diagnostic observation, three directional paradigm shifts for human–AI collaborative creation are outlined: shifting from human substitution to human–machine complementarity, shifting from exhaustive template imagery generation to targeted latent fragment elicitation, and shifting from optimising figurative Ring-tier descriptive labels to pursuing transformative aesthetic fission effects. The study frames imaginative cognition as the central driving force behind fruitful human–AI co-creation, and positions prompt engineering as the actionable operational bridge connecting human imaginative thought to machine-executable generative parameters. Three tentative prompt design tactics are then elaborated: physiological arousal framing, multisensory scenario simulation prompts, and intentional strategic blank-leaving. Additionally, this work discusses the plausible constructive functions of model hallucination phenomena when viewed through the lens of high-tier aesthetic evocation, rather than merely framing such outputs as technical errors. All judgments and tier comparisons raised throughout the paper are framed as unvalidated observational hypotheses open to empirical testing. To facilitate follow-up empirical scrutiny, the paper collates a full set of testable hypotheses derived from its theoretical reasoning and outlines feasible experimental validation pipelines, with an open call for controlled empirical research to corroborate or refine the proposed qualitative framework.

More from our Archive