DOI: 10.3390/app16199590 ISSN: 2076-3417

Controllable Training-Free Diffusion Style Transfer via Latent Initialization and Attention Statistical Alignment

Yanzhi Yuan, Zhiqiang Pan, Le Xia, Ruojun Yang, Xinjie Long, Yingchun Kuang, Lizhuo Zhang

Recent advances in diffusion models have demonstrated remarkable generative capabilities, but their application to style transfer remains limited by inference instability and poor adaptation to diverse styles. Many methods either rely on costly fine-tuning or sacrifice flexibility and controllability at inference time. We present a unified, training-free diffusion framework for image-guided, text-guided, and localized style transfer. Built on Denoising Diffusion Implicit Model (DDIM) inversion, the framework incorporates z-domain initialization, which aligns the initial content latent with style statistics; and key–value statistical alignment (KV-StatAlign), which progressively mixes attention keys and values and calibrates their head-wise logit energy during sampling. It further incorporates a Prompt-to-Prompt branch for reference-free text-guided stylization and mask-based composition for regional editing. Experiments on MS COCO and WikiArt show that the proposed method achieves the best FID score of 18.039 and ranks second in ArtFID, LPIPS, and CFSD among the evaluated methods. Our experiments indicate that z-domain initialization and KV-StatAlign improve stylization stability and balance style fidelity with structural preservation while supporting continuous control over style strength. These results demonstrate that latent initialization and attention statistical alignment provide an effective basis for flexible style transfer without additional training.