Music Generation Model Based on Multi-Encoder Transformer and Improved GAN
Ruoqi WangDriven by the advancement of artificial intelligence in music generation and cross-modal composition, achieving coherence of musical themes and accurate matching of background music in videos has become a research hotspot. To this end, this study constructs a music generation model with a multi-encoder Transformer and an improved generative adversarial network. This model combines a basic encoder and a time-series encoder to enhance the capture of musical themes and achieves selective modeling of theme information through an adaptive theme decoder. Simultaneously, it achieves accurate alignment of video motion information with musical rhythm through a spatial channel attention mechanism and multi-scale convolution optimization of the generator and discriminator. Findings indicate that in four music styles, piano, jazz, electronic, and rock, the model’s average ratio of empty bars is 11–13%, the number of used pitch classes is 44–46, and the ratio of qualified notes is 86–89%. Furthermore, in a listening test, the model’s overall subjective evaluation score is 4.36[Formula: see text] ± [Formula: see text]0.25, close to the 4.53[Formula: see text] ± [Formula: see text]0.22 score of real music samples. Therefore, the model demonstrates high accuracy, naturalness, and robustness in multi-style music generation and cross-modal composition, offering an efficient intelligent solution for intelligent composition and human-computer interactive music creation.