Publications-Periodical Articles

Article View/Open

Publication Export

Google ScholarTM

NCCU Library

Citation Infomation

Related Publications in TAIR

題名 An automated system for generating expressive virtual violin performances from music
作者 劉昭麟
Lin, Ting-Wei;Kao, Hsuan-Kai;Wei, Wen-Li;Liu, Chao-Lin;Lin, Jen-Chun;Su, Li
貢獻者 資訊系
關鍵詞 Generative AI; cross-modal generation; Vtuber; music-to-performance generation; expressiveness
日期 2026-01
上傳時間 13-Aug-2026 09:15:23 (UTC+8)
摘要 Motion capture (MOCAP)-free music for performance generation using deep generative models is emerging as a promising solution for the next-generation animation industry. This technology allows users to create dynamic musical performance animations without relying on MOCAP. However, implementing MOCAP-free content-to-performance generation systems presents significant challenges. First, integrating various standalone models into a cohesive system is essential, as each model governs distinct aspects of the avatar's behavior. For example, a facial expression generation module influences the avatar's facial expressions, whereas a fingering generation module determines hand positions. Second, most applications focus on humanonly performance generation, such as virtual vocalists and virtual dancers, without considering interactions with other objects, such as instruments, referred to as human-instrument performance generation. To our knowledge, comprehensive human-instrument content-to-performance generation systems are still rare. In this paper, we present a complete content-to-performance generation system, demonstrating its capabilities through a web application that enables users to create animations from their music content. This system incorporates four modules to generate parameters for controlling avatars: 1) a facial expression module, 2) a fingering generation module, 3) a body movement generation module, and 4) a video shot generation module. Additionally, we integrate an expressive music synthesis module to generate expressive audio from symbolic music data. Quantitative evaluations confirm the effectiveness of the four modules, whereas a user study provides qualitative insights into the system's performance. The web application is available on our website (https://virtual-musician.iis.sinica.edu.tw:8800).
關聯 IEEE Transactions on Multimedia, pp.1-14
資料類型 article
DOI https://doi.org/10.1109/TMM.2026.3651045
dc.contributor 資訊系
dc.creator (作者) 劉昭麟
dc.creator (作者) Lin, Ting-Wei;Kao, Hsuan-Kai;Wei, Wen-Li;Liu, Chao-Lin;Lin, Jen-Chun;Su, Li
dc.date (日期) 2026-01
dc.date.accessioned 13-Aug-2026 09:15:23 (UTC+8)-
dc.date.available 13-Aug-2026 09:15:23 (UTC+8)-
dc.date.issued (上傳時間) 13-Aug-2026 09:15:23 (UTC+8)-
dc.identifier.uri (URI) https://ah.lib.nccu.edu.tw/item?item_id=184445-
dc.description.abstract (摘要) Motion capture (MOCAP)-free music for performance generation using deep generative models is emerging as a promising solution for the next-generation animation industry. This technology allows users to create dynamic musical performance animations without relying on MOCAP. However, implementing MOCAP-free content-to-performance generation systems presents significant challenges. First, integrating various standalone models into a cohesive system is essential, as each model governs distinct aspects of the avatar's behavior. For example, a facial expression generation module influences the avatar's facial expressions, whereas a fingering generation module determines hand positions. Second, most applications focus on humanonly performance generation, such as virtual vocalists and virtual dancers, without considering interactions with other objects, such as instruments, referred to as human-instrument performance generation. To our knowledge, comprehensive human-instrument content-to-performance generation systems are still rare. In this paper, we present a complete content-to-performance generation system, demonstrating its capabilities through a web application that enables users to create animations from their music content. This system incorporates four modules to generate parameters for controlling avatars: 1) a facial expression module, 2) a fingering generation module, 3) a body movement generation module, and 4) a video shot generation module. Additionally, we integrate an expressive music synthesis module to generate expressive audio from symbolic music data. Quantitative evaluations confirm the effectiveness of the four modules, whereas a user study provides qualitative insights into the system's performance. The web application is available on our website (https://virtual-musician.iis.sinica.edu.tw:8800).
dc.format.extent 104 bytes-
dc.format.mimetype text/html-
dc.relation (關聯) IEEE Transactions on Multimedia, pp.1-14
dc.subject (關鍵詞) Generative AI; cross-modal generation; Vtuber; music-to-performance generation; expressiveness
dc.title (題名) An automated system for generating expressive virtual violin performances from music
dc.type (資料類型) article
dc.identifier.doi (DOI) 10.1109/TMM.2026.3651045
dc.doi.uri (DOI) https://doi.org/10.1109/TMM.2026.3651045