Simulation-based deep learning for IoT-oriented music improvisation optimization

0Citations
Citations of this article
6Readers
Mendeley users who have this article in their library.

This article is free to access.

Abstract

Real-time music improvisation demands continuous audio effect adjustment, yet human control is limited by latency and cognitive load. Conventional model- or rule-based controllers struggle with cross-modal data fusion, often producing suboptimal trajectories. A dynamic architecture addresses this by analyzing multimodal cues. This research proposes a Music Effect Control with Chaotic Wing Suit Flying Search optimized-Seq2Seq (MEC-CWFSO-SeqNet) model, integrating attention-enhanced sequence modeling with a multi-layered MEC-CWFSO engine for adaptive hyperparameter tuning and output refinement. Multimodal inputs, including audio signals and motion sensor data, are processed using an encoder-decoder architecture. During training, the MEC-CWFSO-SeqNet algorithm iteratively adjusts latent embeddings to enhance predictive performance. A diverse dataset of expressive musical improvisations, incorporating instrumental and sensor-enhanced performance, is employed with high temporal resolution to ensure accurate multimodal alignment. Data augmentation techniques such as time stretching, pitch shifting, and Z-score normalization are applied to create balanced, key-agnostic sequences. Chroma features extracted via STFT and CQT were combined with performance and sensor data into 256-dimensional vectors. Transformer-based encoder layers model cross-domain relationships, while the attention-driven decoder predicts effect parameters. CWFSO dynamically optimizes decoder behavior and loss weights through chaotic search. MEC-CWFSO-SeqNet achieved improved performance with PF (+ 0.05), SPD (-0.31), ADP (-0.04), NDG (0.0033), PCG (0.017), and an accuracy of 98.8% compared to existing methods. Simulation-based measurements indicate an average latency of 46 ms per frame on Intel Core i9 CPU, below the ~ 100 ms perceptual threshold for live musical improvisation, demonstrating that attention-based sequence learning with chaotic meta-heuristic search enables a robust, timely expressive effect controller.

Cite

CITATION STYLE

APA

Liu, M. (2026). Simulation-based deep learning for IoT-oriented music improvisation optimization. Discover Artificial Intelligence, 6(1). https://doi.org/10.1007/s44163-026-01064-y

Register to see more suggestions

Mendeley helps you to discover research relevant for your work.

Already have an account?

Save time finding and organizing research with Mendeley

Sign up for free