by Aleksandra Teng Ma
Improvisation has always been a powerful way to create new kinds of music: musicians feed off each other in real time, and new forms emerge from that back-and-forth. Machine learning could push this even further. Since ML models represent musical structure differently from how humans understand it, they have potential to open up creative territory that’s genuinely new, not just music generated “in the style of” existing artists. But realizing that potential in improvisation requires humans and machine learning systems to communicate meaningfully in real time. This is where free improvisation is most demanding: it removes any pre-agreed style, tempo, key, or role; a significant part of the performance has to emerge from interaction and communication itself. Most current real-time AI improvisation systems, however, treat communication as an afterthought, layering interaction strategies onto a generative algorithm through static, predefined modes like trading or call-and-response, rather than building communication into the algorithm design itself. To close this gap, we looked at how expert human musicians actually communicate with each other during free improvisation, in hopes of eventually building that understanding into generative systems at the design level.
Through a six-month co-design process with expert improvisers, we formalized a communication model built from two point-in-time actions, initiation and acknowledgement, which compose into three temporal states: negotiation (initiations without uptake), proposal (an initiation awaiting acknowledgment), and stability (the shared space that follows). The figure below shows the model itself, alongside a real clip from our dataset, annotated independently by both musicians who played it.
To ground the model in real playing, we recorded the H2H (Human-to-Human) Music Improvisation Dataset: six hours of duo free improvisation across five expert musicians, captured with clean per-player audio stems, and synchronized multi-camera video. Each musician annotated their own intentions and their perception of their partner’s intentions to capture the intent-perception gap.
Our project page provides an interactive preview and downloadable dataset, along with the paper: https://h2himprov.github.io/. This work will also be presented at the 2026 International Society for Music Information Retrieval Conference (ISMIR).
