scieee AI-readable full text Open interactive document viewer

two tales from the shadows of the grid

Lindgren, Brian

Abstract

abstract for musical composition 'two tales from the shadows of the grid'

Full text

two tales from the shadows of the grid Brian Lindgren University of Virginia [email protected] I. THE COMPOSITION two tales from the shadows of the grid is a composition for the EV, an augmented string instrument that integrates IRCAM’s RAVE variational autoencoder for real-time neural audio synthesis. The RAVE model used in this work was trained exclusively on the EV’s own recordings, allowing the instrument to perform through a learned version of itself—functioning as an internal “meta-resonator.” This selfreferential architecture transforms the EV from a sound source into a system that listens, interprets, and rearticulates its own material in real time. The piece investigates how physical gesture and machine learning can merge into a single expressive feedback loop, where sound emerges from both human motion and algorithmic response. It unfolds across two contrasting movements that alternate between automation and embodied control. In the first movement, two low-frequency oscillators modulate the opening latent dimensions of the model, producing a slow, cyclical pulse whose contour is shaped by bow pressure and amplitude. A melodic figure on the A string modulates a third dimension, while spectral centroid data control another, coupling pitch and timbral motion. The performer periodically overrides automated behavior by manually steering one latent axis, creating delicate tension between mechanical periodicity and human inflection. The second movement reverses the relationship. The model begins in a near-static resonance that gradually destabilizes as low tones from the C string activate deeper layers of the latent space. Melodic gestures disturb this balance, while interventions on other strings restore it, producing waves of suspension and release. Additional dimensions are gradually introduced, expanding the spectral density and culminating in a dense, timbrally saturated cadence. The composition traces a negotiation between performer and model—a conversation between embodied technique and machine memory. II. THE EV The EV extends the design of a bowed string instrument into a hybrid digital–acoustic ecosystem while preserving its tactile and resonant character. Early prototypes fused the instrument’s acoustic output with synthesized material through convolution and later integrated ambisonic projection for spatial performance contexts [2]. Subsequent iterations added physical modeling and reverberation, forming a cohesive system in which acoustic gesture drives digital transformation. The current Pure Data implementation [3] assigns each string a dedicated engine composed of FFT convolution, granular delay, reverberation, physical modeling, Paulstretch freeze, and ambisonic panning. Each engine supports up to eight voices, and a routing matrix connects seven modulation sources to twenty possible destinations. The EV’s 3Dprinted frame houses four infrared optical pickups—one per string—whose signals are digitized by a Bela board and sent to the host computer. Beyond technical design, the EV explores how instrumental practice can evolve through computation. By embedding the bow–string interface within convolution, spatialization, and neural synthesis, it reframes the instrument as a site of collaboration between embodied action and algorithmic agency. III. RAVE INTEGRATION RAVE [1] is a lightweight variational autoencoder capable of real-time neural synthesis on standard CPUs. It has been adopted in several instruments that emphasize tactile interaction with latent spaces. Sophtar [6] maps finger pressure to a RAVE model trained on vocal material, while Stacco [5] employs magnetic sensors and attractor fields to visualize latent relationships in physical space. The EV’s use of RAVE differs in its self-referential design: the model is derived solely from the instrument’s own corpus, effectively turning the EV into its own training data. Within performance, parameters such as frequency, amplitude, and spectral centroid modulate the latent dimensions, blending automated and gestural control. This integration positions AI not as an external compositional tool but as an embedded collaborator—an entity that listens, transforms, and resonates within the performer’s expressive field. REFERENCES [1] A. Caillon and P. Esling, “RAVE: A Variational Autoencoder for Fast and High-Quality Neural Audio Synthesis,” CoRR, abs/2111.05011, 2021. [2] B. Lindgren, “The EV: An Iterative Journey in Digital–Acoustic String Instrument Augmentation,” Proceedings of NIME, 2025. [3] B. Lindgren, “An Augmented String Instrument Architecture Created with Pure Data (forthcoming),” Proceedings of PdMaxCon25, 2025. [4] B. Lindgren, “Exploration of Spatial Composition with a New Electronic Stringed Instrument,” Organised Sound, 2025. [5] N. Privato, V. Shepardson, G. Lepri, and T. Magnusson, “Stacco: Exploring the Embodied Perception of Latent Representations in Neural Synthesis,” Proceedings of NIME, 2024. [6] F. Visi, “The Sophtar: A Networkable Feedback String Instrument with Embedded Machine Learning,” Proceedings of NIME, 2024.