Program

29th International Conference on Digital Audio Effects
September 1–4, 2026 • MIT, Cambridge, MA, USA

Click any paper title to reveal its authors.

Download All Papers

Overview
TueSept 1
WedSept 2
ThuSept 3
FriSept 4
SatSept 5
Paper Session
Keynote
Tutorial
Poster / Demo
Break / Lunch
Social Event
Awards / Admin
Time Tuesday, Sept 1 Wednesday, Sept 2 Thursday, Sept 3 Friday, Sept 4
8:30 Registration (all day) Registration (all day) Registration (all day) Registration (all day)
8:45
9:00 Tutorial 1
Victor Zappi
Opening Remarks Paper Session 4
Filters & Distortion
Paper Session 7
Reverb 3
9:15
9:30 Paper Session 1
Reverb 1
9:45
10:00 DAFx Challenge Posters
+ Coffee Break
10:15
10:30 Coffee Break Coffee Break Coffee Break
10:45
11:00 Tutorial 2
Sebastian Schlecht & Facundo Franchino
Keynote 1
Alexey Lukin
Keynote 2
Dina Pearlman-Ifil
Keynote 3
Sean Costello
11:15
11:30
11:45
12:00 Lunch Lunch Lunch
12:15
12:30Lunch
12:45
13:00 Paper Session 2
Instruments & Voice
Paper Session 5
Sound Design & Effects
Paper Session 8
Virtual Analog
13:15
13:30
13:45
14:00 Tutorial 3
Georg Essl
14:15 Poster Craze 1 Poster Craze 2
14:30 Poster/Demo Session 1
+ Coffee Break
Poster/Demo Session 2
+ Coffee Break
Coffee Break
14:45
15:00 Paper Session 9
Audio Coding & Benchmarking
15:15
15:30 Coffee Break
15:45
16:00 Paper Session 3
Sound Synthesis
Paper Session 6
Reverb 2
16:15 Tutorial 4
You (Neil) Zhang & Yoshiki Masuyama
16:30 Awards & Closing Ceremony
16:45
17:00 Handover Address
17:15 Board Meeting
17:30
17:45
18:00 Welcome Reception
Presented by Analog Devices
18:15
18:30
18:45
19:00 Banquet
Presented by Eventide
19:15
19:30
19:45
20:00 Concert
Presented by Soundtoys
20:15
20:30
20:45
21:00
21:15
21:30
21:45
Time Session Location
8:30 Registration (all day) W18 Lobby
9:00
Tutorial 1
Victor Zappi
Tull
10:30 Coffee Break W18 Lobby
11:00
Tutorial 2
Sebastian Schlecht & Facundo Franchino
Tull
12:30 Lunch W18 Lobby
14:00
Tutorial 3
Georg Essl
Tull
15:30–16:15 Coffee Break W18 Lobby
16:15–17:45
Tutorial 4
You (Neil) Zhang & Yoshiki Masuyama
Tull
18:00–20:00 Welcome Reception
Presented by Analog Devices
W18 Lobby + Outdoor Plaza
Time Session Location
8:30 Registration (all day) W18 Lobby
9:00 Opening Remarks Tull
9:30–10:30
Paper Session 1: Reverb 1
Session Chair: Gloria Dal Santo
Ambisonic Decoder Equalization in Reverberant Environments via Closed-Hull Crosstalk Inversion
Alex Tung and Mark Rau
In this work, the authors develop a higher order Ambisonic (HOA) decoder that compensates for listening room reverberation by cascading a conventional decode matrix with a crosstalk matrix derived from room impulse responses (RIR) constrained over the listener area, namely by sampling over a spherical boundary according to HOA convention. A decoder is generated for simulated RIRs of mixed specular and diffuse reverberation, and spectral and spatiotemporal energy distributions are shown for directional and diffuse HOA signals rendered through the reverberant room. Results demonstrate that the decoder renders temporally-variant directional sources with higher directivity as compared to a conventional decoder.ing Ambisonic decoders in reverberant environments using closed-hull crosstalk convolution. By modeling the acoustic path from each loudspeaker to the listener's ears via measured room impulse responses, we formulate equalization filters that compensate for room-induced distortions in the decoded signals. The proposed closed-hull approach constrains the solution to preserve perceptually relevant spatial cues while minimizing spectral coloration. Listening tests demonstrate improved externalization, timbral transparency, and localization accuracy compared to standard free-field decoding in reverberant conditions.
Perceptual Optimisation of Loudspeaker-Based Reproduction
Antoine Souchaud, Llado Pedro, Rapolas Daugintis, Annika Neidhardt, Zoran Cvetkovic and Enzo De Sena
This paper proposes POLAR, a framework for the optimisation of loudspeaker signals using end-to-end differentiable perceptual loss functions. The framework optimises multiple perceptual attributes across multiple listeners, offering a versatile method for a range of problems. This versatility stems from the ability to customise the number of loudspeakers, listeners, and the weighting applied to different perceptual attributes. Here, we apply the method to four problems: (a) source panning for a single listener in stereo reproduction, (b) single-listener colouration matching in stereo reproduction, (c) extended sweet spot using stereo pairs beamforming, and (d) multi-attribute perceptually driven panning in stereo reproduction. The first three problems are evaluated against solutions traditionally used for these tasks: solutions of (a) are shown to be similar to those obtained with tangent panning law and vector-base amplitude panning (VBAP), solutions of (b) are shown to be similar to those obtained for cross-talk cancellation, and solutions of (c) are shown to be similar to those obtained in earlier work on directivity pattern optimisation for sweet spot widening. Each of these solutions was previously obtained using fundamentally different methodologies, demonstrating the flexibility and broad applicability of the proposed framework.
Multi-Source Extension and Hyperparameter Optimization of the DiffRIR Framework for Room Impulse Response Synthesis
Luka Fehrmann, Mason L. Wang, Martin Rumori and Peter Plessas
Efficient prediction of Room Impulse Responses (RIRs) is a cornerstone for immersive virtual acoustics and scalable room acoustic modeling. This study extends the DiffRIR framework – proposed by Wang et al. in Hearing Anything Anywhere – by introducing a multi-source training logic and systematically optimizing its convergence behavior to overcome the inherent limitations of the original framework. Our results reveal that multi-source training acts as implicit data augmentation, where the resulting increase in spatial entropy enhances the model's spectral accuracy. Furthermore, we demonstrate that the model exhibits remarkable robustness against geometric inaccuracies, maintaining numerical stability even with source positional offsets of up to 4 m in single-source baseline evaluations. By identifying a learning rate of 3×10⁻², we were able to reduce the training duration to 23% of the original baseline without compromising prediction accuracy. While the increased complexity of multi-source fields necessitates a trade-off in temporal precision – quantified via our newly integrated Energy Decay Convergence (EDC) metric – this research provides an efficient and resilient solution for acoustic simulations in complex environments.
Parametric Resynthesis of Measured Spatial Room Impulse Responses
Anthony Gallien, Benoit Alary and Markus Noisternig
Spatial Room Impulse Responses (SRIRs) are fundamental to immersive audio rendering and have become a key focus of recent machine learning research in acoustics and auralization. Due to the high computational cost of direct convolution, spatial audio systems commonly employ artificial reverberation algorithms. However, these approaches often fail to accurately reproduce the spatial, temporal, and spectral characteristics of early reflections, leading to notable deviations from measured SRIRs. This paper presents a comprehensive framework for the analysis and efficient resynthesis of SRIRs captured with Spherical Microphone Arrays (SMAs). The proposed method accounts for hardware-induced artifacts, including scattering and spatial aliasing. Early reflections are reconstructed using a parametric approach based on the Herglotz analysis method, while late reverberation is synthesized using a Directional Feedback Delay Network (DFDN) with optimized filter-attenuation and correlation-matching. The proposed framework produces signals whose spatial correlation and Energy Decay Relief (EDR) closely match those of measured SRIRs, demonstrating its effectiveness for both real-time spatial audio rendering and realistic dataset generation for machine learning applications.
Tull
10:30 Coffee Break W18 Lobby
11:00
Keynote 1
Alexey Lukin
Tull
12:00 Lunch W18 Lobby
13:00
Paper Session 2: Instrument and Voice Modeling
Session Chair: Marcelo Caetano
Physical Model of the Chinese Yehu for Sound Synthesis
Zhen Zheng, Champ Darabundit and Gary Scavone
The yehu is a Chinese bowed string instrument featuring a resonator carved from a coconut shell, a seashell-based bridge, and two silk strings. This paper proposes a physical model of the yehu and reports on simulations using a finite-difference scheme with measurement-based physical characterization. The proposed model consists of two stiff strings coupled at the bridge, a bow with elastic bow hairs, a stopping finger, and a modal model of the bridge. A non-iterative solver based on energy quadratization is used to model the finger–string contact force, while an iterative solver is used for elasto-plastic bow-string friction force. The bridge-body model is based on a modal characterization obtained from the measured bridge admittance. The measured radiation transfer function is represented as a bank of parallel second-order filters and is applied to the simulated bridge force to incorporate body radiation characteristics. Finally, computational performance tests are conducted, showing that the proposed model is capable of real-time computation.
Eigensystem Realization of Violin Bridge Admittances
Riccardo Giampiccolo, Alessandro Ilic Mezza, Raffaele Malvermi, Mirco Pezzoli, Alberto Bernardini and Fabio Antonacci
Modeling violin bridge admittance is a long-standing problem in musical acoustics, with applications in sound analysis, synthesis, and virtual instrument design. In this work, we investigate the use of the Eigensystem Realization Algorithm (ERA) for deriving reduced-order state-space models directly from measured impulse responses. The proposed approach allows us to extract dominant system dynamics and obtain compact realizations without requiring explicit modal parameterization. We evaluate ERA on a dataset of modern and historical violins and compare it against established modal and state-space identification methods. Experimental results demonstrate that ERA outperforms existing approaches by achieving lower reconstruction errors in both the time and frequency domains while preserving perceptually relevant characteristics of the bridge response. Furthermore, we show that the state-space realizations obtained using ERA reproduce the target frequency-dependent energy decay more accurately than models obtained using the baseline methods. These findings support the use of ERA as an efficient and flexible alternative for modeling violin bridge admittances, with applications that span from audio synthesis and processing to instrument virtualization.
Measurement-Informed Nonlinear Modal Synthesis of 65 Classical Guitars
Michele Ducceschi, Riccardo Russo and Craig J. Webb
When a classical guitar string is plucked, vibration energy flows through the bridge into the body and is radiated as sound. Synthesising this process for a large collection of instruments requires both an efficient nonlinear string model and a robust method for extracting instrument-specific parameters from measurements. This paper addresses both issues. Starting from the publicly available dataset of Mores, which provides impulse-response measurements on 65 classical guitars, modal parameters of the bridge compliance and of the bridge-to-air radiation path are extracted for each instrument. These feed a nonlinear string model in which transverse vibration is governed by a geometrically exact elastic potential coupled at an interior bridge point to the measured body data. The nonlinear potential is quadratised via the Scalar Auxiliary Variable (SAV) method, so that the equations of motion become linear in a scalar variable and a known gradient vector, even at the continuous level. After time discretisation, the coupled system is inverted through two sequential Sherman–Morrison rank-one updates (one for the bridge coupling, one for the SAV nonlinearity), yielding an O(N) algorithm per time step. Two regularisation techniques prevent long-term drift of the auxiliary variable. The complete pipeline is demonstrated by synthesising plucked notes across all frets and strings for each of the 65 guitars.
Differentiable Articulatory Copy-Synthesis of Biphonic Singing
Mateo Cámara, María Pilar Daza, Fernando Marcos and Jose Luis Blanco
Sygyt is a Tuvan style of biphonic singing in which a low vocal drone is sustained while a high harmonic is selectively amplified in the 1–3 kHz region. Copy-synthesizing this effect remains challenging for articulatory models, since it requires fine control of narrowly focused resonances that standard low-dimensional tract parameterizations cannot easily reproduce. We address this problem with a differentiable Kelly–Lochbaum waveguide augmented with a sublingual second source, cubic B-spline tract parameterization, and spatially varying learnable damping, optimized end-to-end by gradient descent from audio. On 20 segments from two independent sygyt datasets (5 singers, 10 pitches), the proposed model reduces log-spectral distance by 30–38% relative to an articulatory baseline, with the largest gains concentrated in the overtone region. Cepstral-envelope analysis further shows more accurate recovery of the merged formant structure characteristic of sygyt production. The model also outperforms a DDSP harmonic-plus-noise baseline with direct per-harmonic spectral control, suggesting that explicit acoustic structure is a useful inductive bias for overtone-singing copy-synthesis.
Adapting Diffusion-Based Music Synthesis to Speech and Singing Voice Conversion
Ben Maman, Frank Zalkow, Hans-Ulrich Berendes, Paolo Sani, Christian Dittmar and Meinard Müller
Recent diffusion-based generative models have achieved strong results in domain-specific audio generation tasks such as speech, singing, and instrumental music synthesis. However, these models are typically specialized and do not generalize well to mixed or intermediate audio types. In this work, we adapt a diffusion-based model originally designed for multi-instrument music synthesis to voice conversion, covering both speech and singing within a unified framework. Specifically, we extend musical note-based conditioning to include phonetic posteriorgrams (PPGs) and pitch contours, and reinterpret timbre conditioning as speaker or singer identity via feature-wise linear modulation. Experiments show that the adapted model matches or surpasses a dedicated voice conversion system in terms of naturalness and performer similarity, while maintaining accurate pitch control across speech and singing. At the same time, we observe limitations in phonetic fidelity and a degradation in vocal quality when incorporating instrumental training data. Furthermore, we demonstrate that off-the-shelf feature extractors provide effective conditioning signals, enabling large-scale self-supervised training without manual annotations. These results highlight the potential of cross-domain model transfer towards unified audio generation systems capable of handling speech, singing, and music.
Tull
14:15 Poster Craze 1 (1-minute poster previews) Tull
14:30 Poster/Demo Session 1 + Coffee Break
Posters
Explicit Wave Digital Model of the Fulltone OCD Pedal Based on Canonical Piecewise-Linear Functions
Riccardo Giampiccolo, Stefano Polimeno, Carlo Macrì, Alice Lenoci, Oliviero Massi and Alberto Bernardini
Virtual Analog (VA) modeling aims at digitally emulating analog audio equipment while preserving its characteristic nonlinear behavior and musical expressiveness. In the context of guitar effects, overdrive pedals represent a cornerstone of many signal chains, as they strongly contribute to the perceived dynamics, articulation, and timbral identity of the instrument. Among these, the Fulltone OCD overdrive is considered a standard in both studio and live environments, being widely adopted across rock and metal genres. In this article, we present an explicit Wave Digital (WD) model of the Fulltone OCD (v2) pedal. By exploiting the circuit topology, the MOSFETs and the germanium diode composing the asymmetric clipping stage are grouped into a single equivalent nonlinear element, enabling an explicit WD realization that avoids costly iterative solvers. The resulting nonlinear characteristic is approximated by means of a Canonical Piecewise-Linear (CPWL) function, yielding a compact and efficient explicit model suitable for real-time implementation. The proposed model is validated against reference simulations and implemented both in MATLAB and as a real-time audio plug-in using the JUCE framework.
Fourier Neural Operators for Sample-Rate-Independent Virtual Analog Modeling
Oliviero Massi, Alessandro Ilic Mezza and Alberto Bernardini
Neural networks that operate directly on time-domain signals are widely used for virtual analog (VA) modeling. A key limitation of these models is their dependence on the sampling rate used during training, which becomes implicitly encoded in the learned parameters, so that changing it generally alters the realized dynamics. Although architectural modifications to recurrent neural networks have been proposed to enable sample-rate independent operation, these approaches are inherently tailored to upsampling and do not accommodate downsampling scenarios. In this manuscript, we present a VA modeling framework based on Fourier Neural Operators (FNOs) adapted to process fixed-duration audio frames. The proposed formulation defines the learned mapping over a fixed temporal support and evaluates it on uniform grids of different densities, so that a model trained at a single sampling rate can be applied at unseen sampling resolutions. Numerical results on a nonlinear transistor circuit show that the proposed model achieves competitive accuracy in upsampling scenarios while remaining directly applicable to downsampling, unlike a sample-rate independent baseline recurrent architecture.
Deep Regularized RNNs for Virtual Analog
Valtteri Kallinen, Lauri Juvela and Thom Sherson
Virtual analog (VA) modeling methods seek to emulate analog audio hardware using digital signal processing (DSP). Modeling approaches fall into three broad categories: white-box methods, which use detailed device knowledge for accurate simulation; gray-box methods that use generic DSP blocks to model the system; and black-box methods, which rely solely on opaque models learned from input–output data. A category of architectures used widely in black-box modeling are recurrent neural networks (RNNs). To model device controls, the control values can be provided as conditioning input to the network. However, when the conditioning is time-varied, the models are susceptible to producing noise artifacts. Regularization of the RNN dynamics significantly reduces these artifacts, though at a loss in modeling accuracy. This paper closes the dynamics regularization quality gap by introducing deep control-conditioned LSTMs and a gammatone filterbank (GFB) loss. Experiments indicate that the proposed method achieves comparable modeling performance as unregularized baselines while avoiding the noise artifacts caused by time-varying control inputs.
FM Parameter Estimation with Low-Order Rational Constraints on Wasserstein Loss Landscape
Ryoya Tabata, Masaki Iwaya and Kazunobu Kondo
Frequency modulation (FM) synthesis has been widely used in music production and sound design due to its ability to generate rich timbres with few control parameters. However, estimating the frequency parameters from a target sound remains challenging because different parameter configurations can yield similar spectra, creating numerous local minima in the loss landscape. In this paper, we analyze the Wasserstein distance loss landscape for two-operator FM synthesis under practical FFT-based spectral representations and show that it exhibits non-differentiable ridges at rational frequency ratios, arising from negative-frequency folding and spectral ordering transitions. Exploiting this structure, we propose a constrained gradient-based optimization strategy that constrains the frequency ratio in each optimization run to an interval bounded by consecutive low-order rational ratios and retains the lowest-loss candidate across intervals. Experimental results from controlled ablations show that maintaining the constraint throughout optimization improves reliability over random initialization and initialization-only constraints, particularly for more complex spectra at higher modulation indices.
Enhancing Automatic Chord Recognition via Pseudo-Labeling and Knowledge Distillation
Nghia Phan, Rong Jin, Gang Liu and Xiao Dong
Automatic Chord Recognition (ACR) is constrained by the scarcity of aligned chord annotations, which are costly to acquire. At the same time, open-weight pre-trained models are more accessible than their proprietary training data. In this work, we present a two-stage training pipeline that leverages pre-trained models together with unlabeled audio. The proposed method decouples training into two stages. In the first stage, we use the pre-trained BTC model as a teacher to generate pseudo-labels for over 1,000 hours of diverse unlabeled audio and train a student model solely on these pseudo-labels. In the second stage, the student is continually trained on ground-truth labels as they become available. To prevent catastrophic forgetting of the representations learned in the first stage, we apply selective knowledge distillation (KD) from the teacher as a regularizer. In our experiments, two models (BTC, 2E1D) were used as students. In Stage 1, using only pseudo-labels, the BTC student achieves about 99% of the teacher's performance, while the 2E1D model achieves about 97% of the teacher's performance across seven standard mir_eval metrics. After continual training with labeled data in Stage 2, the resulting BTC student model consistently surpasses both the traditional supervised learning baseline and the original pre-trained teacher model across all metrics. The resulting 2E1D student model also outperforms the supervised baseline and approaches teacher-level performance, with both models demonstrating substantial gains on rare chord qualities.
Demos
A Perceptually Inspired Single Parameter Auditory Distance Renderer for Music Production
Stefanos Biliousis and Cumhur Erkut
Conveying auditory distance in a digital audio workstation requires balancing several uncoupled tools (reverb, gain, equalization, pre-delay) by hand, a workflow that is cognitively demanding and easily produces spatially incoherent results. We demonstrate a real-time VST3 plugin that derives five correlated distance cues from a single normalized control and keeps them mutually coherent by construction, grounded in the psychoacoustics of auditory distance perception. A headphone listening test with 18 participants showed that the coupled renderer roughly halves distance placement error relative to an uncoupled manual mix and was unanimously preferred on composite spatial quality. Users sweep one knob and hear sources move convincingly from near to far on multitrack material, compare the result against a manual uncoupled mix, and toggle an optional binaural externalization stage. Source code is available on GitHub.
IRIS: Continuous Spatial Navigation of Measured Acoustic Fields via Impulse Response Interpolation
Luna Valentin, Celeste Betancur Gutiérrez and Romain Michon
Impulse response (IR) collections are useful in virtual acoustics, sound design, and field-based acoustic research, but they remain difficult to explore as continuous resources in lightweight real-time plugin workflows. This demo paper presents IRIS, a VST3 plugin for arranging, navigating, and auditioning measured or user-defined IR collections in a two-dimensional navigation plane. Each IR is represented as a node whose position can be imported from metadata or assigned manually. During navigation, nearby responses are combined using Gaussian distance-based weighting, while a bounded active set limits the number of simultaneous convolutions. The system also includes smoothing, hysteresis, optional preprocessing, boundary attenuation, OSC control, and coupled multichannel handling. The demo focuses on workflow and audible behavior rather than perceptual validation. A short timing characterization reports practical real-time limits as a function of IR length, buffer size, and active-set size. IRIS is presented as a practical tool for exploratory, analytical, and creative navigation of IR collections rather than as a physically optimal interpolation method.
FDN Sandbox: Real-Time Experimentation and Analysis of FDNs
Alexandre St-Onge
This work presents sfFDN, a modular and real-time-capable C++ library for Feedback Delay Networks (FDNs), together with the companion FDN Sandbox application designed for interactive experimentation, analysis, and parameter optimization. The library implements the canonical FDN as well as several recent extensions, including filter feedback matrices, velvet-noise decorrelation filters, and two-stage graphic equalizers for attenuation and tone correction. The Sandbox application exposes these features through a graphical interface, providing a suite of real-time visualizations, as well as an optimization framework supporting nine algorithms from the ensmallen library, with built-in loss functions for both colorless reverberation and room impulse response matching. Both the library and the application are open-source.
A Frequency-Domain Reverberator Plug-In
Jonas Roth, Nishanth Kumar, Silvan Krebs, David Wieland and Christoph Studer
We present FDverb, a frequency-domain artificial reverberator, based on the idea of a vocoder with a noise carrier signal. Using a short-time Fourier transform (STFT) for analysis and synthesis, FDverb generates late reverberation by weighting spectral noise components with envelopes. We extend FDverb with early reflections, nonlinear decay, and pitch shifting. These extensions enable creative sound-design applications. We provide FDverb as an open-source DAW plug-in, using the JUCE framework.
Bunkervik Spatial Reverb Demo
Craig Webb and Michele Ducceschi
This paper accompanies a demonstration of a real-time audio plug-in for a dynamic spatial reverb. The reverb is based on acoustic measurements of the Bunkervik creative arts space in Brescia, Italy. A modal synthesis reverberation engine was created based on measured impulse responses from three locations in the tunnel. Using a common set of modal frequencies the position of the receiver can be dynamically moved through the space by interpolating between data sets of residue weights and FIR filter taps. The audio plug-in also allows real-time manipulation of the high-frequency content, damping, and microphone rotation, all of which can be modulated using two LFOs.
Pulsetable Synthesis of Wind Instrument Tones
Christian Dittmar, Simon Schwär, Manuel Peters, Stefan Balke and Meinard Müller
We revisit pulsetable synthesis, an efficient technique for generating plausible and expressive wind instrument tones. Based on the principles of pulse forming theory, this method models sound production as the periodic repetition of shaped pulses characterizing the target instruments' spectral envelope. In this approach, single-cycle waveforms, referred to as pulses, are stored in pulsetables indexed by their corresponding fundamental frequency. During synthesis, the pulses are read from these tables to form a periodic waveform, which is further shaped by time-varying low-pass filtering, amplification, and reverberation. These processes are guided by control signal contours that describe how fundamental frequency, brightness, and loudness evolve over time. Through case studies with real-world wind instrument recordings, we show how the interplay between these control signals gives rise to articulations such as attack transients, vibrato, and growl. Finally, we discuss the potential of this framework for integration into Differentiable Digital Signal Processing (DDSP) models, where neural networks could learn synthesis parameters directly from training data.
PolyMap: A 64-Channel Polyphonic Guitar Pickup System
David Wieland, Jonas Roth and Christoph Studer
In electric guitars, the vibrations of the strings are typically sensed by coils of wire combined with a magnet, called pickups. The pickups and their position along the strings contribute strongly to the instrument's sound. Most guitars feature one to three pickups, each spanning across all strings with fixed positions and generating a single mono output. The work of this Master's Thesis at ETH Zürich introduces a new pickup system called PolyMap, which senses each string individually and at multiple locations. The system is demonstrated with a custom-made eight-string guitar that contains eight pickups per string for a total of 64 pickups. The signals from these 64 pickups are individually digitized inside the guitar and transmitted over a multichannel audio digital interface (MADI), a low-latency digital audio interface, to a computer for further processing. PolyMap enables high-resolution sensing of an electric guitar's strings and enables extensive post-processing capabilities for musicians, audio engineers, and researchers. To the best of our knowledge, this is the first polyphonic guitar pickup system with such a complete feature set.
WaveNet-Style Guitar Amplifier Model Pruning for Real-Time iOS Deployment
Ryota Sato and Eli Silverstein
WaveNet-style convolutional networks emulate tube amplifiers and distortion pedals with high fidelity, but their computational cost has confined them to desktops or dedicated DSP hardware. We present a sparse-enabled WaveNet inference engine for iOS that runs heavily pruned neural guitar amplifier models in real time on iPhones. Aggressive iterative magnitude pruning removes 90% of the network weights with no perceptible loss in quality. A custom sparse C++ engine turns this sparsity directly into compute savings, sustaining low-latency real-time operation on a CPU-only iPhone implementation where the dense model cannot. On-device output matches the trained model to within int16 quantization error. At the demonstration, visitors will play a guitar through the app on iPhone hardware and A/B the on-device pruned model against the physical pedal it emulates. Source code and audio examples are available online.
Praat AudioTools: Analysis Objects as Compositional Controllers for Interpretable Sound Transformation
Shai Cohen
This demonstration presents Praat AudioTools, an open-source hybrid toolkit that repurposes Praat's phonetic-analysis environment for electroacoustic composition, sound design, and offline analysis–resynthesis workflows. Rather than treating analysis data as temporary measurements hidden inside an audio processor, Praat AudioTools exposes pitch contours, formant structures, temporal segmentations, spectral descriptors, phrase boundaries, stochastic trajectories, and host-application exchange files as editable compositional objects. These objects can be inspected, modified, chained, reused, and rendered into new sound transformations. The demonstration focuses on seven offline workflows: Neural Ambient Drone Designer, Praat for Max and Max for Live, Phase-Space Composer, Reich Generator, MCMC Musical Variation, Messagesquisse Opening, and Vector/Full-Chain composition workflows. None of the examples are presented as real-time effects. Instead, they show an "edit-in-the-middle" model in which sound is analyzed, intermediate representations are made visible, compositional decisions are applied to those representations, and the result is rendered as audio. The aim is to demonstrate a transparent alternative to both conventional black-box audio effects and end-to-end generative audio systems: a compositional environment where analysis objects become controllers, traces, scores, and reproducible technical artifacts.
Residual-Driven Adaptive Multi-Rate Quadratic Programming Framework for Nonlinear Analog Audio Circuit Emulation
Miguel Zea and Luis A. Rivera
This work extends our previously proposed Quadratic Programming (QP) approach for the emulation of nonlinear analog audio circuits by formalizing its main numerical ingredients and introducing a residual-driven adaptive multi-rate scheme. Starting from a state-space Differential Algebraic System of Equations (DAE) formulation, the nonlinear algebraic circuit device relations are replaced inside the QP by a first-order surrogate linear constraint, and the post-step nonlinear residual is shown to act as a valid defect indicator for adaptive step-size control. This yields a single-step simulation procedure that avoids the usual combination of nonlinear iterative solves and separate integration updates. The method is evaluated on a diode clipper, a BJT common-emitter amplifier, and a Colpitts oscillator, using SPICE as a baseline reference. The results show that adaptive step sizing considerably improves agreement with the reference solution, that the pseudo-inverse implementation is essentially equivalent to the full equality-constrained QP in the tested cases, and that the proposed formulation remains effective beyond the baseline clipper example, including for a self-oscillating circuit. These results position the proposed method as a promising bridge between SPICE-like interpretability and the efficiency demands of virtual analog (VA) audio applications.
W18-1202 + W18 Lobby
16:00
Paper Session 3: Sound Synthesis
Session Chair: Jeremy Hyrkas
Loopback Frequency Modulation Using a Time-Varying Delay Line
Tamara Smyth
This work examines the use of the time-varying delay line (TVDL) to implement loopback frequency modulation (LBFM), an oscillator that loops back to modulate its own frequency. Digital delay lines are used regularly in sound synthesis/processing to model the pure delay associated with one-dimensional acoustic propagation. When the delay is made time varying, the TVDL time warps the input according to a delay function, altering the input's instantaneous frequency and phase. As a result, the TVDL is well suited for delay-based effects and, in particular, those involving frequency/phase modulation for which the TVDL delay function is oscillatory and thus bounded by a maximum and minimum delay. Sustaining a constant change in sounding frequency however, corresponds to a delay function having a term that is linear in time, making it limited only by the length of the input signal. While TVDLs may still be used when there is a pitch shift, limiting the delay by simple wrapping of the delay function and/or cross fading between multiple TVDLs may not be adequate to avoid audible artifacts. In LBFM, the resulting phase has both linear and oscillating terms and the resulting signal undergoes a sustained shift in the fundamental frequency that makes a TVDL implementation more challenging. An alternate closed-form representation of the LBFM oscillator, however, provides the information necessary for accurately wrapping the TVDL delay function and ensuring it is suitably bounded so that the produced sound is free of phase distortion and audible artifacts.
FM Synthesizer Audio-Parameter Shared Embeddings
David Braun and Adam Finkelstein
Given a target sound, finding the synthesizer preset that best reproduces it remains a core problem in sound design. Existing methods treat synthesis parameters as flat vectors, discarding the signal routing and parameter interactions that produce audio. We make two contributions. First, to learn a representation of parameters including their signal routing, we design a graph neural network whose message passing structure imitates FM signal processing. Second, we adapt the multimodal objective from SLAP to learn joint embeddings of audio and FM synthesizer parameters, enabling preset retrieval from a gallery. We focus on the Yamaha DX7, where six identical sinusoid operators interact according to one of 32 routing topologies. Our graph encoder's message passing weights are shared across all nodes and layers, enabling processing of arbitrary topologies of any size. When every topology is seen during training, the DX7-GNN and two baselines achieve strong audio-to-preset retrieval. When some topologies are held out for testing, the DX7-GNN substantially outperforms both baselines despite having the fewest parameters. Our ablations further support the claim that imitating FM signal flow in a parameter encoder improves generalization to unseen topologies.
Sound Matching with a Differentiable Karplus-Strong Algorithm
Pablo Tablas de Paula, David Marttila, Rodrigo Díaz, Irán Román, Emmanouil Benetos and Joshua Reiss
We present a self-supervised, event-based sound matching model using a differentiable extended Karplus-Strong algorithm. To avoid relying on external onset and fundamental frequency detectors, we explore training methodologies combining parameter losses on synthetic data with audio losses. We demonstrate that time-domain fractional delay interpolation provides gradient accuracy comparable to frequency-sampling while avoiding time-aliasing in highly resonant time-varying scenarios. Through systematic gradient analysis, we reveal that standard spectral losses provide no meaningful directional gradients for onset times, heavily degrading joint training. Training exclusively with parameter losses on synthetic data effectively learns fundamental frequency, timbral parameters, and onset times, but struggles to generalise to monophonic studio recordings of plucked guitar. External detectors combined with audio losses generalise best, isolating the model to timbre optimisation. While our Karplus-Strong decoder recovers interpretable parameters and naturally captures the transient characteristics of plucked guitar, Harmonics plus Noise baselines yield higher reconstruction fidelity by most metrics.
Arbitrary Polygon Oscillator: Generalizing Polygonal Synthesis to Arbitrary Shapes, Morphing, and Three-Dimensional Polyhedra
Antonio Argentieri and Francesco Scagliola
Polygonal synthesis generates audio by traversing the perimeter of a polygon with a phasor; prior work uses a constant angular velocity, whereas the proposed system adopts constant arc-length (perimeter) velocity. Existing formulations operate on regular, parametrically defined polygons, producing smooth timbral transitions within a single family of shapes. This paper generalizes polygonal synthesis around a unified arc-length engine: vertex data of any origin feed the same DSP pipeline. First, we adapt the oscillator to accept arbitrary vertex configurations from an external buffer, opening the possibility for a broad class of closed polygons — regular, irregular, or star-shaped — to function as a waveform generator. Second, a hybrid interpolation algorithm enables smooth morphing between polygons with unequal vertex counts, passing through intermediate shapes that have no parametric description. Third, we extend the paradigm to three dimensions: a convex polyhedron rotated about three axes is sliced by a fixed horizontal plane, and the resulting cross-section yields a continuously variable polygon controlled by the solid's orientation. The system runs in RNBO (Cycling '74) with a geometry caching strategy that avoids per-sample recomputation. Antialiasing combines a four-point polyBLAMP correction derived from runtime Bézier tangents with adaptive oversampling, adapting the correction geometrically to general vertex configurations without per-shape analytical derivation.
Using the Distribution Derivative Method to Model Acoustic Musical Instrument Sounds with Polynomial AM-FM Sinusoids
Marcelo Caetano
The oscillatory modes of musical instrument sounds are commonly modeled with time-varying sinusoids. Several estimation methods model quasi-stationary oscillations accurately, yielding a high-quality representation. However, nonstationary oscillations such as attack transients are still very challenging to model accurately. In this work, we propose to model musical instrument sounds with polynomial modulation sinusoids (PMS) estimated with the distribution derivative method (DDM). DDM gives accurate parameter estimations for PMS with arbitrary order, allowing great flexibility in modeling temporal changes inside analysis frames as amplitude and frequency modulations. We used 39 musical instrument sounds to compare DDM objectively against the standard (SM+) and an adaptive sinusoidal model (eaQHM) using time and frequency error measures. We showed that DDM captures more oscillatory energy than SM+ or eaQHM by modeling PMS more accurately. A MUSHRA listening test with 18 selected sounds confirmed that DDM has higher perceptual quality than both SM+ and eaQHM and that DDM is almost perceptually indistinguishable from the original sounds.
Winding Numbers and Monodromy of Vector Bundles over a Circular Buffer
Georg Essl
The Möbius strip is perhaps the most recognizable topological object of general knowledge. It can be described mathematically in various ways including the formalism of line bundles. In this paper we discuss the bundle idea in the context of digital processing over a circular buffer and show how the idea leads to a more general notion known as monodromy, which describes the effect of the space on traversing a circle once. In this formulation, the monodromy of the Möbius strip is an orientation inversion characterized by a change in sign. This in turns leads to the concept of the winding number, which describes how many windings it takes to return to the original state. We discuss variable monodromy and illustrate that the winding number is robust under this variation. This will allow us to interpret previous disparate results in audio signal processing from Möbius waveguides to chaotic oscillators in delay loops in one unified framework. We close by showing how extending from line to vector bundles opens up the notion of braids to describe monodromy.
Tull
19:00–22:00 Banquet
Presented by Eventide
Symphony Hall
Time Session Location
8:30 Registration (all day) W18 Lobby
9:00
Paper Session 4: Filters and Distortion
Session Chair: Russell McClellan
A Clipping Prevention Method for All-Pass Digital Filters with Time-Varying Coefficients
Federico Fontana, Silvia Pasin, Alberto Bernardini and Stefano D’Angelo
A clipping prevention method is proposed for first- and second-order all-pass filters with time-varying coefficients. Unlike conventional anti-clipping or declipping approaches, the method operates directly on the coefficient dynamics and does not rely on assumptions about internal energy evolution, by just asking that the input signal is not already clipping. The core idea is to control the deviation between the output of the time-varying filter and that of an equivalent static all-pass structure with constant coefficients. By adaptively limiting this deviation at runtime, the output is constrained below a prescribed clipping threshold (typically unit magnitude). The method is active only during short transients where clipping would occur, after which the coefficients are released to reach their target values. This preserves the integrity of the input signal and the numerical properties of the all-pass filter. Experimental results confirm the expected behavior even in scenarios where energy-preserving all-pass structures exceed the clipping threshold, suggesting the proposed approach as a practical solution for robust dynamic filter implementations with limited additional computational cost, suitable especially for embedded digital audio processing hardware.
A DDSP Framework for Adaptive Room Equalization
Fernando Marcos Macías, María Pilar Daza Llin, Mateo Cámara and José Luis Blanco Murillo
Adaptive room equalization remains challenging under time-varying acoustic conditions and complex excitation signals, such as music. In these scenarios, classical filtered-x least mean squares (Fx-LMS) methods falter due to their rigid formulation. We present a modular differentiable digital signal processing (DDSP) framework for closed-loop adaptive room equalization that recovers Fx-LMS as a special case through automatic differentiation. The framework supports interchangeable EQ structures, response estimation methods, loss functions, and optimizers. Experiments with time-varying measured room impulse responses show that frequency-domain objectives provide more stable adaptation than time-domain objectives in the considered scenarios. Relative to the non-equalized response, system distance is reduced by 70% and mel-spectral distance by 13% (worst-case scenario). We further examine how online room response estimation accuracy and frame length affect the trade-off between responsiveness and convergence stability. Overall, the framework provides a unified open-source basis for exploring synergies between classical adaptive filtering and DDSP-based optimization.
Exploring Parallelism and Energy Efficiency in a Multistage Linear-Phase Octave Filter Bank
Jose M. Badia, Jose A. Belloch and Vesa Välimäki
This paper presents a high-performance and energy-aware implementation of a multistage linear-phase octave filter bank for edge system-on-chip (SoC) platforms. The algorithm relies on a cascade of stretched FIR filter stages and complementary band splitting to preserve linear phase across all outputs. While effective, mapping such structures to embedded multicore CPUs introduces significant challenges regarding state management, task synchronization, memory-traffic efficiency, and energy-aware execution. These issues are especially relevant in block-based edge-audio processing, where high throughput must be balanced against the power constraints of mobile and embedded devices. We derive a cache-friendly sequential realization using a blocked streaming schedule and compact circular state. Building on this, we propose a parallel design based on an OpenMP task pipeline with explicit dependencies to preserve the filter-bank semantics without fine-grained synchronization in the filtering tasks. Experimental results on an NVIDIA Jetson Orin Nano module show that the optimized sequential version already sustains more than 1.18 M samples/s, while the task-level pipeline reaches speedups above 4.5× for suitable block sizes. Furthermore, our analysis reveals a clear trade-off between throughput and power, showing that the most energy-efficient operating point does not necessarily coincide with maximum performance on multicore edge SoCs.
PolyADAA: Improving Aliasing Reduction in Memoryless Nonlinearities Using Lagrange Interpolation and Polynomial Approximation
Leonardo Gabrielli and Stefano Squartini
Reducing the aliasing of nonlinear functions is an important problem in digital signal processing. The introduction of the Antiderivative Antialiasing (ADAA) method brought many benefits and is an active area of research. The current bottleneck in terms of aliasing reduction is the initial conversion from discrete- to continuous-time, which was previously done by linear interpolation. In this paper we derive PolyADAA, a method for computing the ADAA output when this conversion is done using higher order Lagrange interpolation. To obtain a viable solution, the nonlinear function is approximated using Chebyshev polynomials, which enable the ADAA integral to be computed. The paper provides numerical examples to show the effectiveness of the approach and discusses the advantages of the method.
Alias-Free Oscillator Synchronization via Additive Synthesis
Jonas Roth, Domenic Keller, Oscar Castañeda and Christoph Studer
Oscillator synchronization is a widely used sound-synthesis technique, but straightforward digital implementations suffer from aliasing artifacts. This paper presents an alias-free method for digital emulation of oscillator synchronization of arbitrary periodic waveforms based on additive synthesis. Starting from a finite set of Fourier-series coefficients representing a bandlimited free-running waveform, we derive linear spectral-resampling transforms that map these coefficients to those of the bandlimited synchronized waveform. Beyond conventional hard synchronization, the proposed approach also supports two additional soft-synchronization modes. To address the high computational complexity of the proposed method, we introduce HASY, a 6 mm² application-specific integrated circuit (ASIC) fabricated in 65 nm CMOS technology. HASY generates one 96 kHz, 24 bit alias-free synchronized waveform with up to 512 harmonics and computes the spectral-resampling transform within only five audio-sample periods.
Evaluating AI Coding Assistants in Audio DSP Education: A Small Scale Study
Leonardo Gabrielli, Michael Fioretti and Giuseppe Bergamino
Recent advances in AI-assisted coding tools raise questions about how programming-intensive subjects such as audio digital signal processing should be taught and how exam projects should be evaluated. This paper presents a small-scale controlled exploratory study conducted in a graduate course on music DSP. As a final project at the end of the course, the students implemented a modular synthesizer plugin in C++. Half of the students had access to AI-assisted coding support, while the other group developed the plugin manually. All students had to follow a protocol and provide data at the end of the project, together with their code, which was discussed with them as part of the course exam. Although the scale of the study is small, the paper shares a qualitative analysis of the results and a few takeaway messages for future reference among lecturers in the field. Overall, AI-assisted coding does provide some advantage to students but only in certain regards. The used AI tools, trained on GitHub repositories, seem to have only partial awareness of the state of the art in digital audio processing (e.g. antialiasing oscillators, virtual analog filters, etc.). Finally, the use of AI seems to not interfere excessively with the ability of the students to learn from their practical experience.
Tull
10:30 Coffee Break W18 Lobby
11:00
Keynote 2
Dina Pearlman-Ifil
Tull
12:00 Lunch W18 Lobby
13:00
Paper Session 5: Sound Design and Effects
Session Chair: Shahan Nercessian
SCAPES: Semantically Conditioned Autoregressive Prior for Environmental Sounds
Esteban Gutiérrez, Lonce Wyse, Frederic Font and Xavier Serra
This paper presents SCAPES, a semantically conditioned autoregressive prior for environmental sound generation. The system models discrete audio representations using an autoregressive architecture conditioned on semantic information, enabling the generation of environmental sounds that follow user-specified concepts. By learning a prior over audio tokens, SCAPES combines high-level semantic control with detailed temporal modeling. Experimental evaluation investigates the quality, diversity, and semantic consistency of generated sounds, demonstrating the potential of autoregressive priors for controllable environmental sound synthesis.
WildFX: A DAW-Powered Pipeline for In-the-Wild Audio FX Graph Modeling
Qihui Yang, Taylor Berg-Kirkpatrick, Julian McAuley and Zachary Novack
We present WildFX, a digital-audio-workstation-powered pipeline for modeling audio-effects graphs from in-the-wild audio. The system uses a DAW environment to construct, render, and evaluate effect-processing graphs, enabling research on realistic effect chains beyond isolated processors or synthetic training settings. WildFX supports the analysis and reconstruction of complex audio transformations by combining flexible plugin routing with data-driven modeling. The pipeline is designed to facilitate scalable dataset creation and experimentation with effect graph inference, parameter estimation, and audio transformation in practical production contexts.
FoleySet: A Multi-Level Human-Annotated Foley Sound Dataset
Sunshiyu Wang and Alexander Lerch
We introduce FoleySet, a human-annotated Foley sound dataset designed to support research on sound-event understanding and Foley sound generation. The dataset provides annotations at multiple levels of granularity, capturing both broad event categories and more detailed semantic or production-related attributes. This multi-level structure supports tasks such as classification, retrieval, captioning, and controllable generation. FoleySet is intended to address the limited availability of systematically annotated Foley material and to provide a common resource for evaluating models across different levels of semantic detail.
Quality Audio Prototyping: A Prototype System for Unified Sound Retrieval and Procedural Generation
Nelly Garcia, Aditya Bhattacharjee, Gabryel Mason-Williams, Israel Mason-Williams, Emmanouil Benetos and Joshua Reiss
This paper presents Quality Audio Prototyping (QAP), a unified prototype system for sound retrieval and procedural generation. The system is designed to support rapid exploration of sound effects through a common interface that combines retrieval from existing audio collections with controllable procedural synthesis. By bringing these two paradigms together, QAP allows users to search for recorded sounds, generate new material, and iteratively refine results within a single workflow. The prototype emphasizes usability, extensibility, and practical sound-design applications, providing a foundation for future work on integrated retrieval and generation systems.
Audio-to-Audio via Diffusion Warm Initialization
Cristobal Andrade and Sebastian Jiro Schlecht
In this paper, we propose diffusion warm initialization as a simple yet effective approach for a range of audio-to-audio transformation tasks. To illustrate the generality of the approach, we demonstrate its use in timbre transfer, MIDI-to-Real synthesis, and multiple audio enhancement tasks. We conduct a detailed empirical analysis on timbre transfer to investigate the role of the initialization time t_init. The effect of t_init is evaluated using pitch-based Jaccard Distance and Fréchet Audio Distance to quantify faithfulness to the input signal and alignment with the target distribution. Our results provide practical guidance for selecting t_init and show that, once properly chosen, a single pretrained diffusion model combined with warm initialization can support multiple transformation objectives without task-specific training or conditioning. Despite its simplicity, this approach already achieves competitive results when compared with more complex pipelines designed specifically for these tasks.
Tull
14:15 Poster Craze 2 (1-minute poster previews) Tull
14:30 Poster/Demo Session 2 + Coffee Break
Posters
SEND: A Spatial Event Neural Detector for Intentional Object Motion in Immersive Music Mixing
Xu Gan, Linhao Zhao and Zhenhai Yan
Deciding exactly when to move audio objects in immersive mixes is a labor-intensive artistic task. Current tools react strictly to instantaneous frequency overlaps, lacking the macroscopic awareness required for musically intentional spatial transitions. To model these decisions, we propose SEND (Spatial Event Neural Detector). Its dual-stream architecture analyzes the target track against its background context, combining a Spec-TNT backbone and a Temporal Convolutional Network (TCN) to capture hierarchical spectral features and precise rhythmic cues. Their dynamic interplay is modeled via a novel Cross-Track Gating Interaction (CTGI) mechanism.
Gauss Circle Lattices with Geometric Convolutions for Synthesizing High Dimensional Image-Source Room Impulse Responses
Yuancheng Luo
The image-source model (ISM) is a widely adopted method for efficiently simulating acoustic room impulse responses (RIRs) under specular reflection assumptions. Acoustic paths between source and receiver are traced to lattice points computed from successive reflections over bounding planes of the room. Rectangular rooms bound the total number of image-sources to be polynomial in the RIR's duration or distance k equivalent, with degree equal the number of room dimensions N. Direct ISM simulations are therefore compute upper-bound by O(k^N), and consider only cases of N≤3 for tractability and real-world applications. This work proposes an alternative computational method that lowers the asymptotic compute bound to O(Nk² log k) for integer coordinates and room dimensions via reducing ISM lattice point counting to the classic Gauss circle problem (GCP). We extend the lattice counting model to frequency-dependent and reflection weighted image-sources in higher dimensions, relating solutions between successive dimensions via the convolution operator. Two constructions for realizing RIRs are presented, along with time-frequency controls, error and run-time analysis, and RIR statistics.
PAEDB: A Synthetic Primary-Ambient Dataset Generation Pipeline for Automatic Upmixing Using Deep Neural Networks
Nicholas Tong and Tom Collins
Automatic blind upmixing aims to convert audio from a smaller channel format (e.g. mono or stereo) into a multichannel format using estimates of direct and diffuse spatial statistics within the signal. Current approaches rely on primary-ambient extraction (PAE) algorithms, which lack real-world context through limited processing windows. Deep learning music source separation (MSS) models have been applied in voice-primary-ambient extraction (VPA) upmixing systems for handling direct components, but still rely on DSP methods of surround channel generation. This work further investigates utilizing source separation within VPA upmixing, focusing specifically on the task of stereo decorrelation and ambience extraction for 5.1 surround. We also release PAEDB (Primary–Ambient Extraction Dataset), a high-quality music dataset derived from MUSDB18-HQ and MoisesDB, comprising 1,809 primary–ambient stem pairs totaling over 550 hours of audio. The performance of selected DNNs trained on PAEDB is then evaluated using signal metrics and a listening study. Our findings indicate that DNNs can effectively model the behavior of PAE algorithms, establishing PAEDB as a strong foundation for ML upmixing systems and underscoring the need for higher-quality multichannel data to advance beyond conventional methods.
A Production-Oriented Framework for Evaluation of SFX Generation
Mélodie Desbos, Yara Bahram, Eric Granger and Mohammadhadi Shateri
Industrial sound design requires audio generation systems that not only produce realistic audio, but also preserve the perceptual identity of a reference, support controllable variation, and remain efficient for practical workflows. Existing evaluations are usually tied to text-to-audio (TTA), unconditional, or task-specific settings, limiting assessment for reference-guided sound effects (SFX) variation. To address this gap, we present a production-oriented evaluation framework for structured comparison of heterogeneous audio generation and editing methods. Our framework identifies nine production requirements and explicitly accounts for differences in model capabilities, enabling comparison under a common production objective. A two-stage protocol is introduced: (1) a reference-guided audio-to-audio (ATA) variation task, in which all methods are evaluated under the same ESC-50 SFX adaptation setup, and (2) capability-specific analyses of native operations such as SFX morphing, temporal and energy alignment, inpainting, and targeted editing. This framework combines objective metrics (including FAD, ImageBind-based reference alignment, and diversity across generated variants), together with a human study of perceptual identity preservation and transient diagnosis. Our study reveals complementary strengths and trade-offs across baselines for different production needs. Among the full-generation baselines evaluated under a shared ATA setting, AudioX provides the strongest overall trade-off between reference alignment and diversity while still supporting SFX morphing. Other baselines remain most suitable for specific editing operations. Our framework establishes a structured evaluation and decision protocol for reference-guided SFX variation and provides a practical basis for designing future unified industrial audio generation pipelines. Audio demos and further details can be found on the accompanying web page.
Sound Effects Dataset Unification With the Universal Category System
Jun Woo Beck and Alexander Lerch
Sound effects (SFX) datasets and libraries often employ distinct tagging schemes, taxonomies, and metadata structures. This creates challenges for research on SFX classification and generation because incompatible taxonomies lead to siloed datasets that might require individualized approaches, result in non-comparable outcomes, and prevent data merging strategies. We propose a modular dataset relabeling framework that adopts the Universal Category System (UCS), an industry-standard hierarchical taxonomy for sound effects, as a shared structural foundation. This open-source framework enables us (i) to convert tags of existing datasets to UCS with a rule-based multi-stage pipeline and conflict resolution to achieve high automatic conversion rates, (ii) to suggest a stratified dataset split for the new labels, and (iii) to combine multiple datasets. To showcase the practical utility, we introduce the EnvSound-UCS dataset, a publicly available unified UCS-compliant dataset of environmental sounds with 58,057 sound clips from three sources: AudioSet, FSD50K, and ESC-50.
Demos
Neural Morphing: Sequence-Optimized Token-Level Morphing in Neural Audio Codecs
Emmanouil Karystinaios
Neural audio codecs were originally developed for high-fidelity compression; however, their latent token representations and expressive decoders also constitute a powerful substrate for controllable audio transformation. This work introduces Neural Morphing, a training-free token-domain audio effect that selects residual-vector-quantized (RVQ) token grains from a user palette and decodes the edited stream through a pretrained codec. The method combines an RVQ-group transfer policy that separates coarse, middle, and fine codebook groups with a continuity-constrained sequence matcher that replaces independent greedy selection with bounded beam search. The intended output is a controlled hybrid: the source preserves rhythmic organization while the palette contributes timbral color and residual detail. We focus on the implementation and real-time behavior of a deployable VST3/AU system, including chunked rendering, palette-size scaling, and backend health checks.
From Arbitrary Audio to EDM: Audio-Conditioned Retrieval of Discrete Rhythm Archetypes
Lindsey Pietrewicz
We present a system for transforming arbitrary audio into Electronic Dance Music (EDM) drum patterns while preserving the timbral identity of the source material. A Vector Quantized Variational Autoencoder (VQ-VAE) trained on 7,999 EDM drum loops learns a discrete codebook of 256 rhythm archetypes, validated through UMAP and hierarchical clustering to exhibit semantically meaningful structure. At inference, spectral features extracted from arbitrary input audio select the nearest archetype via nearest-neighbor retrieval in a shared audio feature space. A training sample from the selected archetype is reconstructed through the VQ-VAE, and a second decoder predicts per-hit velocity dynamics. The user's sounds are then placed at the reconstructed hit positions, scaled by predicted velocity. Applied to 2,000 files from the ESC-50 environmental sound dataset, the system activates 128 of 256 codebook entries (50% coverage), demonstrating broad responsiveness to diverse non-EDM audio.
Compiling Differentiable Audio Graphs to Real-Time DSP
Facundo Franchino and Sebastian Jiro Schlecht
Differentiable audio processors are habitually designed and optimised in machine-learning frameworks, but deploying them as real-time audio effects still often requires non-automatic implementation in a dedicated digital signal processing language. The translation is error-prone, demands an onerous verification process, and detaches research prototypes from usable production tools. That being so, we present ADAC, a compiler that lowers a trained model to a framework-agnostic intermediate representation and emits efficient FAUST code whose impulse response matches the source model to within floating-point arithmetic noise, direct paths included. The optimisation loop is made audible by replacing the model in a running plugin after each gradient step. The exported processor carries a small set of macro-controls that leave its stability intact. A stability certificate computed from the shipped parameters is checked before the plugin is built. At the demonstration, a feedback delay network is trained and exported to a working plugin.
Diffusion-Based Music Audio Editing System Using Differentiable Digital Signal Processing Mixture Model
Kengo Takemoto, Tomohiko Nakamura and Hiroshi Saruwatari
This paper proposes a music audio editing system that enables source-wise editing of harmonic instrument mixtures without explicit source separation. It builds on our previously proposed score-informed method for estimating source-wise synthesis parameters, i.e., time-varying controls used to synthesize each source, such as fundamental frequency and loudness. The method directly estimates these parameters from a mixture signal and the corresponding musical score in an analysis-by-synthesis framework. Using the estimated parameters, the proposed system allows users to edit individual sources by modifying note sequences and instrument types, and then re-synthesizes the edited mixture. Through demonstrations on two-instrument mixtures, we show that the system supports note-level phrasing modification and instrument conversion of selected sources.
CLEAN2FX: Label-Conditioned Modeling for Clean-to-Effect Guitar Audio Transformations
Oliverio Bombicci Pontelli and Iran Roman
We present Clean2FX, a study and demo of label-conditioned clean-to-effect transformation for electric guitar audio. Given a clean guitar input and a target effect label, the task is to synthesize the corresponding effected signal while preserving the musical content. Training and evaluation pairs are constructed from EGFxSet real, single-tone recordings by assembling matched clean/effected chords, melodies, and mixed timelines. This allows for controlled comparison across effects. We evaluate four neural approaches under a common spectrogram-based transformation setting: two variational autoencoders and two U-Net models that differ in whether they operate on linear or log-magnitude representations. Performance is measured using linear-magnitude spectrogram MSE and Fréchet Audio Distance. The U-Net models outperform the variational autoencoder variants. Per-effect results show that distortion effects are most readily improved, whereas delay and reverb effects exhibit weaker FAD gains despite substantial spectral-error reductions. A conditioning-sensitivity diagnostic provides evidence that the best model responds to target labels rather than collapsing to a single transformation. Our demo website compares two models applied on real-world guitar performances outside training and validation data, providing audio and spectrogram examples of the practical clean-to-effect behavior.
Evaluating Tokenization Strategies for Expressive Classical Piano Performance Generation
Qingyang Lyu, Brian Cruz, Jeremy Wagner and Carmine-Emanuele Cella
Expressive piano performance generation needs symbolic pitch, timing, and dynamics. We evaluate six tokenization strategies for a Transformer that generates classical piano performances. Our tokenizations add velocity, beat annotations, and sustain pedal, from note-only to full representations. We pretrain on MAESTRO, then finetune on ASAP with beat-level annotations. The model uses anticipatory-style note encoding with cross-attention on composer and genre. FAD on the ASAP test set shows that note + velocity + pedal and full modes achieve the lowest mean FAD (1.76 and 1.97). Both beat the note-only baseline (3.10). Beat tokens show mixed, category-dependent effects and do not improve the best modes on average.
VoiceFX: CLAP-Based Audio Quality Improvement for Singing and Speech
Elena Georgieva
This project introduces an automatic method for enhancing audio quality in singing and speech. Using recordings from the LibriSpeech and Smule DAMP dataset, I applied a set of degradations and tested a set of audio effect "remedies" designed to reverse them: a high shelf filter, de-esser, noise reduction, and high-pass filter. I used the CLAP (Contrastive Language-Audio Pretraining) model to estimate recording quality and recommend remedies by comparing audio clips to descriptive text prompts in the shared embedding space. To evaluate my method, I conducted a large-scale listener study with 234 participants and 4,600 ratings. While CLAP encoded some relevant information of vocal recording quality, it often favored remedies like noise reduction while listeners preferred the original clips, suggesting that perceptual artifacts introduced by enhancement may not be captured by CLAP. My findings underscore the value of human judgment: embedding models can guide enhancement, but perceptual validation remains valuable. Audio examples are available online on the DAFx demo website.
FPGA-Enabled Real-Time Audio Sampling, Processing, and Recording for an Electronic Drum Set
Matthew Taylor and Mark Rau
Processing and recording multitrack audio from an electronic drum set is demanding of computational power and hardware resources. In this paper, we present a complete musical instrument system capable of up to 16-channel percussion sampling, processing, and recording, all in real time. The system leverages a field programmable gate array (FPGA) for parallel audio processing and includes audio effects such as pitch shift, delay, reverb, distortion, a virtual analog low-pass filter, and bit crush. The FPGA also provides interfaces for other system hardware, including an Ethernet audio interface and various audio effect control interfaces. The final design has a cost of under $500 and utilizes about half of the hardware resources on an entry-level FPGA, providing a future platform for more advanced percussion synthesis using real-time physical modeling.
L-BOW: Gesture-Driven Digital Audio Effects for Augmented Violin in a Unified Csound Environment
Jinlan Zhao and Richard Boulanger
Live performance leaves little room for a sensor pipeline that misfires; when a gesture fails to map correctly to an intended effect, the error is immediately audible. This paper presents L-Bow, a wrist-worn six-degree-of-freedom (6-DoF) inertial measurement unit (IMU) controller for augmented violin performance. By removing intermediate software layers, L-Bow integrates gesture sensing, six performance modes, and a shared digital effects chain within a single, self-contained Csound file. This is achieved using Csound's native arduinoRead opcode for direct serial communication rather than an external Python–OSC bridge. The paper discusses the architecture of this system and its implications for designing dependable, low-maintenance interactive digital audio effects.
InstructFX2FX: A Multi-Turn Text-to-Effect System for Sequential Audio Effect Refinement
Song-Ze Yu, Milan Liessens Dujardin, Yuxuan Cai, Wantong Zhang, Brian Cruz, Jeremy Wagner and Carmine-Emanuele Cella
We present InstructFX2FX, a system for sequential audio effect refinement through multi-turn natural-language instructions. Existing text-to-effect systems are largely single-shot, mapping one textual descriptor to one preset. Real audio engineering is instead sequential: engineers refine an existing effect chain through successive instructions. This poses a stateful problem that single-shot systems do not address: given the current effect parameters state and a new instruction, update the sound while preserving what earlier instructions already achieved. InstructFX2FX addresses this with a hybrid architecture that divides labor between a language model and CLAP-guided optimization. The LLM serves as a high-level planner that selects effects and proposes the initial parameter state, motivated by recent evidence that LLMs can outperform CLAP-based optimization for single-turn text-to-effect mapping; CLAP-guided optimization then refines the existing parameter state, providing a more stable and robust refinement mechanism than LLM reprompting. In the demo, attendees drive a dry recording through successive natural-language instructions: after each turn, they choose how strongly the effect is applied, then issue the next instruction based on what still differs from the sound they intend. In a preliminary evaluation on SocialFX-derived descriptor pairs, CLAP-guided refinement achieves lower DSP-feature MMD than an LLM+LLM initialize-then-reprompt baseline on 9 of 10 pairs. Trajectory analysis further shows that, for differentiable effects, optimization tends to gradually move the audio toward the new target while retaining the effects of the previous instruction, highlighting the potential for gradual refinement. Audio demo and source code are available online.
Keyframe Audio via Extrema Sampling
Matthew Nielsen
Overlap-add (OLA) is the simplest approach to audio time stretching. Methods like the phase vocoder (PV) and waveform-similarity OLA (WSOLA) offer higher quality results but require operations like the FFT or cross-correlation. On low power embedded hardware, this cost adds up quickly. We present a content-adaptive OLA method, an order of magnitude cheaper than PV or WSOLA, whose dominant artifacts are added saturation and some spectral contrast loss. Our method reduces uniformly sampled signals to sets of timestamped local extrema, a sparse representation where the distance between points encodes the signal's information density directly into the buffer. In OLA, the crossfade duration is fixed, but no one value suits both transients and sustained sounds. We use the extrema density to inform the crossfade duration, adapting it to the signal's local content on a sample-by-sample basis. We compare our method against OLA, WSOLA, and PV using objective metrics and a listening test. Our method coherently stretches audio, preserving transients across a wide range of stretch ratios and capturing dense, layered material cleanly.
W18-1202 + W18 Lobby
16:00
Paper Session 6: Reverb 2
Session Chair: Anthony Gallien
A Unified Framework for Real-Time Concatenation-Driven Convolution
Niccolo Abate and Brian Hansen
This work introduces a novel framework for Concatenation-Driven Convolution (CDC), unifying concatenative synthesis and real-time convolution into a single integrated audio processing paradigm. While concatenative synthesis has traditionally been used for corpus-based sound generation and convolution has served as a largely static filtering technique, the proposed approach reconceptualizes impulse responses (IRs) as dynamic, navigable sonic material. In the CDC framework, a corpus of audio segments is analyzed using perceptual features and organized via a self-organizing map (SOM), enabling intuitive, gesture-based traversal of a structured timbral space; the resulting concatenative output is treated as a continuously evolving impulse response and injected directly into a partitioned convolution engine. Its central technical contribution is single-engine frequency-domain kernel interpolation: rather than crossfading the outputs of two convolution engines, the FFT-domain kernels of the current and target IRs are interpolated within a single engine, preserving the internal convolution state across IR transitions and avoiding the warm-up energy loss inherent to dual-engine crossfading.
Shimmer Reverberation with Nonlinear Feedback Delay Networks
Gloria Dal Santo, Xiaojie Pi, Karolina Prawda, Sebastian Schlecht and Vesa Välimäki
Shimmer reverberation is an effect used in music production to deliver ethereal, pitch-shifted textures and evolving ambient soundscapes. This paper explores the synthesis of shimmer effects using the feedback delay network architecture, a popular real-time reverberator. We propose five distinct approaches for integrating nonlinear and time-varying operations into the feedback loop, focusing on expanding the harmonic content while adhering to energy-preservation and stability criteria. Our approach can generate a wide range of sonic characteristics, from harmonically rich distortions to musically coherent pitch-shifted reverberation, while maintaining stability and controllable decay behavior.
Fast Parametric Matrices for Lossless Feedback Delay Networks
Andrea Coppola
This paper presents a framework for designing creative reverbs using parametric orthogonal feedback matrices on Feedback Delay Networks (FDNs) through recursive Kronecker products of 2D rotation and reflection matrices. By parameterizing each 2×2 kernel with a single angle, we construct a family of 2M×2M orthogonal matrices that maintain losslessness while enabling continuous control over network topology. We then exploit their recursive definition to compute the feedback operation with an O(N log₂ N) divide-and-conquer algorithm that matches the Fast Walsh-Hadamard Transform time complexity while offering parametric flexibility. Strategic manipulation of individual kernel angles enables creative sound design applications, such as stereo cross-coupling, selective freeze, and time-varying modulation for resonance breaking.
Gradient Descent Optimization of Room Impulse Responses with Parameter-Efficient Differentiable Feedback Delay Networks
Ilias Ibnyahya and Joshua Reiss
Artificial reverberation can be produced either by convolving a signal with a measured room impulse response (RIR) or by synthesizing it with a parametric algorithm such as a Feedback Delay Network (FDN). The former reproduces a captured space faithfully but is costly to run and offers no control over its acoustic properties, while the latter is efficient and editable but hard to match to a specific room. In this paper we bridge the two by fitting a fully differentiable FDN to a measured RIR through gradient descent. The proposed network uses sixteen delay lines at a sampling rate of 48 kHz and trains all of its components jointly, including the delay lengths, the feedback matrix, the early-reflection taps, and a set of attenuation filters that control the frequency-dependent decay.
Group Delay Manipulation for Creative Sound Transformation with the Giant FFT
Ted Apel
The Giant FFT is a single DFT spanning an entire audio file that produces a spectrum encoding the complete temporal evolution of a sound. Creative manipulations in this domain have produced compelling results, but typically smear discrete events into sustained textures by disrupting the temporal relationships between frequency bins. This paper introduces a framework for coherent spectral manipulation in the "group delay domain", where the derivative of the phase spectrum with respect to frequency makes the temporal center of gravity of spectral energy explicit at every frequency bin. By identifying spectral regions around amplitude peaks and grouping them by group delay similarity, spectral features can be displaced in time through uniform modification of their group delay.
Tull
20:00–22:00 Concert
Presented by Soundtoys
Tull
Time Session Location
8:30 Registration (all day) W18 Lobby
9:00–10:00
Paper Session 7: Reverb 3
Session Chair: Alex Tung
A Corpus-Driven Parametric Modal Reverberator
Michele Ducceschi, Leonardo Gabrielli, Riccardo Simionato and Riccardo Russo
A parametric modal reverberator is presented in which synthesis parameters are derived from a large, curated corpus of room impulse responses (IRs). The collected responses are subjected to modal decomposition, yielding per-mode frequencies, damping coefficients, and residue amplitudes, together with a short early-reflection finite impulse response (FIR) filter. From the decomposed data, a feature table is constructed per IR comprising standard acoustic indices, per-band damping and density statistics, amplitude distributions, and FIR descriptors—50 variables in total. Six acoustically meaningful user controls are selected; since these exhibit substantial pairwise correlations across the corpus, they are orthogonalised via principal component analysis (PCA) prior to regression.
Diagonal Complex-Valued State Space Models for System Identification and Modeling of Metal Plate Reverbs
Matthias Bittner, Matthias Wess, Dominik Dallinger, Daniel Schnöll and Axel Jantsch
Accurate and interpretable modeling of plate reverbs remains an important challenge in virtual analog modeling of audio effects. While existing neural network-based black-box approaches already achieve high-quality synthesis and strong perceptual quality, they often lack the possibility to identify the underlying physically meaningful complex, long-memory modal behavior. In this work, we address this limitation by proposing a restricted complex-valued diagonal State Space Model (SSM), showing its equivalence to a parallel second-order all-pole filter, also utilizing efficient training via parallel state computation using the parallel scan algorithm. Additionally, we propose a Matrix Pencil (MP) guided eigenvalue initialization, improving synthesis quality and system identification performance.
Modal Structure of Plate Boundaries and Klein Bottle Reverberation
Jin Woo Lee and Mark Rau
Physical modeling sound synthesis has achieved remarkable success in terms of its fidelity to reality. In many cases, since modeling of the physical system is performed on the sounding objects that already exist in the real world, observation precedes the model itself. Departing from this convention, this paper aims to physically model the acoustic characteristics of objects that do not necessarily exist in reality. Specifically, we study wave propagation on compact two-dimensional (2D) manifolds that are non-orientable surfaces, such as the Klein bottle that cannot be embedded in three-dimensional Euclidean space without self-intersection. We derive closed-form expressions for the eigenfrequencies and mode shapes of non-orientable 2D topologies and study their acoustic characteristics. The modal structures are verified through comparison with finite-difference time-domain simulations. The results demonstrate how the topological character formed by the boundaries influences the acoustic resonances, and how the quotient-space framework provides a practical route to reverb synthesis on geometries with no physical counterpart.
DAFx Challenge Introduction & Results
Leonardo Gabrielli and Michele Ducceschi (Challenge Chairs)
The 1st DAFx Parameter Estimation Challenge is an open initiative to advance the state of the art in parameter estimation for acoustic modeling. Stated as a system identification problem, this first edition focuses on plate reverberation—an archetypal dense, modal and weakly damped acoustic system. Participants tackled two tasks: (A) estimating the physical parameters of a vibrating plate from its impulse response, and (B) recovering the modal parameters of the same system. Both rest on a simulation framework based on the damped Kirchhoff–Love plate equation, and both are posed and scored entirely on synthetic data produced by that framework: no measurement of a real plate is involved. Two participants solved Task A down to machine precision by different strategies: one a neural network trained on a very large dataset, and one gradient-free optimization with many inexpensive evaluations. Task B proved considerably harder: the best submission attains a relative error of 0.33 on a [0, 2] scale, and every method recovers modal frequencies and decay rates far more accurately than modal gains. A complementary frequency-domain evaluation reorders the ranking and exposes a systematic gain bias to which the per-mode metric is blind.
Tull
10:00–11:00
DAFx Challenge Posters + Coffee Break
Non-iterative Modal Parameter Estimation for Plate Reverbs via Matrix-Pencil-Guided State Space Model Initialization
Matthias Bittner, Axel Jantsch
Modal parameter identification for plate reverbs remains a challenging problem in virtual-analog audio effect emulation. Though neural network-based black-box approaches achieve high modeling accuracy, they generally lack interpretability and do not provide access to physically meaningful modal parameters. In this work, we present our solution to Task B of the DAFx Plate Reverb Parameter Estimation Challenge. Our method first estimates the total number of modes and then employs a Matrix Pencil (MP)-guided eigenvalue initialization strategy for a diagonal complex-valued State Space Model (SSM), which can be interpreted as a bank of parallel second-order all-pole filters. Exploiting the linearity of the resulting system, we compute the state impulse responses and replace gradient-based optimization with a closed-form least-squares estimation of the modal gains. The proposed approach enables accurate recovery of the modal parameters while maintaining an interpretable system representation.
ALAMODE: Automated Learning of Acoustical Modal Parameters via Differential Evolution
Jin Woo Lee, Jatin Chowdhury, Facundo Franchino, Soohyun Kim, Mark Rau
This paper is a technical report on the methodology submitted for Task A of the 1st DAFx Parameter Estimation Challenge. The goal of the challenge’s task is to invert the multi-dimensional physical and geometric parameters of a virtual plate reverberator given a target reference impulse response. To achieve this, we present a multi-stage gradient-free optimization framework. This three-stage optimization is computed using an efficient physics-based simulator, starting with an optimization of only mode frequency-determining physical parameters, followed by a 6-DoF parameter optimization with position-determining ones and a final phase for frequency- and position-independent mode amplitude estimation.
Simulation-based Inference Plate Reverberation Inverse Problems
Dylan Sechet, Matthieu Kowalski, Marc Evrard
We address Task A of the 1st DAFx Parameter Estimation Challenge, which aims to retrieve the physical parameters of a plate model from an impulse response. To do so, we use the Simulation-Based Inference (SBI) framework, in which we train a neural network to estimate a density over plate parameters given an impulse response, using a dataset generated by the simulator. Inference for a new impulse response then requires only a forward pass through the network, without involving the simulator. For each test observation, we fine-tune a specific network: additional simulation rounds are performed by sampling parameters from the current estimated distribution, simulating the corresponding impulse responses, and fine-tuning to produce the specialized network.
Transformer-Based Plate Parameter Estimation with Differentiable and Particle-Swarm Refinement
David Marttila, Rodrigo Diaz, Pablo Tablas de Paula, Ilias Ibnyahya, Chin-Yun Yu
We present two Transformer-based methods for Task A of the 1st DAFx Parameter Estimation Challenge, which requires estimating six effective physical parameters of a synthetic plate-reverb model from its impulse response (IR). Method A1 combines an Audio Spectrogram Transformer encoder and Transformer regressor with differentiable IR refinement. Method A2 uses the same encoder to condition a continuous normalizing flow and refines sampled candidates using particle swarm optimisation (PSO) and gradient polishing. Both methods preserve the absolute IR scale to recover surface density. On a synthetic holdout set of 100 IRs, both refinement procedures reduce waveform and parameter errors by more than three orders of magnitude relative to the unrefined neural outputs. The PSO-based pipeline achieves the lowest errors, indicating near-perfect recovery in this matched synthetic setting.
Count-Density Networks for Modal Plate Parameter Estimation
Rodrigo Diaz, Pablo Tablas de Paula, David Marttila, Ilias Ibnyahya, Chin-Yun Yu
We describe two submissions to Task B of the 1st DAFx Parameter Estimation Challenge, which estimates an unknown number of modal frequency, decay, and gain triples from a synthetic plate-reverb impulse response. The first method combines pooled spectral features with time-domain and absolute-scale conditioning in a real-valued convolutional count-density network, while the second uses a complex-valued Transformer count-density network. Both methods jointly infer the modal count and per-mode attributes directly from the IR. On an independently generated 100-IR comparison set, the two neural estimators achieve lower overall challenge error than the evaluated classical baselines, with frequency and decay estimation substantially more accurate than gain estimation.
Multi-View Subband Autoregressive Pole Harvesting for Modal Plate Identification
Facundo Franchino, Jatin Chowdhury, Jin Woo Lee, Soohyun Kim, Mark Rau
We describe a Task B submission for the 1st DAFx Parameter Estimation Challenge. A matching-pursuit anchor stage seeds a multi-view subband autoregressive (AR) pole harvester on the raw IR and its first two finite differences; since linear filtering preserves pole locations while changing residues, the three views expose complementary subsets of the same pole set. Bands in which the AR order saturates are recursively split, and any remaining under-resolved region is completed from an IR-derived saturation indicator. Gains are assigned by a global ridge least-squares (LS) fit in the damped-biquad atom convention. The pipeline uses only the unnormalized IR; so it does not use plate parameters, analytical modal-frequency or decay laws, .wav files, or Task A information.
A Multi-Resolution Spectrogram Approach for Estimating the Physical Parameters of a Plate Reverb
Jared Lipkin, Meiying Chen, Benjamin Thompson, David Anderson, Andrea Cogliati, Michael Heilemann
The ResNet-18 image classification model is employed to determine the physical parameters of a plate reverb from a recording of the impulse response. The model is adapted to derive parameters using normalized and down-sampled multi-resolution spectrograms computed from the provided impulse responses (IRs). To refine the prediction of the output location, the spectral phase response is also included as an additional input channel to the network since multiple output locations can give the same magnitude response for high-order resonant modes. On a 5000 IR validation set, our model achieves an average normalized mean squared error (NMSE) of 0.02920 across all parameters, with the lowest average NMSE occurring for parameters yo (0.00228), Ly (0.00347), and xo (0.00574).
Parameter Estimation via Differentiable Modal Plate Synthesis
Filippo Garofalo, Alessandro Antonio Lillo, Alessandro Ilic Mezza, Riccardo Giampiccolo, Alberto Bernardini
We present our submission to Task A of the 1st DAFx Parameter Estimation Challenge, which concerns the estimation of the physical parameters of a vibrating plate from a synthetic impulse response. Our approach introduces a differentiable modal plate synthesizer and estimates the plate parameters through inference-time gradient-based optimization of the synthesizer parameters. The six target parameters are recovered by minimizing a multi-scale spectral loss via backpropagation through the differentiable plate model. To handle the non-convexity of the loss landscape, we adopt a two-phase training strategy consisting of multiple short-term probe optimizations, followed by full-scale refinement initialized from the best candidate. We evaluate the approach on eight impulse responses synthesized with the official challenge dataset generator. Compared with a constant-value predictor and the particle swarm optimization baseline provided by the challenge, the proposed method reduces the prediction error by approximately one order of magnitude.
Accurate Plate Reverb Parameter Estimation Using Two-Stage Evolutionary Search
Byunghoo Park, Jayeon Yi, Takyoung Kim, Minje Kim
We describe our submission to Task A of the 1st DAFx parameter estimation challenge. The task is to recover the six physical parameters of a simulated metal-plate reverberator—its dimensions and material properties—from a single impulse response (IR). We treat this as a black-box optimization: candidate parameter sets are fed to the simulator and scored by a loss against the target IR. The method has two stages. The first uses CMA-ES, an evolutionary optimizer, to recover five of the six parameters, comparing IRs under an amplitude-normalized loss. Amplitude normalization makes the search robust but discards the cue to the sixth parameter, the plate’s surface density; a second stage therefore estimates it alone, with a ternary search on the un-normalized loss. As the choice of loss strongly affects the search, we select it beforehand, and analyze why compression in the common multi-scale spectral loss degrades recovery. Finally, we test our method on a validation set of 50 IRs, discuss a pathological failure mode, and ablate to justify having two different stages instead of a unified CMA-ES search.
Simulation-Based Plate-Reverb Parameter Estimation from a Single Impulse Response ★
Minhui Lu, Joshua Reiss
We present a simulation-trained, non-iterative estimator for Task A of the 1st DAFx Parameter Estimation Challenge. Each unnormalized plate-reverb impulse response is summarized by amplitude, spectral, and decay descriptors, and an ensemble of tree regressors estimates the six target parameters in one pass. Across two independent synthetic validation sets, the normalized models outperform the training-set mean and an earlier raw-regression baseline. On a shared set, the final ensemble also outperforms a single run of the official default PSO at substantially lower inference cost. Since the official labels are hidden, parameter accuracy is measured on simulator-matched data, and the released responses support only audio-side consistency checks. The estimator returns point estimates without uncertainty.
Neural Networks for Physical Parameter Estimation of Plate Reverberation from Impulse Responses ★
Jia-Chang Yang
This paper presents our Task A submission to the 1st DAFx Parameter Estimation Challenge. We use the official ModalPlate dataset generator to synthesize 1000 one-second plate impulse responses with randomly sampled parameters inside the public ranges. A time-domain CNN-GRU regressor then estimates the six official Task A parameters from each unnormalised waveform. The model combines three one-dimensional convolutional blocks with a bidirectional gated recurrent unit and is trained with mean squared error on min-max normalised targets. The generated data are split into 700/150/150 train/validation/test examples, and the test split is never used during training or model selection. The implementation follows the official Task A format and exports evaluation-compatible prediction files for both development evaluation and blind-set submission.
Band-Count Dense Modal Estimation with Fixed-Frequency Differentiable Resonator Refinement ★
Minhui Lu, Joshua Reiss
Task B of the 1st DAFx Parameter Estimation Challenge requires estimating the frequencies, decay rates, gains, and number of modes in a dense plate-reverb impulse response. Weak and overlapping modes make sparse peak detection prone to severe undercounting. We train an ExtraTrees regressor on simulator-generated data to predict mode counts in four frequency bands. These counts define dense frequency grids, after which a differentiable all-pole resonator model refines decay and gain while keeping frequency fixed. On two separate synthetic validation sets, the system reduces a local challenge-style error by about 66% relative to the official default peak-picking baseline. The improvement is mainly associated with lower mode-count mismatch, while decay and gain remain the largest error sources. These findings support separating modal-density estimation from continuous parameter fitting.
Peak-Residual Modal Estimation with Learned Calibration and High-Band Density Correction ★
Doohyun Jung
This paper describes two related submissions to Task B of the 1st DAFx Parameter Estimation Challenge. Both estimate modal frequency, decay, and gain directly from an unnormalised plate impulse response without using plate parameters, the excluded analytical modal-frequency law, or official-test ground truth. The primary system constructs a large candidate pool through prominence-graded spectral peak picking, iterative residual analysis, multi-view consensus, band-wise budgeting, and a learned file-level mode-count target. Raw decay and gain estimates are then corrected by a small mode-wise neural network that is not allowed to move frequencies or change the number of rows. A secondary variant addresses suspected high-frequency under-counting with a separately gated, non-oracle density-fill stage in the 6–10 kHz band. The paper reports development diagnostics, reproducibility information, and descriptive statistics for the 16 official outputs. The two variants expose a deliberate precision–recall trade-off: one preserves a visible spectral justification for every row, while the other tests bounded hidden-multiplicity augmentation in densely overlapped regions.
Physics-Inspired Feature Fusion for Plate Parameter Estimation from Acoustic Impulse Responses ★
Zhenyu Guo, Yun Zhang, Liangming Chen, Wei Liu, Gongping Huang
Estimating physical plate parameters from impulse responses is a challenging inverse problem. Task A of the first Digital Audio Effects Parameter Estimation Challenge requires the recovery of six identifiable parameters from displacement impulse responses. In this work, we propose a physics-inspired feature fusion network (PIFFN) that combines a pretrained convolutional backbone with a 15-dimensional physics-inspired feature vector computed from the impulse response. These physics-inspired features describe amplitude scale, temporal decay, and spectral structure without relying on modal-distribution priors. The proposed model is evaluated on the official validation set, achieving an overall normalized mean squared error of 0.00362. Compared with the official particle swarm optimization baseline and backbone-only model, PIFFN shows a clear performance improvement, demonstrating its effectiveness for plate parameter estimation.
A Dual-Stream Framework Combining Audio Spectrogram Transformer and Dynamic Mode Decomposition for Plate Modal Parameter Estimation ★
Yun Zhang, Liangming Chen, Zhenyu Guo, Wei Liu, Gongping Huang
Plate reverberation is characterized by a dense distribution of resonant modes, which makes the estimation of modal parameters from observed responses a challenging inverse problem. To address this problem, we propose a physics guided dual stream framework that integrates an Audio Spectrogram Transformer (AST) with Dynamic Mode Decomposition (DMD). The AST branch models the global temporal and spectral structure of the impulse response, whereas the DMD branch extracts local descriptors associated with modal dynamics. The resulting representations are fused and processed by convolutional prediction heads to jointly estimate mode presence and the corresponding modal parameters. Experiments on the official validation set of Task B in the DAFx Challenge show that the proposed method reduces the overall relative error from 1.976 for the official baseline to 0.867. These results demonstrate that integrating local dynamic information derived from physical modeling with global transformer based representations substantially improves plate modal parameter estimation.

★ Not presenting in the poster session.

W18-1202 + W18 Lobby
11:00–12:00
Keynote 3
Sean Costello
Tull
12:00 Lunch W18 Lobby
13:00
Paper Session 8: Virtual Analog
Session Chair: Fabian Esqueda
A Comparative Study of Kolmogorov-Arnold Networks and Multi-Layer Perceptrons for Virtual Analog Modeling in Wave Digital Filters
Riccardo Giampiccolo, Enrico Torres, Mauro Giuseppe de Bari, Samuel Limier and Alberto Bernardini
The design of Virtual Analog (VA) algorithms has traditionally been divided between white-box (physics-based) and black-box (data-driven) approaches. Recent work has shown that hybrid methods, combining physical modeling with neural networks, can effectively leverage the strengths of both paradigms. In particular, Wave Digital Filters (WDFs) can be coupled with Multi-Layer Perceptrons (MLPs) to model circuits with multiple nonlinearities in a fully explicit manner. In this paper, we present a comparative study investigating the use of Kolmogorov-Arnold Networks (KANs) for VA modeling within the WDF framework. Unlike MLPs, KANs shift the learning paradigm by parameterizing activation functions instead of relying exclusively on learned weight matrices, potentially enabling more compact representations. Results show that, for our case study, KANs achieve accuracy comparable to MLPs while requiring approximately 70% fewer parameters at the cost of increased computational complexity. These findings suggest that KANs may represent a promising alternative in scenarios where memory footprint is a primary constraint, such as embedded audio applications, or when target models feature numerous nonlinear elements.
Performance-Oriented Wave Digital Circuit Emulation
Jatin Chowdhury and Mark Rau
Wave Digital Filters are a circuit-modeling paradigm well-suited for reusable software implementation, but existing software implementations often incur significant overhead due to run-time abstractions and data layout constraints. This paper presents a performance-oriented toolchain for implementing Wave Digital circuit models based on static code generation. The toolchain consists of a declarative circuit description language, a compiler that generates circuit simulation code with minimal persistent state and no run-time abstraction, and a minimal runtime library implementing specialized circuit components as Wave Digital Filters. Performance measurements across several test circuits demonstrate that the generated models consistently outperform existing implementations, and achieve near-ideal performance relative to a theoretical execution bound.
Stability Analysis of Time-Varying Virtual Analog Filters
Russell McClellan
Time-varying virtual analog filters used in digital audio effects and synthesizers are often implemented by discretizing continuous-time state-space systems using trapezoidal integration. When filter parameters such as cutoff frequency or resonance vary over time, as is the norm in musical applications, proving BIBO stability of the resulting time-varying discrete-time system becomes nontrivial. In this paper, we review the technique of common quadratic Lyapunov functions (CQLFs) from the control systems literature and show how a continuous-time CQLF is preserved through discretization. This allows us to prove stability of some time-varying virtual analog filters by working in the often simpler continuous-time domain. We apply this framework to several filters of musical interest, providing new proofs of stability for the state variable filter and the Sallen-Key filter, and new bounds on the stable time-varying parameter range for the Moog ladder and diode ladder filters.
Quantifying Nonlinear Behavior in Digital Moog Ladder Filters: Cross-Implementation Comparison and Common-Core Ablation
Hiroyuki Oyama
Digital Moog ladder filters are often compared by their linear frequency response, even though musically important differences emerge under nonlinear operation. This paper presents an open, reproducible SPICE-referenced evaluation framework that combines two components: a cross-implementation comparison of five digital ladder-filter implementations against the same SPICE reference, and a controlled common-core ablation. The cross-implementation comparison shows that close linear agreement can mask substantial nonlinear divergence, while the ablation shows that ladder nonlinear behavior depends not only on saturator choice and nonlinearity placement, but also on the non-uniform contribution of different ladder stages: in partial ablations of a SPICE-referenced TPT/ZDF core, retaining earlier-stage nonlinearities preserves harmonic behavior better than retaining later-stage nonlinearities alone. Beyond the specific comparisons reported here, the framework provides a reproducible basis for evaluating fidelity–cost trade-offs when nonlinear structure must be simplified.
Residual-Driven Adaptive Multi-Rate Quadratic Programming Framework for Nonlinear Analog Audio Circuit Emulation
Miguel Zea and Luis A. Rivera
This work extends our previously proposed Quadratic Programming (QP) approach for the emulation of nonlinear analog audio circuits by formalizing its main numerical ingredients and introducing a residual-driven adaptive multi-rate scheme. Starting from a state-space Differential Algebraic System of Equations (DAE) formulation, the nonlinear algebraic circuit device relations are replaced inside the QP by a first-order surrogate linear constraint, and the post-step nonlinear residual is shown to act as a valid defect indicator for adaptive step-size control. This yields a single-step simulation procedure that avoids the usual combination of nonlinear iterative solves and separate integration updates. The method is evaluated on a diode clipper, a BJT common-emitter amplifier, and a Colpitts oscillator, using SPICE as a baseline reference. The results show that adaptive step sizing considerably improves agreement with the reference solution, that the pseudo-inverse implementation is essentially equivalent to the full equality-constrained QP in the tested cases, and that the proposed formulation remains effective beyond the baseline clipper example, including for a self-oscillating circuit. These results position the proposed method as a promising bridge between SPICE-like interpretability and the efficiency demands of virtual analog (VA) audio applications.
Evaluating Dynamic Range Compressor Models Using Control-Voltage Measurements: An Approach and Dataset
Benjamin Thompson and Michael Heilemann
The quantity that defines the behavior of a dynamic range compressor is the time-varying gain applied to the signal as a function of the input level. However, models of these devices are typically evaluated using proxy metrics because isolating the gain reduction signal from the audio input–output data included in existing datasets creates an ill-conditioned inverse problem. It is unclear how accurately these metrics describe the behavior the model is tasked with emulating, particularly as waveform-based metrics can be influenced by secondary effects introduced by analog processing and capture, even when those effects are inaudible. We investigate a method of evaluation in which the gain-reduction signal produced by a model is measured directly against a gain-reduction control voltage signal produced by the hardware. To evaluate the efficacy of this metric as a learning objective, a gray-box model is trained using loss computed directly over the gain control signals alongside two models trained using common proxy losses. The models trained using proxy losses did not achieve parity with models trained directly on the gain control signal when evaluated with respect to the underlying control trajectory, and the waveform-domain metrics assigned similar errors to models that were clearly separated by the direct metric. To facilitate further exploration of this method of evaluation, we present a Solid State Logic bus compressor dataset that includes the gain control voltage signal captured alongside the audio output.
Tull
14:30 Coffee Break W18 Lobby
15:00
Paper Session 9: Audio Coding, Representation, and Benchmarking
Session Chair: Jatin Chowdhury
Keyframe Audio via Extrema Sampling
Matthew Nielsen
Overlap-add (OLA) is the simplest approach to audio time stretching. Methods like the phase vocoder (PV) and waveform-similarity OLA (WSOLA) offer higher quality results but require operations like the FFT or cross-correlation. On low power embedded hardware, this cost adds up quickly. We present a content-adaptive OLA method, an order of magnitude cheaper than PV or WSOLA, whose dominant artifacts are added saturation and some spectral contrast loss. Our method reduces uniformly sampled signals to sets of timestamped local extrema, a sparse representation where the distance between points encodes the signal's information density directly into the buffer. In OLA, the crossfade duration is fixed, but no one value suits both transients and sustained sounds. We use the extrema density to inform the crossfade duration, adapting it to the signal's local content on a sample-by-sample basis. We compare our method against OLA, WSOLA, and PV using objective metrics and a listening test. Our method coherently stretches audio, preserving transients across a wide range of stretch ratios and capturing dense, layered material cleanly.
Robust Recovery of Deterministic Timecode Signals Under Analog Degradation
Brady Cruse
This paper presents a controlled evaluation framework for recovering deterministic stereo timecode waveforms used in digital vinyl systems. Three lightweight decoding policies are compared: Fixed-Threshold Symbol Decoder (FTSD), Adaptive-Threshold Symbol Decoder (ATSD), and State-Constrained Temporal Decoder (SCTD). The work provides matched comparison and explicit separation of availability, reference-relative correctness, and temporal behavior across noise, dropout, and clipping conditions.
Real-Time Neural Audio on Apple Silicon: Benchmarking Inference Frameworks Under Realistic DAW Contention
Dharanipathi Rathna Kumar Balasubramaniam and Joseph Timoney
Neural network models are increasingly deployed in audio plugins across a wide range of applications, including amplifier emulation, effects modeling, and synthesis. This paper evaluates widely used inference options including BNNSGraph, RTNeural, LibTorch, ONNX Runtime, and anira on model architectures commonly used in neural audio plugins. The key contribution is moving beyond isolated benchmarks to evaluate performance under realistic DAW contention, constructing mix sessions with configurable plugin loads. Results show that isolated benchmarks can be misleading, and BNNSGraph proves most robust for convolutional models on Apple Silicon.
Benchmarking Integrated GPU Acceleration of Real-Time Neural Audio Inference on Snapdragon
Avery Huang, Gautham Srinivasan, Akito van Troyer and Victor Zappi
This paper investigates whether integrated GPUs on Qualcomm Snapdragon SoCs can accelerate streaming inference of neural audio models. Five models spanning three orders of magnitude in parameter count are benchmarked across three inference approaches (best available CPU, QNN CPU, and QNN GPU). Results reveal when GPU acceleration offers meaningful gains, when per-call overhead negates benefits, and how model size and architecture determine GPU suitability.
Perceptually Motivated Alignment and Interpolation of Pitch-Aligned Time-Frequency Representations
Shahan Nercessian, Jeff Sontag and Alejandro Koretzky
This paper proposes methods for alignment and interpolation of pitch-aligned time-frequency representations, building on the tonal interval vector. Extensions reformulate it as an invertible operator, enabling alignment via permutation search under perceptually weighted distances and interpolation via optimal transport with a circular formulation respecting harmonic structure. The work also develops a geometric scale representation factorizing scale structure into root, density, and color.
Probing Low-Level Acoustic Attribute Encoding in CLAP Audio Embeddings
Hector Martel, Joe Hennessy-Priest and Taemin Cho
This work analyzes CLAP audio embeddings through a probing framework, studying the encoding of reverberation (RT60), loudness (LUFS), spectral content (SC), and relative pitch (RP). Results show that all attributes are reliably recoverable from CLAP embeddings, with RT60, LUFS, and RP approximately linearly encoded, while SC requires non-linear probes. The identified patterns generalize across eight additional audio foundation models.
Tull
16:30 Awards & Closing Ceremony Tull
17:00 Handover Address Tull
17:15 Board Meeting W18 1311
Time Session Location
11:00–13:00 Duck Boat Tour Museum of Science