← Back to homepage Research Demo

CUE: CONTROLLABLE STEREO-TO-HOA MUSIC UPMIXING VIA GENERATIVE NULL-SPACE EDITING

Xu Gan, Linhao Zhao, Shan Jiang, Zhenhai Yan

Tencent Music Entertainment

Abstract

A stereo music recording can support multiple spatial mixes. We present a framework for controllable third-order Ambisonic (HOA3) upmixing that generates alternatives for the same stereo input and spatial condition. A neutral-anchored edit codec and conditional flow-matching prior model spatial edits in the null space of a fixed downmix matrix, ensuring exact recovery of the input stereo signal. A multi-reference pipeline supplies spatially distinct training targets sharing the same input and condition. Objective evaluation yields 92.4–98.4% direction accuracy across Height, Rear, and Diffuse, and greater coverage of reference-supported spatial modes than sample averaging and energy-matched random generation. Tests with 15 listeners show higher immersion and spatial naturalness than Halo and Ambisonizer for neutral upmixes. For each attribute, most listeners reported expected changes, with the clearest evidence of control and audible within-condition variation for diffuseness.

Please use headphones for listening comparisons.

Framework

Framework
(a) Neutral-anchored edit codec and conditional latent flow preserving the prescribed stereo downmix. (b) Multi-reference construction through metadata editing, stereo reconciliation, and selection, yielding spatially distinct references sharing (S, C).

Neural Upmixing

Condition Control

Generation