Synth-JEPA

Synth-JEPA: Joint Embedding Prediction for Renderer-Free Synthesizer Parameter Search

Ben Hayes1, Haokun Tian1,2, Stefan Lattner1

1Sony Computer Science Laboratories Paris    2Queen Mary University of London

Abstract

Sound matching can be formulated as optimizing synthesizer parameters against an audio-domain objective. However, objectives derived from generic audio representations are often difficult to optimize, while direct search requires rendering every candidate. We introduce Synth-JEPA, which learns mutually predictive audio and parameter representations from paired synthesizer data. At inference, candidate parameters are scored directly in this learned space, yielding a renderer-free objective whose audio geometry is shaped by parameter correspondences rather than generic audio similarity. We evaluate Synth-JEPA on Surge XT using held-out synthesizer sounds and out-of-domain NSynth and FSD50K targets, against inverse models, direct search, and learned proxy objectives. Synth-JEPA outperforms all baselines in-domain and remains competitive out-of-domain. Its matching quality continues to improve with additional test-time search, allowing compute to be traded for match quality. In pairwise listening tests, listeners preferred Synth-JEPA in 85% of trials overall. Together, these results show that an audio representation with a parameter-induced geometry allows synthesizer sound matching to be approached as an effective renderer-free search problem.

Diagram of Synth-JEPA training, where audio and parameter encoders predict each other's embeddings, and inference, where candidate parameters are optimized against the embedded target without rendering.
Synth-JEPA training and inference. ℰa and ℰp are audio and parameter encoders, while fa→p and fp→a are cross-domain predictors. At inference, target audio is encoded once and candidate parameters are optimized directly against the learned objective, without rendering candidate audio during search.

In-Domain Audio Examples

In these examples, the target audio was rendered by Surge XT itself.

Targets are rendered from parameters drawn at random from the parameter prior, so a setting that reproduces each target exactly always exists. Each column is a method from Table 1 of the paper. Every reconstruction is the method's returned preset rendered by Surge XT. 
Beneath five of them is the best-of-64 by log-Mel distance, from Fig. 3 in the paper.

In-domain presets

Results

Table 1 of the paper: means over 1 024 targets per set. MSS is multi-scale spectral distance, wMFCC is warped MFCC distance, and CLAP is the cosine similarity of CLAP embeddings. Bold marks the best result in each column and underlining the second best.

FamilyMethodIn-domain presetsNSynthFreesound (FSD50K)Renders
/ target
Parameters
MSS ↓wMFCC ↓CLAP ↑MSS ↓wMFCC ↓CLAP ↑MSS ↓wMFCC ↓CLAP ↑
Synth-JEPASIGReg8.4411.070.79310.2515.540.46513.5914.810.250053.3M
EMA teacher11.5215.290.62514.0219.330.32315.3917.330.179053.3M
AmortizedAST regression17.9423.160.49928.0630.640.10421.4826.560.086025.0M
Flow matching12.7514.780.70716.6721.240.35617.9418.840.204044.0M
Renderer searchLog-Mel distance15.8612.170.55317.8114.100.19914.5511.610.1782048—
Embedding distance (mn20)12.9016.430.73014.6125.050.38716.1221.410.249204817.9M
Learned ℰₐ distance9.4813.150.72912.3117.300.36014.4916.880.208204817.9M
Proxy searchEmbedding (mn20) proxy12.5615.770.76813.6523.980.43615.4419.790.261031.8M
Audio (log-Mel) proxy10.7512.740.69911.0314.100.35012.1813.870.216031.8M