1Sony Computer Science Laboratories Paris 2Queen Mary University of London
Sound matching can be formulated as optimizing synthesizer parameters against an audio-domain objective. However, objectives derived from generic audio representations are often difficult to optimize, while direct search requires rendering every candidate. We introduce Synth-JEPA, which learns mutually predictive audio and parameter representations from paired synthesizer data. At inference, candidate parameters are scored directly in this learned space, yielding a renderer-free objective whose audio geometry is shaped by parameter correspondences rather than generic audio similarity. We evaluate Synth-JEPA on Surge XT using held-out synthesizer sounds and out-of-domain NSynth and FSD50K targets, against inverse models, direct search, and learned proxy objectives. Synth-JEPA outperforms all baselines in-domain and remains competitive out-of-domain. Its matching quality continues to improve with additional test-time search, allowing compute to be traded for match quality. In pairwise listening tests, listeners preferred Synth-JEPA in 85% of trials overall. Together, these results show that an audio representation with a parameter-induced geometry allows synthesizer sound matching to be approached as an effective renderer-free search problem.
In these examples, the target audio was rendered by Surge XT itself.
Targets are rendered from parameters drawn at random from the parameter prior, so a setting that reproduces each target exactly always exists. Each column is a method from Table 1 of the paper. Every reconstruction is the method's returned preset rendered by Surge XT.
Beneath five of them is the best-of-64 by log-Mel distance, from Fig. 3 in the paper.
In these examples, the target audio did not come from the synthesizer.
This means that, in general, the synthesizer can not exactly recreate these sounds. No model saw anything but Surge XT audio in training.
NSynth contains notes from acoustic, electronic and synthetic instruments; FSD50K contains musical and non-musical sounds from Freesound, trimmed to their first 3 s.
Table 1 of the paper: means over 1 024 targets per set. MSS is multi-scale spectral distance, wMFCC is warped MFCC distance, and CLAP is the cosine similarity of CLAP embeddings. Bold marks the best result in each column and underlining the second best.
| Family | Method | In-domain presets | NSynth | Freesound (FSD50K) | Renders / target | Parameters | ||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|
| MSS ↓ | wMFCC ↓ | CLAP ↑ | MSS ↓ | wMFCC ↓ | CLAP ↑ | MSS ↓ | wMFCC ↓ | CLAP ↑ | ||||
| Synth-JEPA | SIGReg | 8.44 | 11.07 | 0.793 | 10.25 | 15.54 | 0.465 | 13.59 | 14.81 | 0.250 | 0 | 53.3M |
| EMA teacher | 11.52 | 15.29 | 0.625 | 14.02 | 19.33 | 0.323 | 15.39 | 17.33 | 0.179 | 0 | 53.3M | |
| Amortized | AST regression | 17.94 | 23.16 | 0.499 | 28.06 | 30.64 | 0.104 | 21.48 | 26.56 | 0.086 | 0 | 25.0M |
| Flow matching | 12.75 | 14.78 | 0.707 | 16.67 | 21.24 | 0.356 | 17.94 | 18.84 | 0.204 | 0 | 44.0M | |
| Renderer search | Log-Mel distance | 15.86 | 12.17 | 0.553 | 17.81 | 14.10 | 0.199 | 14.55 | 11.61 | 0.178 | 2048 | — |
| Embedding distance (mn20) | 12.90 | 16.43 | 0.730 | 14.61 | 25.05 | 0.387 | 16.12 | 21.41 | 0.249 | 2048 | 17.9M | |
| Learned ℰₐ distance | 9.48 | 13.15 | 0.729 | 12.31 | 17.30 | 0.360 | 14.49 | 16.88 | 0.208 | 2048 | 17.9M | |
| Proxy search | Embedding (mn20) proxy | 12.56 | 15.77 | 0.768 | 13.65 | 23.98 | 0.436 | 15.44 | 19.79 | 0.261 | 0 | 31.8M |
| Audio (log-Mel) proxy | 10.75 | 12.74 | 0.699 | 11.03 | 14.10 | 0.350 | 12.18 | 13.87 | 0.216 | 0 | 31.8M | |