Loss Functions for Audio Generation
A loss specifies the numerical objective optimized during training. Its effect on perceived quality, stability, or controllability depends on the architecture, data, weighting, optimization procedure, and evaluation protocol; a loss name alone does not guarantee a particular audible result.
Pointwise reconstruction lossesβ
For predicted representation and target :
L2 penalizes large numerical errors quadratically; L1 grows linearly. Statements such as βL1 preserves transients betterβ are experiment-dependent and should be supported by a controlled comparison on the representation and task being discussed.
For audio, pointwise losses can be applied to waveform samples, magnitudes, log magnitudes, mel features, latents, or other representations. These objectives are not equivalent.
Adversarial objectivesβ
The original GAN minimax value function is
Practical audio GANs often use different generator/discriminator losses, including non-saturating, least-squares, hinge, or Wasserstein-style objectives. Therefore the original minimax equation should not be presented as the universal βaudio adversarial loss.β
HiFi-GAN, for example, combines adversarial objectives with feature-matching and mel-spectrogram reconstruction terms. Its multi-period discriminators are designed to inspect periodic structure at several periods, while multi-scale discriminators inspect waveforms at multiple resolutions. Claims about rhythm, timbre, or long-range musical form should be tied to an experiment rather than inferred from the discriminator name.
Feature matching and perceptual lossesβ
A generic feature-space loss can be written as
Here may be an internal discriminator activation or a separate pretrained representation. Feature matching has improved synthesis results in published systems such as HiFi-GAN, but it does not universally βreduce metallic artifactsβ or outperform every reconstruction objective. Its behavior depends on the features, layers, normalization, weights, and data domain.
When a pretrained network supplies , document its checkpoint and training domain; the resulting loss inherits that model's biases and invariances.
Diffusion objectivesβ
For a variance-preserving forward process, one common parameterization is
with a noise-prediction objective
Noise prediction is only one parameterization. Diffusion and flow systems may instead predict , a velocity variable , a score, or a flow/vector field, with weighting choices that alter the effective objective. Do not infer a system's training target from the word βdiffusionβ alone.
KL terms in variational modelsβ
For a diagonal Gaussian approximate posterior and standard-normal prior,
In a VAE, the KL term is part of the evidence lower bound and regularizes the approximate posterior toward the prior. It does not prevent posterior collapse. With an expressive decoder or an overly strong KL pressure, the model can instead learn a posterior close to the prior and use little information from βthe phenomenon commonly called posterior collapse or KL vanishing.
Mitigations studied in the literature include KL annealing, free bits, modified objectives, architectural changes, and constraints on the inference network. Their effectiveness is task-dependent.
Multi-objective trainingβ
Audio systems frequently combine several losses:
The coefficients are part of the model specification. A raw loss magnitude is not necessarily comparable across objectives, datasets, or implementations, so report weights, reduction conventions, units, and any adaptive loss-balancing scheme.
When claiming that a loss improves quality, report an ablation with the same data, model capacity, training budget, inference settings, and evaluation protocol.
Primary referencesβ
- Goodfellow et al., Generative Adversarial Nets (2014).
- Kingma & Welling, Auto-Encoding Variational Bayes (2013).
- Ho, Jain & Abbeel, Denoising Diffusion Probabilistic Models (2020).
- Kong, Kim & Bae, HiFi-GAN (2020).
Sources checked: 2026-09-05.