Skip to main content

GAN Architectures for Audio

GANs have been widely used as neural vocoders: models that convert acoustic features such as mel spectrograms into waveforms. They are one family among several waveform-generation approaches; a modern audio pipeline is not necessarily GAN-based.

GAN fundamentals​

A GAN trains a generator GG against a discriminator DD:

min⁑Gmax⁑Dβ€…β€ŠEx∼pdata[log⁑D(x)]+Ez∼pz[log⁑(1βˆ’D(G(z)))]\min_G \max_D \; \mathbb{E}_{x \sim p_{\text{data}}}[\log D(x)] + \mathbb{E}_{z \sim p_z}[\log(1 - D(G(z)))]

Audio GANs frequently modify this original objective and add reconstruction or feature losses. Do not assume that every audio GAN uses the same adversarial loss.

HiFi-GAN​

HiFi-GAN was published at NeurIPS 2020 as an efficient, high-fidelity speech vocoder. Its generator upsamples mel-spectrogram features with transposed convolutions and Multi-Receptive Field Fusion blocks.

A representative V1 upsampling schedule is:

Mel spectrogram
-> ConvTranspose1d x8
-> MRF
-> ConvTranspose1d x8
-> MRF
-> ConvTranspose1d x2
-> MRF
-> ConvTranspose1d x2
-> MRF
-> waveform

The paper's Multi-Period Discriminator uses periods {2,3,5,7,11}\{2,3,5,7,11\} and is paired with a Multi-Scale Discriminator. The published generator objective combines adversarial, feature-matching, and mel-spectrogram losses:

LG=Ladv+Ξ»fmLfm+Ξ»melLmel\mathcal{L}_G = \mathcal{L}_{\text{adv}} + \lambda_{\text{fm}}\mathcal{L}_{\text{fm}} + \lambda_{\text{mel}}\mathcal{L}_{\text{mel}}

with Ξ»fm=2\lambda_{\text{fm}}=2 and Ξ»mel=45\lambda_{\text{mel}}=45 in the paper's experiments.

The NeurIPS paper reports 22.05 kHz synthesis 167.9Γ— faster than real time on a single V100 GPU for its reported configuration. That figure is a benchmark result, not a hardware-independent property of HiFi-GAN.

BigVGAN​

BigVGAN was released as a 2022 preprint and published at ICLR 2023. It introduces periodic activations and anti-aliased representations, and scales the vocoder to substantially larger parameter counts than earlier systems in its comparison.

The original paper reports strong zero-shot / out-of-distribution results for unseen speakers, languages, recording environments, singing, music, and instrumental audio. Those claims should be read as results under the paper's evaluation setup rather than a permanent statement that BigVGAN is the best vocoder on every benchmark.

BigVGAN-v2, released in 2024, changed the training recipe and discriminator stack and added optimized CUDA inference support. Keep version names explicit when comparing results.

MelGAN and Multi-Band MelGAN​

MelGAN (2019) demonstrated efficient adversarial waveform generation without an autoregressive decoder. Multi-Band MelGAN reduces computation by predicting sub-band signals that are combined through a synthesis filter bank.

These models remain useful historical and implementation references, but performance comparisons should cite a specific dataset, checkpoint, sample rate, hardware configuration, and listening protocol.

UnivNet​

UnivNet uses location-variable convolutions in the generator and a multi-resolution spectrogram discriminator. It is another example of a GAN vocoder whose quality and speed depend on the checkpoint and evaluation setup rather than a universal ranking.

Vocos​

Vocos (2023) is not simply another time-domain GAN vocoder. It predicts Fourier spectral coefficients and reconstructs audio with an inverse STFT. Its paper reports competitive audio quality and substantial computational gains over the time-domain systems evaluated there.

Representing the output as complex STFT coefficients can be written as:

x^=iSTFT⁑(X^)\hat{x}=\operatorname{iSTFT}(\hat{X})

where X^\hat{X} contains the predicted spectral coefficients.

Comparing vocoders correctly​

Avoid tables such as β€œHiFi-GAN = 80Γ— real time, BigVGAN = 40Γ—, Vocos = 200×” unless every number comes from the same benchmark. Speed varies with:

  • GPU/CPU model and numerical precision;
  • batch size and sequence length;
  • sample rate and hop size;
  • model/checkpoint size;
  • CUDA kernels and implementation version;
  • whether preprocessing and I/O are included.

Quality also depends on the conditioning distribution. A vocoder tested on ground-truth mel spectrograms may behave differently when driven by predicted or out-of-domain features.

Practical evaluation checklist​

For a defensible comparison, record:

  1. exact repository commit and checkpoint;
  2. dataset and licence;
  3. sample rate, channels, mel/STFT configuration;
  4. hardware, precision, batch size, and warm-up procedure;
  5. real-time factor definition;
  6. objective metrics and listening-test protocol;
  7. whether conditioning features are ground truth or generated by another model.

Primary sources​

Checked 5 September 2026:

Use the paper and checkpoint associated with the result you cite; do not turn a historical benchmark win into a timeless β€œstate-of-the-art” label.