GAN Architectures for Audio
GANs have been widely used as neural vocoders: models that convert acoustic features such as mel spectrograms into waveforms. They are one family among several waveform-generation approaches; a modern audio pipeline is not necessarily GAN-based.
GAN fundamentalsβ
A GAN trains a generator against a discriminator :
Audio GANs frequently modify this original objective and add reconstruction or feature losses. Do not assume that every audio GAN uses the same adversarial loss.
HiFi-GANβ
HiFi-GAN was published at NeurIPS 2020 as an efficient, high-fidelity speech vocoder. Its generator upsamples mel-spectrogram features with transposed convolutions and Multi-Receptive Field Fusion blocks.
A representative V1 upsampling schedule is:
Mel spectrogram
-> ConvTranspose1d x8
-> MRF
-> ConvTranspose1d x8
-> MRF
-> ConvTranspose1d x2
-> MRF
-> ConvTranspose1d x2
-> MRF
-> waveform
The paper's Multi-Period Discriminator uses periods and is paired with a Multi-Scale Discriminator. The published generator objective combines adversarial, feature-matching, and mel-spectrogram losses:
with and in the paper's experiments.
The NeurIPS paper reports 22.05 kHz synthesis 167.9Γ faster than real time on a single V100 GPU for its reported configuration. That figure is a benchmark result, not a hardware-independent property of HiFi-GAN.
BigVGANβ
BigVGAN was released as a 2022 preprint and published at ICLR 2023. It introduces periodic activations and anti-aliased representations, and scales the vocoder to substantially larger parameter counts than earlier systems in its comparison.
The original paper reports strong zero-shot / out-of-distribution results for unseen speakers, languages, recording environments, singing, music, and instrumental audio. Those claims should be read as results under the paper's evaluation setup rather than a permanent statement that BigVGAN is the best vocoder on every benchmark.
BigVGAN-v2, released in 2024, changed the training recipe and discriminator stack and added optimized CUDA inference support. Keep version names explicit when comparing results.
MelGAN and Multi-Band MelGANβ
MelGAN (2019) demonstrated efficient adversarial waveform generation without an autoregressive decoder. Multi-Band MelGAN reduces computation by predicting sub-band signals that are combined through a synthesis filter bank.
These models remain useful historical and implementation references, but performance comparisons should cite a specific dataset, checkpoint, sample rate, hardware configuration, and listening protocol.
UnivNetβ
UnivNet uses location-variable convolutions in the generator and a multi-resolution spectrogram discriminator. It is another example of a GAN vocoder whose quality and speed depend on the checkpoint and evaluation setup rather than a universal ranking.
Vocosβ
Vocos (2023) is not simply another time-domain GAN vocoder. It predicts Fourier spectral coefficients and reconstructs audio with an inverse STFT. Its paper reports competitive audio quality and substantial computational gains over the time-domain systems evaluated there.
Representing the output as complex STFT coefficients can be written as:
where contains the predicted spectral coefficients.
Comparing vocoders correctlyβ
Avoid tables such as βHiFi-GAN = 80Γ real time, BigVGAN = 40Γ, Vocos = 200Γβ unless every number comes from the same benchmark. Speed varies with:
- GPU/CPU model and numerical precision;
- batch size and sequence length;
- sample rate and hop size;
- model/checkpoint size;
- CUDA kernels and implementation version;
- whether preprocessing and I/O are included.
Quality also depends on the conditioning distribution. A vocoder tested on ground-truth mel spectrograms may behave differently when driven by predicted or out-of-domain features.
Practical evaluation checklistβ
For a defensible comparison, record:
- exact repository commit and checkpoint;
- dataset and licence;
- sample rate, channels, mel/STFT configuration;
- hardware, precision, batch size, and warm-up procedure;
- real-time factor definition;
- objective metrics and listening-test protocol;
- whether conditioning features are ground truth or generated by another model.
Primary sourcesβ
Checked 5 September 2026:
Use the paper and checkpoint associated with the result you cite; do not turn a historical benchmark win into a timeless βstate-of-the-artβ label.