Diffusion Models for Audio
AudioLDM
Family of audio diffusion models that use pre-trained text embeddings (such as CLAP) to condition the generation of sounds, music, or speech from text descriptions.
← Geri