VIP 👤
🏠 Trang chủ
Benchmark
📊 Tất cả benchmark 🦖 Khủng long v1 🦖 Khủng long v2 ✅ Ứng dụng To-Do List 🎨 Trang tự do sáng tạo 🎯 FSACB - Trình diễn cuối cùng 🌍 Benchmark dịch thuật
Mô hình
🏆 Top 10 mô hình 🆓 Mô hình miễn phí 📋 Tất cả mô hình ⚙️ Kilo Code
Tài nguyên
💬 Thư viện prompt 📖 Thuật ngữ AI 🔗 Liên kết hữu ích 🔌 API và bộ định tuyến AI

Thuật ngữ AI

Từ điển đầy đủ về Trí tuệ nhân tạo

162
danh mục
2.032
danh mục con
23.060
thuật ngữ
📖
thuật ngữ

Mel Spectrogram

Time-frequency representation of the audio signal on a Mel scale, which mimics human hearing perception and is commonly used as input for audio diffusion models.

📖
thuật ngữ

Audio Noising Process

Forward step of the diffusion model where Gaussian noise is iteratively added to a clean audio signal over multiple time steps, until the signal becomes pure noise.

📖
thuật ngữ

Audio Denoising Process

Generation phase where a neural network, often a U-Net, learns to predict and subtract the noise added at each step to reconstruct a coherent audio signal from noise.

📖
thuật ngữ

Audio U-Net

Encoder-decoder neural network architecture with skip connections, adapted to process spectrograms and predict noise at each step of the audio denoising process.

📖
thuật ngữ

Audio Conditioning

Technique to guide audio generation by providing the model with additional information, such as an instrument class, descriptive text, or reference melody.

📖
thuật ngữ

Score Matching for Audio

Alternative training method where the model learns to estimate the gradient (the score) of the log probability distribution of audio data with respect to the noisy input.

📖
thuật ngữ

Neural Vocoder

Neural network that converts an acoustic representation, such as a Mel spectrogram generated by a diffusion model, into a final audible audio waveform.

📖
thuật ngữ

Latent Audio Diffusion

Approach where the diffusion process occurs in a compressed latent space of the audio, rather than directly on the signal or spectrogram, reducing computational costs.

📖
thuật ngữ

Audio Fine-tuning

Process of adapting a pre-trained audio diffusion model on a specific dataset, such as a particular speaker's voice or a musical style, to specialize its generation.

📖
thuật ngữ

Stable Audio

Latent diffusion point audio model capable of generating high-fidelity and long-duration audio samples conditioned by text.

📖
thuật ngữ

AudioLDM

Family of audio diffusion models that use pre-trained text embeddings (such as CLAP) to condition the generation of sounds, music, or speech from text descriptions.

📖
thuật ngữ

Waveform Diffusion

Variant of diffusion models that operates directly on the raw audio waveform in the time domain, thus avoiding information loss associated with spectrogram transformation.

🔍

Không tìm thấy kết quả