Słownik AI
Kompletny słownik sztucznej inteligencji
Temporal Diffusion Model
Neural network architecture that applies the diffusion process along the temporal axis to generate coherent video sequences frame by frame.
Spatial Diffusion
Process of noise addition and denoising applied to the spatial dimensions (height and width) of each individual video frame.
3D Video Tensor
Multidimensional data structure representing a video with three dimensions: time, height, and width, used as input for diffusion models.
Latent Video Diffusion
Approach that performs the diffusion process in a compressed latent space rather than directly in pixel space to reduce computational costs.
Video Tokenization
Process of converting raw video data into discrete representations (tokens) in a latent space for more efficient processing.
Motion Dynamics Modeling
Learning motion patterns and temporal transformations to generate realistic and physically plausible animations.
Video Denoising Diffusion
Iterative denoising process applied sequentially to video frames to reconstruct a clear video from initial noise.
Cross-Frame Attention
Mechanism allowing each frame to attend to information from other frames to maintain temporal consistency and spatial relationships.
Video Generation Pipeline
Sequential chain of computational steps including encoding, latent diffusion, denoising, and decoding to produce videos.
Video Latent Space
Compressed vector space where videos are represented as latent codes, facilitating efficient manipulation and generation.
Video Diffusion Sampling
Iterative sampling process in time and space to generate video frames from learned probability distributions.
Conditional Video Synthesis
Generation of videos controlled by multiple conditional inputs such as poses, masks, or motion trajectories.
Video Frame Prediction
Task of predicting future frames of a video from past frames, using diffusion models for generation.