VIP 👤
🏠 Trang chủ
Benchmark
📊 Tất cả benchmark 🦖 Khủng long v1 🦖 Khủng long v2 ✅ Ứng dụng To-Do List 🎨 Trang tự do sáng tạo 🎯 FSACB - Trình diễn cuối cùng 🌍 Benchmark dịch thuật
Mô hình
🏆 Top 10 mô hình 🆓 Mô hình miễn phí 📋 Tất cả mô hình ⚙️ Kilo Code
Tài nguyên
💬 Thư viện prompt 📖 Thuật ngữ AI 🔗 Liên kết hữu ích 🔌 API và bộ định tuyến AI

Thuật ngữ AI

Từ điển đầy đủ về Trí tuệ nhân tạo

162
danh mục
2.032
danh mục con
23.060
thuật ngữ
📖
thuật ngữ

Multi-modal Model

AI architecture capable of processing, understanding, and simultaneously generating multiple types of unstructured data such as text, images, audio, or video in a common representation space.

📖
thuật ngữ

DALL-E

Text-to-image generation system developed by OpenAI, combining a VQ-VAE transformer with a diffusion model to create photorealistic images from text descriptions.

📖
thuật ngữ

Vision-Language Transformer (ViLT)

Transformer architecture that jointly processes image patches and text tokens without feature pre-extraction, enabling end-to-end learning for vision-language tasks.

📖
thuật ngữ

Multi-modal Foundation Model

Large-scale pre-trained model on diverse multi-modal data, capable of being fine-tuned for numerous downstream tasks without requiring training from scratch.

📖
thuật ngữ

Multi-modal Perceiver

Unified architecture that processes different modalities (text, image, audio) as token sequences in a single latent space, using cross-attentions to model their interactions.

📖
thuật ngữ

GPT-4V(ision)

Multi-modal version of GPT-4 capable of accepting images and text as input, using a hybrid architecture to reason about visual content and generate contextualized text responses.

📖
thuật ngữ

Audio-Visual Speech Recognition (AVSR)

Multi-modal system that combines audio signals and video (lip movements) to improve speech recognition robustness, particularly in noisy environments.

📖
thuật ngữ

Visual Language Model (VLM)

Class of multi-modal models that extend language transformer architectures to understand and reason about visual content, often through cross-attention mechanisms.

📖
thuật ngữ

Contrastive Pre-training Multi-modal

Self-supervised training paradigm that learns aligned representations between modalities by maximizing the similarity of positive pairs and minimizing that of negative pairs in the latent space.

🔍

Không tìm thấy kết quả