🏠 Trang chủ
Benchmark
📊 Tất cả benchmark 🦖 Khủng long v1 🦖 Khủng long v2 ✅ Ứng dụng To-Do List 🎨 Trang tự do sáng tạo 🎯 FSACB - Trình diễn cuối cùng 🌍 Benchmark dịch thuật
Mô hình
🏆 Top 10 mô hình 🆓 Mô hình miễn phí 📋 Tất cả mô hình ⚙️ Kilo Code
Tài nguyên
💬 Thư viện prompt 📖 Thuật ngữ AI 🔗 Liên kết hữu ích

Thuật ngữ AI

Từ điển đầy đủ về Trí tuệ nhân tạo

162
danh mục
2.032
danh mục con
23.060
thuật ngữ
📖
thuật ngữ

Policy

Strategy or mapping that defines the action to take in each possible state, representing the agent's behavior in a reinforcement learning process.

📖
thuật ngữ

Multi-Armed Bandit Problem

Sequential optimization problem where an agent must choose among several options with unknown rewards to maximize cumulative reward over time.

📖
thuật ngữ

Cumulative Reward

Sum of expected future rewards that the agent seeks to maximize, often calculated with a discount factor to give less weight to distant rewards.

📖
thuật ngữ

SARSA Algorithm

On-policy reinforcement learning algorithm that updates Q-values based on the State-Action-Reward-State-Action sequence, unlike Q-learning.

📖
thuật ngữ

Deep Q-Network

Deep neural network architecture used to approximate the Q-function in complex state spaces, combining deep learning and Q-learning.

📖
thuật ngữ

Deep Reinforcement Learning

Approach integrating deep neural networks into reinforcement learning to handle high-dimensional state or action spaces.

📖
thuật ngữ

Epsilon-Greedy Policy

Action selection strategy where with probability ε the agent explores (chooses a random action) and with probability 1-ε it exploits (chooses the best known action).

📖
thuật ngữ

Policy Optimization

Class of methods in reinforcement learning that directly optimize the policy without going through a value function, often using policy gradient techniques.

📖
thuật ngữ

Policy Gradient Algorithm

Optimization method that directly adjusts policy parameters by following the gradient of the expected reward with respect to these parameters.

📖
thuật ngữ

Multi-Agent Reinforcement Learning

Extension of reinforcement learning where multiple agents learn simultaneously, often in competition or cooperation, in a shared environment.

📖
thuật ngữ

Experience Replay Memory

Data structure storing transitions (state, action, reward, next state) for resampling during training, improving data usage efficiency.

📖
thuật ngữ

Actor-Critic Algorithm

Architecture combining an actor that selects actions according to a policy and a critic that evaluates these actions, enabling more stable and efficient learning.

🔍

Không tìm thấy kết quả