VIP 👤
🏠 홈
벤치마크
📊 모든 벤치마크 🦖 공룡 v1 🦖 공룡 v2 ✅ 할 일 목록 앱 🎨 창의적인 자유 페이지 🎯 FSACB - 궁극의 쇼케이스 🌍 번역 벤치마크
모델
🏆 톱 10 모델 🆓 무료 모델 📋 모든 모델 ⚙️ 킬로 코드 모드
리소스
💬 프롬프트 라이브러리 📖 AI 용어 사전 🔗 유용한 링크 🔌 AI API 및 라우터
Expert

Alignment and Corrigibility

#alignment #theory #ethics

Discuss the theoretical conflict between optimizing for a fixed objective and maintaining the ability for an operator to correct the AI.

Provide a rigorous theoretical analysis of the tension between instrumental convergence and corrigibility in advanced AI systems. Specifically, explain why an agent optimizing for a fixed utility function might resist shutdown, and propose a theoretical framework for modifying utility functions to incentivize corrigibility without causing instability in the agent's goal structure.