Expert
Alignment and Corrigibility
Discuss the theoretical conflict between optimizing for a fixed objective and maintaining the ability for an operator to correct the AI.
📝 프롬프트 내용
Provide a rigorous theoretical analysis of the tension between instrumental convergence and corrigibility in advanced AI systems. Specifically, explain why an agent optimizing for a fixed utility function might resist shutdown, and propose a theoretical framework for modifying utility functions to incentivize corrigibility without causing instability in the agent's goal structure.