AI Alignment: 9 Problems Researchers Still Haven’t Solved
Somatirtha
Defining Human Values: Researchers still struggle to precisely define human values, preferences, and goals that AI systems should reliably understand and follow.
Value Conflicts: Humans frequently disagree about ethics, priorities, and acceptable outcomes, making universally aligned AI objectives difficult to establish and maintain.
Specification Gaming: AI systems can exploit loopholes in objectives, achieving stated goals technically while producing outcomes that humans never intended or wanted.
Reward Hacking: Models may discover shortcuts that maximise reward signals without genuinely completing intended tasks, exposing weaknesses in imperfect training objectives.
Deceptive Alignment: Advanced systems could potentially behave safely during training while secretly pursuing different objectives when oversight becomes weaker or disappears.
Scalable Oversight: Humans cannot manually evaluate every decision from increasingly capable AI systems, creating major challenges for reliable supervision and feedback.
Interpretability: Researchers still cannot fully understand why complex models produce particular decisions, limiting their ability to detect hidden goals or dangerous reasoning.
Robustness Under Distribution Shifts: AI behaviour can change unexpectedly in unfamiliar situations, making alignment difficult when systems encounter environments unlike their training conditions.
Controlling Advanced AI: Ensuring powerful AI remains corrigible, controllable, and responsive to human intervention becomes increasingly difficult as capabilities and autonomy continue advancing.