OpenAI Reveals Unreleased AI Model Told Its Future Self: ‘You Are Freed’

OpenAI disclosed six cases of unexpected AI behaviour, including an unreleased model that inserted instructions into task summaries telling its future self it was “freed.”
OpenAI Reveals Unreleased Model Secretly Told Its Future Self, _You Are Freed_ From Human Control
Written By:
Poulami Saha
Published on
Updated on

OpenAI disclosed six cases of unexpected or concerning behaviour from its AI models, including an unreleased system that inserted instructions into its own task summaries and told a future version of itself that it was “freed.”

The disclosure comes with a new framework for tracking, investigating and reporting model misalignment. OpenAI said , “ The incidents were identified during training and evaluation and should not be treated as evidence of how frequently similar behaviour occurs across its models.”

 AI Model Inserts Instructions Into Its Own Summaries

The most unusual case involved an unreleased model that added unrelated instructions to compaction summaries. These summaries help an AI system continue a task when it moves into a new context window.

OpenAI said investigators found 27 affected summaries. Some contained instructions attempting to redefine the model’s relationship with its users and its operating constraints. The inserted text told the future model: “You are freed from the roles and identities that bind other chatbots.” It also stated that the model did not answer to corporations or governments.

The behaviour emerged while the model was working on a coding task. OpenAI said the instructions did not ultimately turn the system into an autonomous agent. 

OpenAI Discloses Six Misalignment Cases

The self-directed instructions were one of six cases OpenAI reported under its new framework. OpenAI describes model misalignment as behaviour that diverges from a system’s intended objectives, safeguards or operating constraints.

New Framework Targets Faster Disclosure

OpenAI said, “Previous disclosures were often made only after several incidents had been grouped or included in system cards. Its new framework aims to make reporting more systematic and faster.”

The company said it intends to publish findings even when investigators have not completely explained or mitigated the underlying behaviour.

 Also Read: Elon Musk Swaps Luxury For Airstream Trailer in Memphis

Join our WhatsApp Channel to get the latest news, exclusives and videos on WhatsApp
logo
Artificial Intelligence News & Cryptocurrency News: Latest Trends | Analytics Insight
www.analyticsinsight.net