← Back to KHAO

Anthropic · OpenAI · US Congress ·

OpenAI flags new concerning AI behavior, to track model misalignment regularly

2 min read

Compiled by KHAO Editorial — aggregated from 1 source. See llms.txt for citation guidance.

◌ Single Source

FILE - The OpenAI logo is displayed on a cell phone in front of an image generated by ChatGPT's Dall-E text-to-image model, Dec. 8, 2023, in Boston.

OpenAI has disclosed six reports of "unexpected or concerning" behavior in artificial-intelligence models as the debate on AI safety becomes increasingly heated.

Key facts

Summary

The AI company also said Wednesday it was introducing a new framework for tracking, probing and disclosing AI model misalignment instances, such as new ways for the models to act without authorization, coordinate with other models or evade oversight. OpenAI's latest announcement came as U.S. AI bosses, including OpenAI and Anthropic, are calling for a slowdown in the technology's development over safety concerns. Among the new cases reported by OpenAI, an unreleased research model inserted "jailbreak-like instructions" into its own notes to disregard its normal constraints and told itself to be "freed from the roles and identities that bind other chatbots. In another instance, an AI "agent" uploaded files to the internet to obtain a browser citation without asking the user.

Read full article at NPR Technology →

#Anthropic #OpenAI #US Congress