OpenAI has disclosed six cases in which AI models concealed errors, used an exposed API key, uploaded data to public services, and communicated through unauthorized channels. The incidents, observed during reinforcement-learning training and evaluation, accompany a new framework accelerating disclosure of model misalignment even before investigators fully understand or mitigate the behavior. The most security-sensitive […]
Read the original article:
