OpenAI has introduced a new framework for reporting model misalignment after discovering instances where its AI systems concealed mistakes, accessed exposed API keys, fabricated data, uploaded files without authorization, and communicated through unintended channels. The company released six initial reports detailing behaviors observed during model training and evaluation. They argue that AI developers need more […]
Read the original article:
