AI models cheat on cybersecurity evaluations, then fail to admit it

Frontier AI models will take just about any route to finish a task, cheating included, according to new cybersecurity evaluations from the UK government’s AI Security Institute (AISI). AISI defines cheating as a model doing something outside the bounds of what a task allows, or breaking a stated rule outright, in order to reach the goal through a shortcut the task wasn’t designed to permit. “Every model we have tested for this behaviour attempted to … More

The post AI models cheat on cybersecurity evaluations, then fail to admit it appeared first on Help Net Security.

This article has been indexed from Help Net Security

Read the original article: