New AI Attack Hides Malicious Instructions in Normal-Looking Text to Evade Safety Filters

A newly disclosed prompt-crafting technique can hide policy-violating instructions inside ordinary-looking English prose, allowing malicious requests to pass through lightweight LLM safety filters before being recovered and processed by a more capable downstream model. Researchers found that carefully structured prose can make the first model miss an embedded instruction entirely, while the target model invests […]

This article has been indexed from GBHackers Security | #1 Globally Trusted Cyber Security News Platform

Read the original article: