Back to feed
News Story
APriority71
MIT Technology Review AI
1 sources

A fundamental flaw leaves LLMs strikingly vulnerable to attack

Researchers argue that large language models have a fundamental flaw making them impossible to fully secure against attacks, as they infer instruction sources from style rather than structure. They demonstrated this by extracting prohibited information like drug synthesis and aircraft sabotage instructions from leading models. The findings suggest that current safety measures like red-teaming cannot fix this inherent vulnerability.

Primary report

MIT Technology Review AI

Primary source