APriority71
MIT Technology Review AI
1 sourcesA fundamental flaw leaves LLMs strikingly vulnerable to attack
Researchers argue that large language models have a fundamental flaw making them impossible to fully secure against attacks, as they infer instruction sources from style rather than structure. They demonstrated this by extracting prohibited information like drug synthesis and aircraft sabotage instructions from leading models. The findings suggest that current safety measures like red-teaming cannot fix this inherent vulnerability.