A
AI & Machine Learning
Artificial intelligence and machine learning content
Safety Alignment Should be Made More Than Just a Few Tokens Deep (Paper Explained)
This paper demonstrates in a series of experiments that current safety alignment techniques of LLMs, as well as corresponding jailbreaking attacks, are in large part focusing on modulating the distribution of the first few tokens of the LLM response. Paper: https://openreview.net/forum?id=6Mxhg9PtD...



