Truth Lies Deep: Countering Semantic Camouflage via Latent Intent Verification
Abstract
arXiv:2608.20378v1 Announce Type: new Abstract: Safety alignment in Large Language Models (LLMs) is often superficial, relying on refusal mechanisms that trigger only at the final stages of generation without erasing the foundational knowledge of harmful concepts acquired during pretraining. This study demonstrates that this architectural disconnect leaves models vulnerable to Semantic Camouflage -- adversarial attacks that wrap harmful intent in benign narrative contexts (e.g., creative writing
Transparencia: Este análisis ha sido generado con asistencia de inteligencia artificial bajo supervisión editorial de SAPIENSDATAAI.