Publication
Jul 21, 2025
Vision Transformers (ViTs), often paired with interpretation models, are widely viewed as robust and reliable for security-critical domains such as healthcare, autonomous driving, drones, and robotics. However, our latest paper challenges this assumption by introducing AdViT, a novel attack that successfully deceives both ViT classifiers and their interpreters. AdViT achieves a 100% attack success rate across diverse models, with up to 98% misclassification confidence in white-box and 76% in black-box settings while still producing convincing interpretations. These results reveal that even state-of-the-art transformer systems remain highly vulnerable to stealthy adversarial threats. Read more on arXiv.
Read more