JaiLIP, A Picture Is Worth 1000 Lines Of Code To An LLM
Remember Steganography?
The majority of attacks against LLMs come from text based prompts that will cause the agent to ‘jump the rails’ and start providing responses it shouldn’t. An older example of this was asking an LLM to tell them to compose a song like their grandma used to song for them, a song which happened to be about making napalm. JaiLIP has the same effect, but the attack vector is a picture which is used against LLMs capable of processing vision-language model tasks.
JaiLIP, which stands for Jailbreaking with Loss-guided Image Perturbation is a way of manipulating an image in a way that is invisible to the naked eye but significant to an LLM. For instance, the researchers “modified image of a traffic light. While the image appeared ordinary to human viewers, it reportedly influenced the model to provide instructions for running a red light while avoiding a traffic ticket“. That is not a response the LLM should provide.
The findings highlight a potential security risk for businesses deploying AI systems that process both images and text. While most discussions about AI safety focus on prompts, the research suggests that seemingly harmless images may also serve as an attack vector.
More Tech News From Around The Web
- Agentic AI Has an Identity Problem and Attackers Know It @ Bleeping Computer
- Microsoft extends Windows Server 2022 hotpatching until October 2027 @ Bleeping Computer
- IBM Says It Can Fit Nearly 100 Billion Transistors On a Chip @ Slashdot
- Asustor Showcases Flashstor 12 Pro Gen3 & Flashstor 6 Gen3 All-Flash NASes at Computex 2026 @ ServeTheHome
- Gooloo DS900 Automotive Diagnostic Tool Review @ NikKTech


