JaiLIP, A Picture Is Worth 1000 Lines Of Code To An LLM

Source: Slashdot JaiLIP, A Picture Is Worth 1000 Lines Of Code To An LLM

Remember Steganography?

The majority of attacks against LLMs come from text based prompts that will cause the agent to ‘jump the rails’ and start providing responses it shouldn’t.  An older example of this was asking an LLM to tell them to compose a song like their grandma used to song for them, a song which happened to be about making napalm.  JaiLIP has the same effect, but the attack vector is a picture which is used against LLMs capable of processing vision-language model tasks.

JaiLIP, which stands for Jailbreaking with Loss-guided Image Perturbation is a way of manipulating an image in a way that is invisible to the naked eye but significant to an LLM.  For instance, the researchers “modified image of a traffic light. While the image appeared ordinary to human viewers, it reportedly influenced the model to provide instructions for running a red light while avoiding a traffic ticket“.  That is not a response the LLM should provide.

You can read the entire research paper here.

The findings highlight a potential security risk for businesses deploying AI systems that process both images and text. While most discussions about AI safety focus on prompts, the research suggests that seemingly harmless images may also serve as an attack vector.

Video News

About The Author

Jeremy Hellstrom

Call it K7M.com, AMDMB.com, or PC Perspective, Jeremy has been hanging out and then working with the gang here for years. Apart from the front page you might find him on the BOINC Forums or possibly the Fraggin' Frogs if he has the time.

Leave a reply

Your email address will not be published. Required fields are marked *

This site uses Akismet to reduce spam. Learn how your comment data is processed.

Latest Podcasts

Archive & Timeline

Previous 12 months
Explore: All The Years!