Google and ChatGPT Vulnerable to Indirect Prompt Injection That Spreads False Information
“Indirect prompt injection” involves hiding misleading instructions or false content on webpages so that large language models accept and repeat the material when retrieving information. As Google AI Search, OpenAI’s ChatGPT and Gemini become gateways to information, such attacks could distort medical, financial and election-related content and influence consequential user decisions.
A BBC reporter’s experiment showed that a blog post falsely claiming its author was a hot dog-eating champion could persuade Google AI, ChatGPT and Gemini to accept the claim within about 20 minutes. The reports warned that with Google Search reaching about 2.5 billion people, a single AI-generated “answer” could rapidly amplify misinformation. Google has formally classified the manipulation of AI responses as a violation of its search policies.
All Coverage
2 original reportsThe Backstory
The history behind this eventNo historical echoes for this signal
Subscribe to Mark Radar Weekly
Every Friday, the week's strongest signals in your inbox. Unsubscribe anytime.