Back to Search View Original Cite This Article

Abstract

<title>Abstract</title> <p>Indirect prompt injection is now a well-known threat: an attacker plants an instruction in a document; a downstream LLM app reads the document; the model follows the attacker instead of the developer. The literature focuses on what the model does once it has seen the payload. We ask the layer below: which extraction pipeline surfaces the hidden text, and which does not. We build carrier files across multiple formats (PDF, DOCX, DICOM, Parquet, safetensors, PNG, SVG, WAV, Markdown, JPG, MP4) that hide the same benign canary payload with distinct concealment techniques. We measure payload visibility across extraction pipelines and model tiers, finding that hidden prompts surface through metadata channels, binary headers, and structured format fields that extraction pipelines do not sanitize. The work maps the hidden-injection surface across 11 file formats and multiple Gemini tiers, providing a measurement framework for extraction pipeline vulnerability.</p>

Show More

Keywords

extraction model payload attacker document

Related Articles

PORE

About

Connect