Most people think of PDFs as straightforward documents — pages of text, images, and maybe some forms. But the PDF format is surprisingly flexible. Behind the scenes, there's a entire layer of data that most users never see: metadata, embedded objects, JavaScript, and hidden content streams. Attackers have figured out how to exploit this flexibility for steganography — the art of hiding data within other data.
In this post, we're going into the world of PDF steganography. You'll learn how attackers hide malicious code, exfiltrate sensitive data, and use PDFs as covert channels to bypass security controls.
What Is PDF Steganography?
Steganography comes from Greek words meaning "covered writing." Unlike encryption, which makes data unreadable but noticeable, steganography hides the existence of the data itself. A PDF file is the perfect carrier because it's complex, widely used, and rarely inspected closely.
The PDF format allows multiple streams, objects, and metadata fields. Attackers can hide data in places that most PDF readers ignore entirely. This makes PDFs an attractive option for anyone trying to smuggle data past network filters, email scanners, or endpoint protection.
Common Hiding Techniques in PDFs
1. Metadata Manipulation
Every PDF has metadata — author name, creation date, software used, and more. Most people never look at this data, which makes it an ideal place to stash information. You can embed Base64-encoded strings, URLs, or even small pieces of executable code in metadata fields. The content won't affect how the document looks, but anyone inspecting the file closely will find something unexpected.
2. Unused Object Streams
PDFs organize content into objects, numbered and referenced throughout the file. Not all objects are actually used when rendering the document. Attackers can create extra objects containing hidden data, then reference them in ways that never get executed. These ghost objects sit in the file, invisible to normal viewing but easily readable by anyone who knows where to look.
3. Image-Based Hiding
PDFs often contain images — sometimes dozens in a single file. You can hide data in images using techniques like least significant bit (LSB) modification. By tweaking the tiniest details of pixel colors, you can embed data without any visible change to the image. When extracted and processed with the right tools, these modified images reveal their hidden contents.
4. Incremental Updates
Here's one that catches even experienced analysts: PDFs can be updated incrementally. When you save changes to a PDF, the old version isn't necessarily deleted — it's just marked as obsolete. This means earlier versions of the file, with their original (potentially malicious) content, can still exist inside the current file. forensic analysts call this "carving" — extracting the hidden previous versions from inside the current PDF.
5. Invisible Layers
PDF supports optional content groups (OCGs), which let you show or hide content based on user preferences. Attackers can create layers that are invisible by default but contain malicious URLs, scripts, or instructions. These layers aren't rendered on screen, so the victim never sees them, but the data is still there in the file.
Why Attackers Use PDF Steganography
The primary motivation is evasion. Security products that scan attachments often look for obvious threats — known malware signatures, malicious attachments, or suspicious file types. A PDF that looks completely normal but carries hidden instructions can slip right through.
Data exfiltration is another major use case. An attacker inside a compromised network can encode stolen data into PDFs and send them out via email or file transfer. The data looks like a harmless document to network monitors, making outbound data theft harder to detect.
Some advanced threats use PDFs as command-and-control channels. The PDF contacts a server, receives instructions hidden in seemingly innocent content, and executes accordingly. This technique lets attackers maintain persistence using a file type that most security tools trust.
How to Detect Hidden Data in PDFs
Finding steganography in PDFs requires looking beyond what renders on screen. Here are the techniques that work:
- Use pdfid.py — This tool from Didier Stevens scans a PDF and reports all the suspicious elements: JavaScript, OpenAction, JS, Launch, and other potentially dangerous objects
- Extract all objects — Tools like peepdf can parse a PDF and list every single object, even the unused ones
- Check metadata — Use pdfinfo or exiftool to dump all metadata fields and look for anything unusual
- Compare incremental versions — If a file has been saved multiple times, extract each version and compare them
- Analyze images separately — Extract embedded images and run steganography detection tools on them
Defending Against PDF-Based Covert Channels
Detecting steganography is hard because the goal is literally to hide in plain sight. But you can reduce your risk:
- Disable JavaScript in PDF readers — Most PDF steganography attacks rely on embedded scripts to activate hidden content
- Use advanced PDF analysis tools — Enterprise security solutions that unpack and inspect PDFs beyond surface-level scanning
- Monitor network traffic for unusual PDF uploads — Large outbound PDFs from sensitive systems should trigger alerts
- Implement DLP controls — Data loss prevention can flag PDFs containing sensitive patterns even if hidden
The Bottom Line
PDF steganography is real, it's used in the wild, and it's notoriously difficult to detect. The PDF format's complexity — the very feature that makes it so useful — also makes it an ideal vessel for hidden data. Security professionals need to understand these techniques not just to detect threats, but to appreciate how sophisticated attackers have become.
If you're analyzing suspicious PDFs, never judge a file by what you see on screen. The real threat might be hiding in the parts you'll only find if you look deeper.