Every security analyst has been there. An email lands in the queue with an attachment that just doesn't feel right. The sender claims it's an invoice, but the filename is weird — INVOICE_9843721.pdf.exe or maybe something subtler like Invoice_SCAN_2026.pdf. Your gut says malicious, but you need proof before escalating.

That's where PDF analysis comes in. PDFs are surprisingly complex documents — they're almost mini filesystems with their own scripting, multimedia handling, and network capabilities. Attackers love this complexity because it gives them plenty of room to hide. Let me walk you through how to safely analyze suspicious PDF attachments and figure out if they're dangerous.

First: Never Open It Directly

This should go without saying, but it bears repeating: don't double-click that attachment. Modern malware often uses PDF reader zero-days, meaning even an up-to-date system can get pwned just by opening a specially crafted file. Your analysis environment should be isolated — a VM with no network access, snapshots you can roll back, and definitely no integration with your host system.

Set up a dedicated analysis workstation. Snapshot it clean. Disconnect from the network. Only then start poking at the file.

Step 1: Basic Static Analysis

Start with the easy stuff. Before you even open the file, you can learn a lot from its metadata and structure.

Check the file header. A PDF starts with %PDF-. If it starts with something else — like MZ (executable) or a ZIP signature — but claims to be a PDF, that's your first red flag. Attackers sometimes disguise executables as PDFs to trick users and security tools.

Run pdfid.py. This tool (from Didier Stevens) scans a PDF and reports counts of suspicious elements: /JavaScript, /AA (auto-action), /OpenAction, /AcroForm, /URI, and more. A blank invoice shouldn't have any of these. If you see JavaScript or OpenAction, dig deeper.

Look at object count. Use peepdf -i suspicious.pdf to load the PDF interactively. Check how many objects it has. A normal PDF might have 50-200 objects. Malware-laced PDFs can have thousands — especially if they're packing additional payloads or using obfuscation techniques that create redundant objects.

Step 2: Extract and Inspect Streams

PDFs store content in streams — compressed data chunks that can contain anything from text to images to JavaScript code. This is where the action happens.

Use peepdf to list all streams and their filters. Look for streams with /Filter /FlateDecode (compression) — that's normal. But streams with multiple nested filters or unusual encoding should raise eyebrows. Attackers sometimes layer filters to hide malicious code.

Extract suspicious streams using peepdf's stream command or write them to disk for further analysis. If you find JavaScript, deobfuscate it. Look for common malware patterns: app.launchURL, this.getURL, eval, or strings that look like URLs pointing to suspicious domains.

Step 3: Check for Embedded Files

PDFs can contain other files inside them — executables, archives, additional PDFs, anything. This feature has legitimate uses (technical manuals with datasheets attached), but it's also a popular malware vector.

Run pdf-parser.py with the --search /EmbeddedFiles flag to find embedded content. If your suspicious PDF is an invoice but contains an .exe file, you have your answer. The user doesn't even need to "open" the PDF — just previewing it in some mail clients can trigger automatic extraction.

Step 4: Analyze Behavior in a Sandbox

Static analysis gets you far, but sometimes you need to see what happens when the file actually runs. That's what sandboxes are for.

If your organization has a malware sandbox that accepts PDF files, submit the sample there. Look for:

If you don't have a commercial sandbox, you can do basic behavioral analysis manually in an isolated VM. Use tools like Process Monitor and Wireshark to watch what happens when you open the PDF in a vulnerable reader. Just make sure you're completely isolated first.

Step 5: Look for Exploit Signatures

PDF exploits typically target vulnerabilities in Adobe Reader, Foxit, or other popular viewers. These exploits often leave fingerprints:

Identifying these manually requires deep expertise, but you can use tools like pdfid, peepdf, and commercial sandboxes to flag potential exploits. If the PDF triggers an exploit in a sandbox — even if the sandbox catches and blocks it — treat the file as malicious.

A Quick Checklist

When you're in a hurry, here's what to check in order:

  1. File header — does it actually start with %PDF-?
  2. pdfid.py output — any JavaScript, OpenAction, or URI elements?
  3. peepdf entropy — unusually high entropy suggests packing or encryption
  4. Embedded files — any surprise attachments?
  5. Stream content — does anything look like malicious script?
  6. Sandbox behavior — does it do anything suspicious?

Final Thoughts

PDF malware isn't going away. It's one of the most reliable delivery mechanisms for targeted attacks and mass campaigns alike, precisely because PDFs are so common and because users trust them. As a security professional, knowing how to analyze these files is essential.

The techniques above will help you handle most suspicious PDFs you encounter. Start with static analysis — it's safe and fast. Escalate to behavioral analysis only when needed. And when in doubt, treat the file as malicious until proven otherwise. That's how you stay ahead of the attackers.