When you open a PDF in your favorite reader, you see text, images, and maybe some forms. What you don't see is the underlying data structure — streams of compressed content, encoded objects, and potentially hidden payloads. Here's the uncomfortable truth: attackers frequently hide malicious code inside PDFs using compression and obfuscation techniques that make traditional signature-based detection useless.

That's where entropy analysis comes in. It's a statistical technique that measures the randomness of data in a file. High entropy usually means compressed or encrypted content. Low entropy suggests plain text or repetitive data. By analyzing entropy patterns across a PDF, you can spot sections that stand out — and those anomalies often point to malware.

What Is Entropy, Anyway?

Entropy, in the context of information theory, measures how unpredictable the bytes in a file are. A value of 0 means the data is completely predictable (all the same byte). A value of 8 represents maximum randomness. Most normal PDF content — text, images, layout instructions — falls somewhere between 2 and 5. Compressed streams jump to 6-7. Encrypted content or packed executables can hit 7.5 or higher.

Think of it like this: if you read a PDF and every other word was random gibberish, you'd notice. Entropy analysis does the same thing mathematically — it flags data that doesn't match what you'd expect from a typical document.

Why Attackers Pack PDF Payloads

Malware authors don't want their code to be found. They use packers — tools that compress and sometimes encrypt the malicious payload. The packed result looks like random noise to signature scanners. When the victim opens the PDF, the unpacking routine runs, extracts the real payload, and executes it.

In PDF attacks, you'll see this with embedded JavaScript. The script might be compressed using the PDF's built-in FlateDecode filter, then further obfuscated with XOR operations or other transforms. The result looks like garbage bytes — high entropy — sitting inside what appears to be a normal document.

This technique works because most PDF scanners check for known bad signatures. If the malicious code is hidden behind compression and never hits the disk in its raw form, traditional antivirus often misses it.

How to Perform PDF Entropy Analysis

1. Use peepdf

The go-to tool for interactive PDF analysis. Load a suspicious file and use the entropy command. Peepdf visualizes entropy as a graph, showing you which objects have high randomness. Objects with entropy above 6.0 deserve closer inspection.

2. Calculate Entropy Per Object

PDFs are built from objects — numbered chunks that contain streams, dictionaries, and references. Extract each object separately and calculate its entropy. Normal text objects stay low. Compressed streams with hidden payloads spike high.

3. Look for Section Boundaries

Entropy analysis works best when you scan the entire file as a byte stream. Plot entropy values along the file's length. Normal PDFs show consistent, moderate entropy with occasional spikes in compressed image areas. A malware-laden PDF often has one or more sections with unusually high entropy — the packed payload.

4. Combine with Other Techniques

Entropy alone won't give you a definitive verdict. Combine it with other checks: look for suspicious JavaScript with pdfid.py, search for OpenAction objects that auto-execute on open, and examine any streams using the /RichMedia or /Launch keys. High entropy is a red flag, but context matters.

Practical Example: Spotting a Compressed Payload

Let's say you download a PDF that claims to be an invoice. You run pdfid.py and notice it has JavaScript and OpenAction — two things invoices rarely need. Then you run peepdf and see entropy values hovering around 7.2 in object 5.

That's your signal. Normal PDF text rarely exceeds 5.0 entropy. Something in object 5 is heavily compressed or encrypted. You extract that object, decompress it if possible, and find a JavaScript payload that downloads additional malware from a remote server.

This is exactly how security researchers catch PDF-based attacks. The entropy spike is the tip-off that something's hidden.

Tools for Entropy Analysis

Limitations to Know About

Entropy analysis isn't foolproof. Some legitimate PDFs use compression heavily — especially scanned documents converted to PDF/A with embedded images. You'll get high entropy in those cases even though there's nothing malicious. The key is comparing entropy against what you know about the file type and looking for anomalies within the specific PDF structure.

Advanced attackers also know about entropy detection. Some packers add padding or structured noise to normalize entropy values. This makes analysis harder but not impossible — determined analysts still find the unpacked payload through behavioral analysis and code inspection.

The Bottom Line

PDF entropy analysis is a powerful technique in your security toolkit. It doesn't replace signature-based detection or behavioral analysis, but it excels at finding things those methods miss — especially packed malware hiding inside compressed streams. When combined with other PDF forensics techniques, entropy analysis helps you dig deeper and spot threats that would otherwise slip through.

If you're handling suspicious PDFs — whether in a SOC, during incident response, or just triple-checking an attachment — run an entropy check. That spike in randomness might be the only signal you get before the malware executes.