PDF JavaScript Extractor
Extract and analyze JavaScript embedded in PDF files. Upload a PDF to scan for OpenAction scripts (auto-execute on document open), document and page-level Additional Actions (AA), Names tree JavaScript entries, and raw JS streams. View extracted code with source context, copy to clipboard, or download for further analysis. All processing is local and private with no signup required.
Drop your PDF here
or click to browse
Unencrypted PDF files only
Upload a PDF file to scan for embedded JavaScript. The tool detects OpenAction scripts (auto-execute on open), Additional Actions, named JavaScript entries, and raw JS streams. All processing is local — your file never leaves your device.
Why Use Our PDF JavaScript Extractor?
Multi-Source JavaScript Detection
Detects JavaScript from all PDF sources: OpenAction (auto-execute on document open), document and page-level Additional Actions (AA), Names tree entries, and raw embedded JS streams. No script goes unnoticed.
Code Viewer with Context
View extracted JavaScript in a syntax-highlighted code viewer with source location, trigger event, and page number. Each entry shows its origin (OpenAction, Page AA, etc.) and trigger condition for full context.
Obfuscation & Threat Detection
Automatically flags obfuscated or minified JavaScript. OpenAction scripts are highlighted as high-priority threats since they execute automatically. Suspicious patterns like eval(), ActiveXObject, and Base64 encoding are detected.
100% Browser-Local & Private
All PDF parsing and JavaScript extraction runs entirely in your browser using pdfjs-dist and raw binary analysis. Your PDF files and extracted code never leave your device. No uploads, no servers, no tracking.
Common Use Cases for PDF JavaScript Extractor
Malware & Phishing PDF Analysis
Security analysts can scan suspicious PDFs for malicious JavaScript that executes on open. OpenAction scripts are a common vector for drive-by downloads, credential phishing, and malware delivery. Quickly identify and extract the embedded script for further analysis.
Security Auditing & Forensics
Audit PDF documents for hidden or unauthorized JavaScript actions. Detect auto-execute scripts, form validation triggers, mouse-event handlers, and print/document actions that could exfiltrate data or modify document behavior without user knowledge.
Reverse Engineering PDF Forms
Extract and analyze JavaScript from PDF forms with complex validation, calculation, or formatting scripts. Understand how form fields interact, what data is processed, and whether the form performs any network requests or file operations.
Vulnerability Research & Testing
Researchers can examine PDF-based JavaScript for browser plugin exploits, sandbox escape attempts, and unusual API usage. Extract and study real-world PDF JavaScript samples for vulnerability research and detection rule development.
Document Sanitization Prep
Before sanitizing a PDF, use this tool to locate and review all embedded JavaScript. Understand what each script does so you can decide whether to keep or remove it. Pairs perfectly with PDF sanitization tools for thorough document cleaning.
Education & Training
Learn about PDF JavaScript internals by examining real examples. Understand OpenAction, Additional Actions, and named JavaScript entries. A valuable resource for security courses, digital forensics training, and PDF format education.
Understanding JavaScript in PDF Files
What Is PDF JavaScript?
PDF JavaScript is a scripting feature defined in the PDF specification that allows developers to embed JavaScript code within PDF documents. This code can respond to document events (opening, closing, printing), page events (entering, exiting), form events (keystrokes, value changes, formatting, validation), and mouse events. PDF JavaScript uses a subset of the JavaScript language with additional objects specific to Acrobat (app, doc, console, security, etc.). While legitimate PDFs use JavaScript for form validation, auto-calculation, or interactive features, the same capability is frequently exploited by malware authors to create malicious PDFs that execute automatically upon opening — making JavaScript detection a critical security capability.
How Our PDF JavaScript Extractor Works
The tool uses a two-pronged approach combining pdfjs-dist parsing with raw binary scanning for thorough JavaScript extraction:
- pdfjs-dist structured extraction: Loads the PDF using the pdfjs-dist library and traverses the document structure to find JavaScript entries. This includes checking the catalog for OpenAction entries (scripts that run when the document opens), scanning the Names dictionary for named JavaScript actions, examining document-level Additional Actions (AA) for save/print events, and checking each page for page-level AA entries.
- Raw binary scanning: Parses the raw PDF bytes using regex-based pattern matching to find JavaScript-related keywords and structures that pdfjs-dist might not expose through its high-level API. This includes searching for /JavaScript references, /OpenAction entries, /AA dictionaries containing /JS actions, and stream content with JavaScript-like code patterns.
- Analysis and classification: Each extracted entry is classified by source type (OpenAction, DocumentAA, PageAA, NamesTree, RawStream), trigger condition, and page number. The code is analyzed for obfuscation patterns and suspicious API usage. Results are displayed in a sortable list with a code viewer for detailed inspection.
Types of PDF JavaScript Sources
- OpenAction (Critical): The most dangerous source — JavaScript that executes automatically when the PDF is opened, without any user interaction. Commonly used for malware delivery and phishing attacks. Instantly flagged as high-priority.
- Document Additional Actions (AA): JavaScript triggered by document-level events like WillSave, DidSave, WillPrint, DidPrint, and DocumentClose. These scripts run in response to user actions but can be used for data exfiltration or persistent behavior.
- Page Additional Actions (AA): JavaScript attached to specific pages, triggered by page-level events such as PageOpen, PageClose, and mouse enter/exit. Can execute as the user navigates through the document.
- Named JavaScript (Names Tree): JavaScript actions registered in the document's Names dictionary with unique identifiers. These can be referenced by other actions or triggered programmatically. Often used for complex form logic.
- Raw Stream JavaScript: JavaScript detected through raw binary scanning of compressed or obfuscated streams. These may be embedded in non-standard locations or use encoding that hides them from high-level parsers.
Privacy & Security
This tool runs entirely in your browser using client-side JavaScript with pdfjs-dist for PDF parsing. Your PDF files, all extracted JavaScript code, scan results, and any data you copy are never uploaded to any server, stored in any database, or transmitted over the network. All PDF parsing, pattern matching, obfuscation detection, and code extraction execute locally on your device. There are no API calls, analytics tracking, cookies, or data collection of any kind. This makes it completely safe for analyzing sensitive, proprietary, or potentially malicious PDF documents.
Related PDF & Security Analysis Tools
PDF Sanitizer / Remove Hidden Content
Remove JavaScript, embedded files, metadata, annotations, and form fields from PDFs. Sanitize potentially malicious PDF documents before distribution.
PDF Stream Decoder
Browse and decode PDF content streams, view object structures, and inspect compressed data. Low-level PDF inspection tool for forensic analysis.
Extract Embedded Files from PDF
Scan PDF documents for embedded file attachments and extract them. Download individual files or all attachments as a ZIP archive.
PDF Metadata Extractor
View and analyze PDF metadata including author, creation date, software, XMP data, and document statistics from any PDF file.
JavaScript Deobfuscator / Unpacker
Deobfuscate JavaScript code: detect hex strings, Unicode escapes, Base64, number obfuscation. Side-by-side cleaned output with detected techniques.
PDF Structure Analyzer
Analyze PDF internal structure: pages, fonts, images, annotations, form fields, cross-reference table, and document health assessment.
Frequently Asked Questions About PDF JavaScript Extractor
PDF JavaScript allows embedded scripts to run when a PDF is opened, printed, or interacted with. While legitimate uses include form validation and calculations, cybercriminals frequently weaponize PDF JavaScript for malware delivery, phishing, and drive-by downloads. Detecting and extracting this JavaScript is a critical first step in PDF security analysis.
An OpenAction script executes automatically the moment a PDF is opened, with zero user interaction required. This makes it the most dangerous type of PDF JavaScript. Attackers use OpenAction to launch phishing pages, download malware, exploit browser vulnerabilities, or redirect users to malicious sites without any visible indication.
No. Encrypted PDFs cannot be parsed because their internal structures are inaccessible without the password. You must decrypt or unlock the PDF first before using this tool. The tool will report an extraction failure for encrypted documents.
No. The tool only reads and displays the extracted JavaScript code — it never executes it. The code is displayed as plain text in a read-only viewer for analysis. You are completely safe as long as you do not copy and run the extracted code outside of this tool.
The tool scans for common obfuscation patterns including eval() calls, ActiveXObject creation, String.fromCharCode usage, excessive Unicode escapes, Base64 encoding, and minified code with very high character-to-newline ratios. Entries matching multiple suspicious patterns are flagged as obfuscated with a warning badge.
JavaScript support was introduced in PDF 1.3 (Acrobat 4.0) and expanded in later versions. Any PDF version 1.3 or higher may contain embedded JavaScript. The tool works with all PDF versions up to 2.0 and can detect JavaScript across the entire version range.
The tool provides context for each entry including its source (OpenAction, AA, etc.), trigger event, and page number. OpenAction scripts are automatically flagged as high-priority. Obfuscation detection and suspicious pattern matching help identify potentially malicious code. However, you should analyze the actual script logic to determine intent definitively.
Some PDF producers use non-standard structures, compression, or encoding that hides JavaScript from standard parsers like pdfjs-dist. The raw binary scan uses regex pattern matching directly on the PDF bytes and can catch scripts that are obfuscated, compressed in unexpected ways, or placed in non-standard locations within the PDF structure.
Yes — 100% free with no signup, no account, and no usage limits. Scan and extract JavaScript from as many PDFs as you want. All processing is local in your browser with no data uploads.