Let an agent read your PDFs
@sezzlee/pdf-mcp gives an agent read-only access to the PDF files in one folder, and to nothing
outside it. It has no tool that writes, so a document is never modified, and it never opens a network
connection.
You need Node.js 22 or later on macOS with Apple silicon, Linux with glibc or Windows x64.
1. Pick a folder
Everything under the folder is readable, subfolders included, so pick the narrowest one that holds the documents the agent needs. To follow along, create one and save annual-report.pdf into it: four pages of text, one of them a table.
mkdir -p ~/sezzlee-pdf2. Add the server to your client
Pass the folder as an absolute path. A JSON configuration is not read by a shell, so ~ is not
expanded there.
claude mcp add pdf -- npx -y @sezzlee/pdf-mcp /Users/you/sezzlee-pdfClaude Desktop reads claude_desktop_config.json (Settings → Developer → Edit Config, then
restart). Cursor reads .cursor/mcp.json in the project or ~/.cursor/mcp.json. VS Code reads
.vscode/mcp.json. Any other client that starts a stdio server works with the command npx and the
arguments -y, @sezzlee/pdf-mcp and the folder.
If the server does not start
Run the same command in a terminal. A relative path or a ~ in JSON exits at once with
The PDF source root '…' does not exist. On an Intel Mac the PDF engine cannot load: the server
prints that it could not load it on darwin-x64 and exits with code 3. On Alpine and other
musl-based Linux, every read fails with unsupported_platform, because file access goes through a
native module that has no build there. On Windows, escape the backslashes in JSON:
"C:\\Users\\you\\sezzlee-pdf".
3. Ask
Ask the agent: Which region had the most revenue in annual-report.pdf?
Expect three calls: list_documents to find the file, describe_document to learn that page 2
holds a table, and read_pages for that one page. The table comes back as Markdown, so the agent
reads the numbers by column. The answer is North, at EUR 1,240,000.
See what the agent receives
The recipes call the tools the way an agent does, through the MCP Inspector, so you can see each answer exactly. They need the server installed once and jq:
npm install -g @sezzlee/pdf-mcpDefine this helper in your shell; it reads ~/sezzlee-pdf:
pdf() {
npx -y @modelcontextprotocol/inspector --cli sezzlee-pdf ~/sezzlee-pdf \
--method tools/call --tool-name "$@" | jq '.content[0].text | fromjson'
}This is the call behind the answer above. Page numbers start at 1:
pdf read_pages --tool-arg filePath=annual-report.pdf 'pages=[2]' \
| jq -r '.pages[0].markdown'# 2. Revenue by region
Revenue is reported in euros, net of returns.
|Region|Revenue|Orders|Growth|
|---|---|---|---|
|North|1,240,000|3,410|+8.2%|
|South|980,500|2,875|+3.1%|
|East|1,105,250|3,020|+11.4%|
|West|612,800|1,940|-2.6%|
|Total|3,938,550|11,245|+5.9%|
West is the only region that shrank; see section 3 for the late delivery issue.What it can read
- Pages as Markdown, numbered from 1, with headings and tables kept as Markdown structure.
- A summary first: page count, whether the document is text or scanned, and exactly which pages have no readable text.
- Literal search over the text, with the page, the line and the surrounding text of each match.
- Scanned pages, when you start the server with an OCR binding. OCR is off unless a call asks for it.
A page with no readable text is never shown as an empty page: it is marked needsOcr. See
Read scanned pages with OCR.
Why it cannot read outside the folder
The server opens the folder once at startup and resolves every path against that handle in a native
module, not by comparing strings. A .. that climbs out, an absolute path, or a symbolic link that
leads outside is refused with path_outside_root. That check runs before the existence check, so
error codes cannot be used to map your disk, and the folder's absolute path is removed from every
error.
The server has no network code, so the links and references inside a PDF are never followed. A
document is read as one snapshot of its bytes; a file that changes mid-read fails with
file_changed. A file over 32 MiB, a document over 2,000 pages and an extraction that runs past 20
seconds are refused, and a password-protected document fails with encrypted_pdf. These are budgets
that keep one file from exhausting the server, not a sandbox around the PDF engine.
Two things are outside what it can defend: a hard link inside the folder is the same file as its other name, and someone who can mount a filesystem over the folder can change what it holds.