Say you have a folder of contracts and need the first page of each, or the signature page, for a review. An AI agent with the GroupDocs.Merger.Mcp server does it in one prompt, locally on your machine, from Claude Desktop, Claude Code, Cursor or VS Code with GitHub Copilot:
Extract page 1 from every PDF in my documents folder and list what you produced.
The step-by-step version with config and troubleshooting is in the documentation: How to extract pages from many documents at once.
What does the agent do with a folder of files?
For each file, the agent calls the split tool once. Each call writes a single-page document. A folder of 40 invoices therefore yields 40 cover pages. The tool has no folder mode; the agent supplies the loop, and the engine does the page extraction for every file. Files are resolved by name inside the folder configured in GROUPDOCS_MCP_STORAGE_PATH, so the agent passes invoice-014.pdf, not a path from your machine.
How do you keep the extracted pages out of the originals’ folder?
Forty extracted pages beside forty originals is a folder nobody can navigate. Set GROUPDOCS_MCP_OUTPUT_PATH to a separate directory, and results land apart from sources. Without it, output goes to the storage folder. This is a one-time setting, and every bulk prompt after it stays tidy.
How do you extract the last page, or a page that moves?
“The last page” is a different number in every file. Have the agent call get_document_info for each document first. It returns pageCount, so the agent can compute the page number per file:
For each PDF in the folder, check the page count, then extract the last page of each one.
An agent that guesses fails on the files where the guess is wrong. The error is per file, not fatal.
What variations come up most?
Extract the last page of each document. Check the page count first.
Pull page 2 out of every contract, then merge all of them into one review file.
For each file, extract the signature page and put the results in the output folder.
The second prompt chains into merge in rounds of up to four documents, which gives you a single combined document built from one page of each source. More than four sources means more rounds. The signature-page prompt works when you tell the agent which page holds the signature; the tool extracts the page numbers you give it and does not search for content.
When is one prompt per folder the right size?
Use it when the same rule applies to every file: first page, last page, a known signature page. The prompt states the rule once, and the agent applies it file by file. Archive and legal-review work fits this shape: a cover sheet from each file for an index, or the page that carries signatures from each contract for a check. If the rule differs per file, split the folder into groups and run one prompt per group, so each prompt still states a single rule. Keep the instruction “list what you produced” in every bulk prompt; the list is your record of which files were handled.
Example session (abridged)
An illustration based on the documented tool behavior, not a captured screenshot.
You: Extract page 1 from every PDF in my documents folder and
list what you produced.
Agent: Found 3 PDFs in the folder.
[calls split on each file with pages = "1"]
Extracted page 1 from all 3 files: 3 single-page documents
in your output folder. No failures.
What goes wrong in a batch of many files?
- Failures are per file. A corrupt or password-protected document fails its own call. Ask for the failures as a list so a batch of 40 does not hide two errors. The
splittool accepts apasswordargument for protected documents. - A short document. Asking for a page that does not exist fails for that file; the other files are unaffected.
- Metered usage adds up. Under metered licensing, every call is billed as processing.
Honest limits
- Evaluation mode trims merged output to 3 pages. The result is trimmed to the first 3 pages, with a trial badge on each page. Extraction of single pages is unaffected, but the moment you merge the results the cap applies silently. A merged review file built from forty first pages comes back as three. Check
get_license_statusbefore the run. - No content search. The agent picks pages by number. Finding “the page with the signature” is your instruction, not a tool feature.
- No folder mode in the tool. One
splitcall handles one file, so a big folder means many calls.
FAQ
How do I extract the first page from all PDFs in a folder? Point the server at the folder, then ask: “Extract page 1 from every PDF in my documents folder and list what you produced.” The agent calls split for each file.
Can AI batch split documents and collect the pages in one file? Yes. After extraction, ask the agent to merge the single-page files in rounds of four. In evaluation mode that merged file is trimmed to 3 pages.
Do the files leave my machine? No. The server reads and writes the folders you configure, over local stdio transport, with no inbound ports.
Go deeper
- Documentation, canonical how-to: How to extract pages from many documents at once
- Documentation hub: GroupDocs.Merger MCP Server
- Start here: Why your AI agent should merge files with an engine, not by regenerating them
- Related: Enforce document pack assembly order with automated AI workflows using MCP
- Related: 3 ways to rebuild a document from just the pages you need via MCP
- On-premise and security model: 3 architectures for AI document processing, and the one that keeps files inside your network
- Questions: GroupDocs Merger forum
- Source: GroupDocs.Merger.Mcp on GitHub