To bulk extract data from documents with an AI agent, write the schema into the prompt, point the agent at a folder, and require an exceptions list at the end. The GroupDocs.Parser.Mcp server does the reading locally in a Docker container, so every file passes through the same extraction tools and no document leaves the machine. The LLM decides the order of work; the engine reads the files.

For every PDF in my documents folder: extract the tables as JSON, total the amount column, and give me one table of file name, total, page count.

The step-by-step version with config and troubleshooting is in the documentation: How to extract data from a folder of documents at once.

Why does manual extraction produce inconsistent data?

When people copy values out of documents one at a time, each file is handled slightly differently: one person includes tax, another does not, a third skips the file with the odd layout. The inconsistency is invisible until the totals are compared. A workflow fixes this by making the rules explicit and applying them to every file in the same way. With an agent the rules live in one prompt, and the extraction tools behave identically on file one and file two hundred.

How do you build the workflow?

Mount the folder into the container as /data, set GROUPDOCS_MCP_STORAGE_PATH=/data, and give the agent these steps.

  1. Check the license. Prompt: “Run get_license_status and tell me the mode.” Unlicensed extraction is limited, and a partial read across a folder is a quietly wrong dataset.
  2. Triage the folder. Prompt: “For each file, report format, page count and whether text extraction returns anything.” get_document_info supplies type, page count and size, and a one-page extract_text call shows whether text exists.
  3. Fix the schema. Prompt: “For each readable file, extract the line-items table as JSON. Columns: item, quantity, amount. Total the amount column.” This uses extract_tables.
  4. Collect failures explicitly. Prompt: “List every file that failed, was password-protected, or returned no table, with the reason.” A batch that reports only successes hides the exceptions.
  5. Write the output. Prompt: “Write one row per file to a CSV in the output folder.”

Files are resolved by name inside the container, so the agent passes invoice-014.pdf, not a host path.

What does the agent decide, and what does the engine do?

Step The agent decides The engine does
Triage Which files to extract from Reports type, pages and size; extract_text shows whether text exists
Capture Which tool each file needs Returns text or table rows from the file
Validation Whether a total matches the text Nothing; the agent compares
Exceptions Which files go on the list Returns the failure message for a bad file
Output The shape of the CSV Nothing; the agent writes it

The split is the point. Capturing the values is deterministic engine work; judging and assembling them is the agent’s.

Example session (abridged):

You:    Check the license, triage the folder, then extract line items from every readable PDF and list the exceptions.
Agent:  [calls get_license_status, get_document_info x12, extract_text, extract_tables x10]
        License mode: evaluation. Results below are a smoke test.
        12 files: 10 with a text layer, 1 scan, 1 password-protected.
        Totals written for 10 files. Exceptions: scan-007.pdf (no text layer), hr-020.pdf (password required).

What happens to files that do not fit?

A scan has no text layer, so table extraction returns nothing for it. Set those aside, or let the agent branch to extract_barcodes when the page carries a tracking number or order code. The fall-through list doubles as an inventory of which parts of the archive are scans. A password-protected file fails its own call, so ask the agent to report it and continue; the tools accept a password argument when you have one.

How do you keep a big run affordable?

  • Restrict by page when the data has a known location, such as the first page of each invoice.
  • Avoid extract_images in bulk. It writes a file for every image, and a folder of illustrated reports produces hundreds.
  • Watch metered usage. Under metered licensing every call is billed processing, so a targeted sweep costs less than an exhaustive one.

What are the limits?

Without a license, extraction is limited; the exact limits are on the library’s licensing page. The server has no OCR, so scanned text cannot be captured. It is delivered as a Docker image only; there is no dnx command.

How do you register the server?

A minimal manual entry for Cursor, Claude Desktop or Windsurf, using the container paths from the Parser configuration docs:

{
  "mcpServers": {
    "groupdocs-parser": {
      "command": "docker",
      "args": ["run", "--rm", "-i",
               "-v", "/path/to/documents:/data",
               "-e", "GROUPDOCS_MCP_STORAGE_PATH=/data",
               "ghcr.io/groupdocs-parser/parser-net-mcp:latest"]
    }
  }
}

FAQ

Can I process a folder of invoices with AI without uploading them? Yes. The container reads the mounted folder and talks to your client over stdio, so the invoices stay on your machine.

How do I batch extract text from PDFs locally? Ask the agent to loop over the folder and call extract_text per file, page by page for long documents, because very large outputs are truncated with a marker.

Can the agent automate data entry from documents? It can assemble the extracted rows into a CSV or a summary table. Review the exceptions list before anything is entered into another system.

Go deeper