An AI agent should not read a table. It should call a tool that returns the table, and then do arithmetic on the rows. With GroupDocs.Parser.Mcp running locally in Docker, an AI client such as Claude Desktop or Cursor can extract tables from PDF, Word and Excel files and receive them as JSON:
Extract the line items from invoice.pdf as JSON and total the amount column.
The step-by-step version with config and troubleshooting is in the documentation: How to extract tables from documents into structured data.
What goes wrong when a model reads the table itself?
The common workflow is to paste a PDF into a chat and ask for the table. The model sees a flattened stream of text, infers where the columns were, and writes the rows back out. Every cell is then a prediction. A dropped column, a merged row or a shifted decimal does not produce an error; it produces a confident, well-formatted table that is slightly wrong. For a list of fruit that does not matter. For invoice totals it matters a great deal.
The failure is hard to catch because the output looks right. Nothing in a chat reply tells you which cells were read from the file and which were reconstructed.
What is the alternative to letting the model guess?
Move the reading out of the model. The extract_tables tool detects tables with the document engine and returns them as data:
- Markdown by default, as
### Table N (page N, R×C)sections that render directly in chat, VS Code and GitHub. - JSON with
format: "json", as an array of{ table, page, rows, columns, data }objects, wheredatais a grid of strings.
The model never transcribes a cell. It receives data and works with it, so its job shrinks to what it does reliably: summing, comparing, filtering and explaining. The tool supports PDF, DOCX, XLSX, PPTX and HTML, so the same prompt covers a supplier invoice, a Word report and a spreadsheet.
How do you get the table out as JSON?
State the format in the prompt.
Extract the line items as JSON and total the amount column.
Example session (abridged)
You: Extract the line items from invoice-014.pdf as JSON and total the amount column.
Agent: [calls extract_tables with format "json"]
{"table": 1, "page": 1, "rows": 4, "columns": 3,
"data": [["Item", "Qty", "Amount"],
["Paper A4", "10", "48.00"],
["Toner", "2", "170.00"],
["Delivery", "1", "22.00"]]}
Total of the Amount column: 240.00 across 3 line items.
Because the agent holds rows, the next steps are ordinary questions: which line is the largest, does the total match the figure printed in the text, write the rows to a CSV in the output folder. When the arithmetic matters, ask the agent to list the rows it summed so you can check them.
When should you ask for Markdown instead of JSON?
Use Markdown when a person reads the result next, and JSON when code does. Markdown is the default and is the fastest way to look at a table and judge whether it was found correctly. Switch to JSON for totals, cross-checks and anything that feeds a spreadsheet or a database. If you know where the table is, narrow the call:
Extract the table on page 4.
The page parameter restricts detection to one page, which is faster and, under metered licensing, cheaper.
Which cases still need a human look?
Honest limits matter more with tables than with plain text.
- Scans have no tables to find. Without a text layer there is no table structure, and the extraction comes back empty. There is no OCR in this server.
- Complex layouts vary. Merged cells, nested tables and tables that continue across pages do not always survive as cleanly as a simple grid. Spot-check one document before running two hundred.
- Evaluation mode is limited. Without a license, extraction is limited; the exact limits are on the library’s licensing page. Ask the agent to call
get_license_statusbefore a real run.
Install it in one command
The server is distributed as a Docker image only, so there is no dnx command. Register it in Claude Code like this, replacing the path with your documents folder:
claude mcp add groupdocs-parser -- docker run --rm -i \
-v /path/to/documents:/data ghcr.io/groupdocs-parser/parser-net-mcp:latest
Cursor, Claude Desktop, VS Code with GitHub Copilot, Windsurf, Cline and Codex use the same container with a client-specific config file; each is listed in Register in AI clients.
FAQ
Does Claude read PDF tables accurately?
That depends on the file and on how the client flattens it, so check it rather than assume. With extract_tables the rows come from the document engine’s table detection, returned as a grid of strings, and Claude works from those.
How do I extract a table from a PDF to Excel with AI? Ask for the table as JSON, then ask the agent to write the rows to a CSV in the output folder, which Excel opens. The conversion from rows to CSV is the agent’s step; the cell values come from the tool.
Can it extract tables from Word and Excel files locally?
Yes. extract_tables handles DOCX and XLSX as well as PDF, and the container reads files from the folder you mount.
Go deeper
- Documentation, canonical how-to: How to extract tables from documents into structured data
- Documentation hub: GroupDocs.Parser MCP Server
- Tool reference:
extract_tablestool reference - Start here: 3 ways to extract document data with AI agents and MCP (text, tables, images)
- Related: Enforce consistent data capture with automated extraction workflows using MCP
- Related: Scan-to-record automation with AI agents
- On-premise and security model: 3 architectures for AI document processing, and the one that keeps files inside your network
- Questions: GroupDocs Parser forum
- Source: GroupDocs.Parser.Mcp on GitHub