You can let an AI agent look at a document page instead of reading a text dump of it. The GroupDocs.Viewer.Mcp server renders any page of a PDF, Word, Excel or PowerPoint file as a PNG and returns it to the agent, and it does so locally on your machine:

How many pages does report.pdf have? Then show me the last one.

The step-by-step version with config and troubleshooting is in the documentation: How to let an AI agent see document pages.

Why does text extraction fail on real documents?

A text extractor returns characters in reading order. It does not return the bar chart between two paragraphs, the “APPROVED” stamp in the corner, the column layout of a price table, or a page that is a scan and has no characters at all. An agent working from that output answers confidently about a document it has only half seen, and the missing half is usually the part you asked about.

The fix is not a better prompt. It is a different input: the page as an image, given to a model that accepts image input.

Wrong approach: pipe everything through a text extractor

The common default is to convert the document to text, paste the text into the context window and ask questions. It works for prose. It breaks in four ways that you can verify on any document you own:

  • Layout is gone, so “is the total in the right-hand column?” cannot be answered.
  • Figures inside charts and diagrams are gone, because they are drawn, not typed.
  • Stamps, signatures and handwriting are gone.
  • A scanned PDF yields no text.

A model cannot repair information that was dropped before it saw the document.

Right approach: render a PDF page to an image for the LLM

render_page renders one page as a PNG, returns the image inline to the agent and saves a copy to your output folder as <name>_page<N>.png. The server supports 170+ formats, including PDF, DOCX, XLSX, PPTX, ODT, RTF, HTML, EML and MSG. A vision-capable model then sees the page the way you would.

The agent should not guess page numbers, so it pairs render_page with get_view_info, which returns the file type, the page count and the size of every page without rendering anything. A third tool, get_license_status, reports the active licensing mode.

How do I run the server for Claude, Cursor or Copilot?

The Docker image is self-contained:

docker run --rm -i -v $(pwd)/documents:/data \
  ghcr.io/groupdocs-viewer/viewer-net-mcp:latest

With the .NET 10 SDK installed, dnx GroupDocs.Viewer.Mcp --yes runs the same server without Docker. Register it in Claude Desktop, Claude Code, Cursor, VS Code with GitHub Copilot, Windsurf and other clients; the hub page lists the exact config for each. The client starts the server as a child process and talks to it over standard input and output, so there is no inbound port and no endpoint to secure.

Example session

This is an illustration built from the documented behavior of the two tools, not a captured screenshot.

You:   How many pages does report.pdf have? Then show me the last one.

Agent: [calls get_view_info on report.pdf]
       report.pdf is a PDF with 2 pages, both 612 x 792 points (US Letter).
       [calls render_page with page = 2]
       Page 2 of 'report.pdf' Saved to: <output folder>/report_page2.png
       Here is what page 2 shows: ...

The image appears inline in clients that display it. The text reply is the agent’s description of what it sees on the page.

What does the agent decide, and what does the engine do?

The agent decides which page answers the question and what to ask the model about it. The engine does the part that has to be exact: opening the file, paginating it the way the document paginates, and producing the PNG. The model never has to guess at page layout, because it receives the rendered page.

Honest limits

  • Evaluation mode. Without a license, rendered pages carry an evaluation watermark and a server process opens at most 15 documents. The watermark can cover the very detail you ask about. Ask the agent “Is the viewer server licensed?” so it calls get_license_status before you rely on what it reads.
  • PNG only. The output is a PNG image; there is no format option.
  • No page-number check. render_page does not reject a page number past the last page. Check the count with get_view_info first.
  • The model must accept images. With a text-only model the agent still gets the saved file path, but it cannot tell you what is on the page.
  • Images are large. Each page is an inline PNG, often a few hundred kilobytes. Render the pages the task needs, not the whole file.
  • One thing leaves the machine. The rendered image goes to the agent and from there to whatever model it uses. A locally hosted model keeps it all on the machine.

FAQ

Can Claude see my PDF page and its layout? Yes, if the PDF page is rendered to an image and the Claude model in use accepts image input. render_page produces that image locally and returns it to the client.

Do I have to upload the document anywhere to render it? The server reads the file from the folder you mount and writes the image next to it; the document is not uploaded by the server. The page image does go to your agent’s model.

Does it work for Excel and PowerPoint, not only PDF? Yes. PowerPoint renders one image per slide, and Excel sheets are split into pages, with the count visible through get_view_info.

Go deeper