An AI agent can read a chart in a PDF, or a scanned page with no text layer, if the page is rendered as an image first. The GroupDocs.Viewer.Mcp server does that rendering locally on your machine and hands the PNG to Claude, Cursor or GitHub Copilot, and a vision-capable model reads the image:
On page 7 of the annual report there is a revenue chart. What were the figures for each quarter?
The step-by-step version with config and troubleshooting is in the documentation: How to have an AI agent read charts, diagrams, and scanned pages.
The situation: the answer exists only as a picture
You open an annual report and the number you need is a bar height. You open a contract that came from a scanner and the PDF has no selectable text. You open a slide where the decision is a diagram arrow. An agent that works on extracted text finds nothing in any of these three cases, because there are no characters to extract. The information is real, and it is drawn.
What can the agent read from a rendered page?
- Charts and graphs. Values, trends, labels and legends.
- Scanned pages. A PDF made by a scanner is a picture of a page. The image is the only way in.
- Stamps, signatures and handwriting. Whether a page is signed, and what a stamp says.
- Layout questions. “Is the logo in the header?” or “Does this slide overflow?”
Each of these is a question about what the page looks like. A rendered image answers them directly.
Step 1: Find the page without rendering the document
A long report should not be rendered in full to find one chart. Ask the agent to locate the page cheaply:
How many pages does the report have? Render the contents page and tell me where the revenue chart is.
The agent calls get_view_info to learn the page count, then render_page for the contents page. That is two small calls rather than fifty renders. Each rendered page returns as an inline PNG, often a few hundred kilobytes, so the saving is real.
Step 2: Render the page and read it
With the page number known, the agent calls render_page for that page and passes the image to the model. The image is also saved as <name>_page<N>.png, so you can open the same picture yourself and compare it with the answer.
Step 3: Ask how sure the model is
Read the revenue chart on page 7. Say which values you are certain of and which you had to estimate from the bar heights.
A model reading an image can misread a small number or a faint scan. For figures that go into a decision, have the agent state its confidence, then check the source.
Example session
An illustration based on the documented behavior, not a captured screenshot. The bracketed text stands for what a model would read from the image.
You: On page 7 of annual-report.pdf there is a revenue chart. What were
the figures for each quarter? Say how sure you are.
Agent: [calls get_view_info on annual-report.pdf, then render_page with page = 7]
Page 7 of 'annual-report.pdf' Saved to: <output folder>/annual-report_page7.png
The chart shows [quarterly figures read from the image].
Values printed on the chart: [list]. Values estimated from bar height: [list].
Why not OCR?
Rendering is not OCR. OCR turns an image into text; this path gives the model the image and lets it interpret what is there, including the parts that are not text. The server does not produce verified text from a scan. If you need a searchable text layer, that is a separate step with a separate tool.
Honest limits
- A reading is an interpretation, not extraction. The model may misread a small number, a faint scan or a crowded legend. Rendering does not turn a scan into verified text.
- The model must accept image input. With a text-only model the agent can pass on the saved PNG path, but it cannot tell you what the chart says.
- Evaluation watermark. Without a license, every rendered page carries an evaluation watermark that sits on top of the page and can hide the detail you asked about. A server process also opens at most 15 documents. Ask “Is the viewer server licensed?” so the agent calls
get_license_statusbefore you rely on what it reads. - One page per call, PNG only.
- The image leaves the server. It goes to your agent and its model. A locally hosted model keeps everything on the machine; a hosted model receives the image as part of the conversation.
FAQ
Can Claude read a chart in a PDF? Claude can read a chart if the page is rendered to an image and the model in use accepts images. GroupDocs.Viewer.Mcp renders the page locally with render_page and returns the PNG to Claude.
My scanned PDF has no text, how can AI read it? Render the page to an image and give the image to a vision-capable model. The model reads the picture rather than a text layer, and its reading should be checked for important figures.
Is this the same as extracting data from a chart image? No. The model describes and estimates what it sees; it does not return exact underlying data. Use the source spreadsheet when exact values matter.
Go deeper
- Documentation, canonical how-to: How to have an AI agent read charts, diagrams, and scanned pages
- Documentation hub: GroupDocs.Viewer MCP Server
- Start here: Multimodal agents need pixels, not just text
- Related: 3 questions an AI agent should ask a document before rendering it
- Related: Automate screenshot-grade page images for reports with MCP
- On-premise and security model: 3 architectures for AI document processing, and the one that keeps files inside your network
- Questions: GroupDocs Viewer forum
- Source: GroupDocs.Viewer.Mcp on GitHub