An AI agent can verify its own annotations by looking at the page it just marked up. With GroupDocs.Annotation.Mcp the LLM in Claude, Cursor or GitHub Copilot annotates a document, asks the server to render the page as a PNG with the annotation included, and judges whether the placement is right. The server does the work locally; the prompt is one sentence:

Highlight the payment clause on page 2, then show me the page.

The step-by-step version with config and troubleshooting is in the documentation: How to let an AI agent see the document page it annotated.

Note: generate_pages_preview currently fails for PDF files on Linux, including the Docker image; see Honest limits before you build on it.

Why does an agent that cannot see its output report success wrongly?

Most document tools are write-only. The agent calls them and says “done” because the call returned. For an annotation that statement can be false in a way no error message shows: the call succeeded, and the highlight sits on the wrong paragraph. The coordinates x and y are document coordinates, and an agent asked to highlight “the payment clause” has to guess where that clause is unless something tells it.

The difference is between “I added a highlight” and “the highlight is on the payment clause”. Only the second statement is useful, and only a look at the page can support it.

How does an agent render an annotated page image and check it?

The tools form a short loop.

  1. add_annotation writes the annotation and saves contract_annotated.pdf. The original file is not modified.
  2. generate_pages_preview renders pages of that file as PNG images, with annotations baked in, and returns them inline. Pass pages as 2, 1-3 or 1,3,5. Omit it and you get page 1.
  3. The agent looks at the image. If the highlight covers the wrong text, update_annotation changes its message or its position and size without starting over.

The preview returns one text block describing what was rendered, followed by one image block for each page. A call renders at most five pages.

Set GROUPDOCS_MCP_OUTPUT_PATH to a folder other than the storage folder before chaining edits on a produced file. With the default (outputs land in the storage folder) the second write fails with being used by another process; with a separate output folder add_reply, update_annotation and remove_annotations on the produced file all succeed. With the Docker image, pass -e GROUPDOCS_MCP_OUTPUT_PATH=/data/output. The session below assumes that setting, so the produced file sits under output/.

What does the agent need for this to work?

  • A client that shows image content. Claude Desktop does. A terminal client may only report that an image came back. The tool behaves the same either way; the images may simply not be displayed.
  • A vision-capable model. The images help only a model that accepts image input.
  • A small page range. Rendering is the most expensive call this server makes. Ask for the page you need, not a fifty-page file.

Which other uses does the preview have?

  • Locate before annotating. Render first, let the model read the page, then place the annotation with coordinates it derived instead of guessed.
  • Check a merge. After import_annotations onto a re-flowed draft, a preview shows whether the imported notes still sit where they should.
  • Show a reviewer the result in the chat without opening the file.
  • Spot the evaluation badge. In evaluation mode a trial badge is stamped on every page, so it is visible in the rendered image.

Example session (abridged)

This is an illustration built from the documented behavior of the tools, not a captured screenshot. It assumes a platform where the preview renderer works for PDF files, such as Windows with dnx.

You:   Highlight the payment clause on page 2 of contract.pdf, then show me the page.

Agent: [calls add_annotation: type=highlight, page=2, x=100, y=250,
        text="Payment clause"]
       Added highlight annotation to contract.pdf on page 2.
       [calls generate_pages_preview on output/contract_annotated.pdf, pages="2"]
       (image: page 2, rendered with the highlight)
       The highlight sits on the "Delivery" paragraph, one block above the
       payment clause. Moving it down.
       [calls get_annotations, then update_annotation: y=330]
       Updated annotation 1 in output/contract_annotated.pdf.
       (The agent passes the saved file name from each result to the next call.)

Where does the rendering run, and what leaves the machine?

The engine produces the PNG on your machine. The image then travels to your model provider like any other message content, so with a cloud-hosted model a preview sends a picture of that page. For sensitive documents, use a locally hosted model.

Honest limits

  • Preview on Linux. generate_pages_preview currently fails for PDF files on Linux, including the Docker image (TypeInitializationException for Gdip: System.Drawing.Common is not supported on non-Windows platforms); Word and other Office documents render, and on Windows with dnx every format renders. The other ten tools are unaffected.
  • Evaluation mode. A trial badge is stamped at the top of every page of the annotated file and of the previews. Check get_license_status before you rely on the output.
  • Approximate placement stays approximate. A preview lets the model notice a misplaced annotation; it does not turn guessed coordinates into exact ones.
  • A text-only model cannot check. It still receives the text block, but it cannot tell you what the page looks like.

FAQ

Can an AI check where a comment landed on the page? Yes, if the model accepts images. The agent renders the annotated page with generate_pages_preview and examines the picture, then adjusts the annotation with update_annotation if needed.

How do I preview an annotated PDF page for the agent? Ask for the page by number, for example “Render pages 1-3 of the annotated contract so I can check the comments”. The agent passes pages as 1-3 and the PNGs come back inline.

Why does the preview fail in my Docker container? For PDF files the tool fails on Linux, including the Docker image, with a Gdip type initialization error. Word and other Office documents render, and on Windows with dnx every format renders.

Go deeper