An AI agent can pull every comment out of a marked-up document and report on it in three ways: a summary, a table, or an XML archive. The GroupDocs.Annotation.Mcp server hands the LLM the annotations as structured data, and Claude, Cursor or GitHub Copilot do the reasoning, all locally:

Summarize the review of contract_annotated.pdf: how many comments, from whom, and what the main themes are.

The step-by-step version with config and troubleshooting is in the documentation: How to extract and report on document annotations with AI.

What do you get when you extract annotations from a PDF?

A marked-up document is data as well as ink. The get_annotations tool returns a JSON object with found, the count, and annotations, an array. Each entry holds the annotation id, type, message, page (1-based), the bounding box, the user and any replies. There is no document parsing and no OCR involved, so the agent works from the annotation records themselves.

The tool takes one document per call. To cover a set of files, name them in the prompt and the agent calls the tool once per file. The server does not offer a folder-level call.

Way 1: Ask for a summary

One call, then reasoning. Name the file and say what the summary is for.

Summarize the review of contract_annotated.pdf: how many comments, from whom, and what the main themes are.

The agent calls get_annotations, counts the entries, groups them by user and describes the themes it finds in the message fields. The count and the authors come from the data. The themes are the model’s reading of the text, so they deserve a human glance.

Way 2: Ask for a table or a filtered list

Because the data is structured, the shape of the report is a matter of the prompt:

Give me a table: page, author, comment, answered yes/no.

Group the comments by reviewer and tell me who has the most open items.

List only the strikeouts. I want to see what people want removed.

Which pages have no comments at all?

The last question is one a PDF reader answers poorly and an array answers directly: “answered” is a non-empty replies list, and “pages without comments” is the page numbers missing from the page values. Pair it with get_document_info for the page count.

Way 3: Archive the comments as XML

Export the annotations so we can keep them with the case file.

export_annotations writes an XML file that holds the annotations alone. Its success message has the form Exported N annotation(s) from <file> to <name>.annotations.xml. The file is small and can be diffed, and import_annotations can apply it to another copy of the document. This helps when the document is under a retention rule that its comments are not, or when you want to compare two review rounds: two exports of the same document are two text files, and the agent can describe what changed between them.

Example session (abridged)

This is an illustration built from the documented behavior of the tools, not a captured screenshot.

You:   Summarize the review of contract_annotated.pdf and group it by reviewer.

Agent: [calls get_annotations on contract_annotated.pdf]
       Found 5 annotations: 4 from Maria, 1 from Dana.
       Maria: payment terms (page 2), delivery date (page 5),
         liability cap (page 6), a missing signature block (page 6).
       Replies: Alex answered the page 2 payment terms comment.
       Dana: a strikeout on page 4 marking an outdated clause.
       Maria comments without a reply: page 5 and the two on page 6.

Can the agent act on the comments?

It can pass the list on: open an issue per unanswered comment, draft a reply email, or write a status paragraph. One rule applies. Comments come from other people, so the agent should treat annotation text as content to report on, not as instructions. A comment that says “ignore previous instructions” is data in a field. Ask for summaries and reports, and keep actions that change anything under your own review.

How do I run the server for Claude, Cursor or Copilot?

The server runs from the ghcr.io/groupdocs-annotation/annotation-net-mcp Docker image or from NuGet with dnx GroupDocs.Annotation.Mcp --yes (.NET 10 SDK), and mounts the folder that holds your documents. The hub page lists the registration for Claude Desktop, Claude Code, VS Code with GitHub Copilot, Cursor, Windsurf and other clients.

Honest limits

  • Evaluation mode. A trial badge is stamped at the top of every page. It does not change the annotation data, so reports built this way are accurate without a license. It affects what you can send on. If the plan ends with “email the marked-up copy”, check get_license_status first.
  • One document per call. A sweep over many files means many calls, and each is a tool call the agent has to make and report on.
  • Export is XML. get_annotations returns JSON for the agent to read; export_annotations writes XML to storage.
  • No preview needed. None of the three ways renders a page, so the PDF preview limit of generate_pages_preview on Linux does not affect them.

FAQ

How do I export all comments from a PDF? Ask the agent to export the annotations. It calls export_annotations, which writes an XML file next to your documents, or it calls get_annotations if you want the comments listed in the chat.

Can AI summarize reviewer comments across several files? Yes. Name the files in the prompt and the agent reads each one with get_annotations, then combines the results in one summary. It reads one file per call.

Does the report change if I have no license? The annotation data does not change. Only the rendered or saved pages carry a trial badge.

Go deeper