A redaction is complete only when you have checked the file you are about to release, and an AI agent can run that check locally in a few calls. With GroupDocs.Redaction.Mcp connected to Claude Code, Cursor or GitHub Copilot, start by reading the redacted copy back as text and searching it:
Extract the text of case-file_redacted.pdf and tell me whether any email address appears in it.
Expect none. The step-by-step version with config and troubleshooting is in the documentation: How to verify that a redaction is complete.
Why is “the tool ran” not evidence?
Because a redaction tool that finishes without an error says nothing about the pattern, the file, the license state or the other places data hides. A redaction you have not verified is a claim, not a control. The usual way this goes wrong is not exotic. The failure classes are few:
- The data was covered, not removed. Opaque blocks drawn over text leave the text in the file, and readers can copy it out. The Wikipedia article on redaction describes a 2005 US military report on the death of Nicola Calipari, published as a PDF with blocked-out portions that could be retrieved by copying and pasting them into a word processor (see also BBC News, 2 May 2005).
- The pattern missed a variant: a different separator, a line break, a different case.
- A step was applied to the wrong file, so a later step undid an earlier one.
- Comments or document properties still named the person.
- An unlicensed run silently stopped at its replacement cap.
What is the wrong way to verify?
Asking the agent “did you get everything?” and accepting “yes”. The agent has no new information; it is summarising its own calls. A second wrong way is to verify the file you think you produced. The chaining rule means it is entirely possible to verify a document that skipped a step.
What is the right way? Four checks
Run them on the exact file you plan to release.
1. Read the redacted copy back as text. The prompt at the top does this with a tool that returns text, such as extract_text in GroupDocs.Parser.Mcp (docs), which is delivered as a Docker image. Any hit means the pattern missed variants, or the run hit the evaluation cap. Do not use redact_text for this check: it reports a saved file, not a count, and writes another copy, so a clean-looking reply proves nothing.
2. Check the other three places. Annotations and properties do not show up in a text search. The GroupDocs.Annotation.Mcp server lists annotations with get_annotations, one entry per comment with its page and message.
Now check the redacted copy for annotations and read its document properties.
For a deeper properties check across every package, the GroupDocs.Metadata.Mcp server’s read_metadata tool returns all metadata properties as JSON, and it also reads EXIF, XMP and IPTC. This is the cross-server loop: the Redaction server removes, a second server reads back.
3. Look at the pages. Area redactions cannot be verified by text search. Open the document, or have the agent render every page with the GroupDocs.Viewer.Mcp render_page tool, and confirm each box covers what it was meant to cover, completely, on the right page.
4. Confirm the license state of the run that produced it.
What is the license status of the redaction server?
If the answer is evaluation, the redaction was capped at four replacements. Verification will usually show it, but a document with exactly four matches would pass and still be wrong. License first, verify second, every time. get_license_status answers the first question, and get_document_info returns page count and sizes when an area redaction is involved.
Example session (abridged)
You: What is the license status of the redaction server?
Agent: -> get_license_status() mode: licensed
You: Extract the text of case-file_redacted.pdf and tell me whether any email address appears.
Agent: -> extract_text (Parser server) No email addresses in the extracted text.
You: Now check it for annotations and read its document properties.
Agent: -> get_annotations (Annotation server) no annotations
-> read_metadata (Metadata server) Author, Company: empty
You: Summarize what you redacted, in which file, and what the verification found.
How do you keep the evidence?
The verification transcript, meaning patterns used, files produced and what the read-back found, is the record that the redaction was performed and checked. Ask the agent for it as a summary and keep it with the case file. This is suitability for an audit trail, not a certification of the outcome.
What can verification not tell you?
It cannot tell you whether your patterns were the right ones. Everything above shows that what you asked to be removed is gone. Deciding what should be removed is your judgement, and no tool takes that off your hands.
What are the limits of evaluation mode?
One document per process, one redaction, capped at 4 replacements, plus trial badges. The cap is a demo limit; an unlicensed run looks redacted and can still contain the data, so check get_license_status before you trust any verification.
- PDF on Linux. On Linux, including the Docker image,
redact_image_areaanderase_metadatacurrently fail on PDF files (the PDF engine’s image handling depends onSystem.Drawing, which .NET supports only on Windows); both work on Word documents, andredact_textworks on PDF. Run PDF area redaction and metadata erasure on Windows withdnx.
FAQ
Is my redacted PDF really redacted? Only if a text read-back of the redacted copy finds none of the removed values, annotations and properties are clear, and you have looked at area redactions. Checking all four takes a few calls.
Can an AI agent check a redacted PDF for leaks? It can read the text back, read properties through a metadata server and render pages for you to view. It cannot judge whether you chose the right patterns.
Does metadata survive a text redaction? Yes, unless you also run erase_metadata. Text redaction changes the text; author, company and date properties are separate.
Go deeper
- Documentation, canonical how-to: How to verify that a redaction is complete
- Documentation hub: GroupDocs.Redaction MCP Server
- Start here: 3 ways to redact sensitive data with AI agents and MCP
- Related: Enforce disclosure redaction policy with automated AI workflows using MCP
- Related: Redacting scanned documents when there is no text to find
- On-premise and security model: 3 architectures for AI document processing, and the one that keeps files inside your network
- Questions: GroupDocs Redaction forum
- Source: GroupDocs.Redaction.Mcp on GitHub