A disclosure redaction policy can be enforced by writing it as a prompt that an AI agent runs, in a fixed order, against each file, locally on your machine. With GroupDocs.Redaction.Mcp connected to Claude Code, Cursor or GitHub Copilot, the whole policy fits in one instruction:

Take case-file.docx and prepare it for disclosure: redact all email addresses and the name “Jane Smith”, delete all annotations, clear the document properties. Apply each step to the file produced by the previous one, and name the file each step produced.

The step-by-step version with config and troubleshooting is in the documentation: How to prepare documents for disclosure with an AI agent.

Why does a manual disclosure checklist fail?

Disclosure work is not hard, it is sequential, and it is repetitive. A records officer preparing a public-records release, or a legal team preparing documents for discovery, repeats the same passes on every file: names and identifiers in the text, images, review comments, document properties. Skipping one pass is easy at the end of a long day, and a partial redaction is the worst outcome because the file looks finished. A policy that lives in a prompt is applied the same way each time, and the transcript records what ran.

This is suitability, not certification. The server removes what you tell it to remove, in the file you give it. Whether the result satisfies a FOI regime, a discovery protocol or an audit remains your decision.

Which steps does the agent run, and in what order?

The policy is five steps. Each of the first four writes a new file.

  1. Run redact_text once for each pattern: names, emails, identifiers, account numbers.
  2. Run redact_image_area for signatures, photographs, stamps and scanned regions. Use it on image files such as PNG or JPEG scans.
  3. Run redact_annotations, usually with deleteAll, because internal review notes are rarely disclosable.
  4. Run erase_metadata for author, company and dates.
  5. Verify the final file, which is the subject of a separate post on verification.

What does the agent decide, and what does the engine do?

Decision Who makes it
Which patterns express the policy You, with the agent proposing regular expressions
Which tool handles each location The agent
Whether a text match is removed The engine, exactly as the pattern says
Which file is passed to the next step The agent, which is why you ask for file names
Whether the policy was complete You

There is no classifier. The engine applies the patterns you approve, and that is what makes the run auditable.

How do you keep every step in the final file?

Each tool writes a new _redacted-style file rather than editing in place. An agent that passes the original file into a later step produces a document that has only that step applied and the earlier removals missing. The instruction “apply each step to the file produced by the previous one” is the control, and the file names in the transcript are the audit trail.

Ask for patterns one at a time instead of one opaque instruction, with the file name each call produced. The redaction tools report the saved file, not a match count; any number the agent gives is its own reading of the text. Several named files make a record, and one reply of “done” does not.

Example session (abridged)

You:    Before you start: what is the license status of the redaction server?
Agent:  -> get_license_status()   mode: licensed
You:    Take case-file.docx and prepare it for disclosure: redact all email addresses and the
        name "Jane Smith", delete all annotations, clear the document properties.
        Apply each step to the file produced by the previous one, and name each file.
Agent:  1. redact_text        on case-file.docx       -> saved file
        2. redact_text        on the file from 1      -> saved file
        3. redact_annotations on the file from 2      -> saved file
        4. erase_metadata     on the file from 3      -> saved file
        Final file: the output of step 4, named in the transcript.

How do you run the policy across many files?

Repeat the same instruction per file, or name a folder and ask the agent to work through it one document at a time, naming the output file for each. Keep the policy text identical between runs; the sameness is the point. Because evaluation mode allows one document per process, a batch like this needs a license file or metered keys.

How do you keep the originals from being released?

The original sits next to its redacted copy with a near-identical name, and the wrong one is easy to attach to an email. Move finished originals out of the working folder, or set a separate GROUPDOCS_MCP_OUTPUT_PATH so redacted copies land in their own folder. A release of the wrong file defeats every earlier step.

What are the limits of evaluation mode?

Evaluation mode makes a disclosure workflow dangerous, and a full disclosure run is not possible without a license. Without one the server allows one document per process, one redaction, capped at 4 replacements, plus trial badges. The result is a document that looks processed and still contains the data, and nothing in the responses says so. Make get_license_status the first step, as in the session above.

  • PDF on Linux. On Linux, including the Docker image, redact_image_area and erase_metadata currently fail on PDF files (the PDF engine’s image handling depends on System.Drawing, which .NET supports only on Windows); both work on Word documents, and redact_text works on PDF. Run PDF area redaction and metadata erasure on Windows with dnx.

FAQ

How do I prepare documents for disclosure with an AI agent, for example a FOIA request? It can run the redaction passes: text by pattern, image areas, annotations and metadata. You still define the patterns, review the result, and decide what is released.

Can I use this for e-discovery review? The same passes apply to documents prepared for legal disclosure, and the server runs locally. It does not decide relevance or privilege, and it is not a review platform.

Where do the redacted copies go? To the output folder you configure. Set GROUPDOCS_MCP_OUTPUT_PATH so they stay apart from the originals.

Go deeper