You can remove metadata before sharing with a single prompt to an AI agent that runs GroupDocs.Metadata.Mcp locally, but a stripped file is not automatically a safe file. The prompt:

Strip all metadata from the files in my documents folder and tell me what you removed.

The step-by-step version with config and troubleshooting is in the documentation: How to strip metadata from documents before sharing them.

What is the common wrong approach?

The common approach is “strip it and send it”: run one removal, assume the result is clean, attach the file. This fails in three ways, and none of them is a bug. Each is a boundary of what metadata removal means.

Why is the original still a risk?

remove_metadata removes metadata from a document and saves a cleaned copy to storage. It does not rewrite the original. The original stays on disk with everything in it, so attaching the wrong file defeats the exercise. This is a familiar incident: a tender document that arrives with a competitor’s name in the Company field because the uncleaned file was the one in the email.

The fix is procedural. Make the agent name the file it produced, and send only that file.

Why is “no metadata” not the same as “no hidden content”?

Tracked changes, comments, embedded objects and earlier revisions live in the document body, not in the property set. A metadata tool cannot remove what is not metadata. For comments and annotations the docs point to the GroupDocs.Annotation.Mcp server, and for content-level redaction to GroupDocs.Redaction. Which server to use for the metadata step: the Redaction server’s erase_metadata clears named document-property fields (author, title, company, keywords, dates and similar), while this server’s remove_metadata strips by category (gps, author, comments, company, dates, software, copyright, keywords, personal) and covers the EXIF, XMP and IPTC packages. For hidden text and annotation concerns, see 3 ways to redact sensitive data with AI agents.

Why can an unlicensed run not be trusted?

In evaluation mode only the first 5 document properties are readable and writable, and XMP and EXIF are partly unavailable. A “strip everything” pass under evaluation cannot honestly claim that everything was removed, and the verification read is partial too: you would be confirming a truncated list. Call get_license_status first. Apply a license file or metered keys, or use the free 30-day temporary license, before a real pre-share run.

What is the right way: read, strip, verify?

Use a loop in which every step is a tool call and the last one is a read of the copy:

  1. Capture. read_metadata on the original records what was there.
  2. Strip. remove_metadata produces the cleaned copy. With no categories it removes all removable metadata, which means document properties, EXIF, XMP and IPTC packages and similar; information that is structurally part of the document is not metadata and stays. To keep the title and dates, pass categories instead: gps, author, comments, company, dates, software, copyright, keywords or personal (a best-effort bundle covering people, company and location).
  3. Verify. read_metadata on the cleaned file shows what remains.

Step 3 is the one people skip, and it turns “I think it is clean” into a before and after. A reusable prompt:

Before I send anything from this folder: read the metadata, strip it, verify the copy is clean, and give me a one-line report per file.

That is three tool calls per file and a table at the end, which is routine work for an agent and work that a person forgets under deadline.

When should the strip run: once, or every time?

Treat it as a checkpoint, not a cleanup project. The incidents this prevents are rarely dramatic: a press release with the drafter’s name in the author field, a contract with a former employer in the company field, a site photo with coordinates attached. Each is an ordinary file that left the folder without anyone asking what it carried. An agent is a good last checkpoint because the question costs nothing to ask. Put the standing instruction in your project rules or in the prompt you reuse, so the read, strip and verify sequence runs on the outgoing folder every time and not only when someone remembers.

The checkpoint only works if the license state is known. In a licensed or metered run the verification read sees the full property set, and the one-line report per file is a statement you can rely on. In an evaluation run, treat the report as a rehearsal of the workflow and nothing more.

Example session (abridged)

You:    Remove the author and company from tender-draft.pdf, keep title and dates, then verify.
Agent:  [calls read_metadata on tender-draft.pdf]            Author, Company, Title, dates found
Agent:  [calls remove_metadata: categories = author, company]  2 properties removed
Agent:  [calls read_metadata on the cleaned copy]            Title and dates remain; Author and Company absent
        Send the cleaned copy named in the remove step, not tender-draft.pdf.

This is an illustration of the flow, not a captured screenshot. remove_metadata returns a saved-path message with the number of properties removed.

Frequently asked questions

Can I remove the author from a PDF before sending it? Yes. Ask the agent to remove the author category, and it will call remove_metadata with that category and save a cleaned copy. Check the copy with a second read.

Should I remove metadata before uploading a document to an LLM? If the document will leave your machine, strip a copy first, because the properties travel with the file. With GroupDocs.Metadata.Mcp the stripping itself happens locally.

Does stripping also remove tracked changes and comments in the body? No. Those are content, not metadata. The comments category refers to comment properties in the metadata, and body-level comments and annotations belong to the Annotation server, and hidden text and other content-level redaction to the Redaction server.

Go deeper