An LLM can only retrieve what the ingestion step preserved, so the first job of a RAG pipeline is turning Word, PDF, Excel and e-book files into clean Markdown. With GroupDocs.Markdown.Mcp an AI agent does that locally, from Claude Desktop, Claude Code, Cursor or any other MCP client, and the files never leave your machine:

Convert every PDF in this folder to Markdown, images as files, with front matter.

The step-by-step version with config and troubleshooting is in the documentation: How to convert documents to Markdown for RAG with an AI agent.

Why do most RAG pipelines start with the wrong conversion step?

The usual habit is to treat conversion as plumbing: send the files to whatever parser is nearest, accept whatever text comes back, and spend the effort on embeddings. That reverses the order of importance. A chunker splits on headings, a table that arrives as a stream of cell values answers no question, and a corpus uploaded to a third-party parser is a disclosure you did not plan for. When retrieval is poor, the cause can be in the first file the pipeline touched.

This post argues for three corrections: keep the step on your own machine, choose output options on purpose, and verify the result instead of trusting it.

Is a cloud parser the right place for document ingestion?

For an internal corpus, usually not. Ingestion is the one stage that sees every document you own: policies, price lists, post-mortems. A parsing service that receives them all has seen the whole archive, not one file.

GroupDocs.Markdown.Mcp uses the MCP stdio transport. Your AI client starts the server as a child process, and the data path is agent, local server, local filesystem. The server has no inbound ports, no external endpoints and no telemetry. Only what the agent writes in its replies travels to the model provider. Metered licensing is the one case that needs outbound traffic: it reports usage, never document content, and a license file works offline.

Start it with Docker:

docker run --rm -i -v $(pwd)/documents:/data \
  ghcr.io/groupdocs-markdown/markdown-net-mcp:latest

or, with the .NET 10 SDK, dnx GroupDocs.Markdown.Mcp --yes. Point GROUPDOCS_MCP_STORAGE_PATH at the folder that holds the corpus.

Which GroupDocs server converts to Markdown: Markdown or Conversion?

Both can produce Markdown, and they answer different questions.

You need Use
Markdown that keeps structure, with control over images, flavor, page selection and front matter GroupDocs.Markdown.Mcp (this server)
Markdown as one of many targets, or DOCX to PDF to XLSX in the same workflow GroupDocs.Conversion.Mcp, which covers 100+ formats in both directions

If Markdown is the destination and its quality matters, choose the dedicated engine. If you convert between many formats and Markdown is only one of them, choose the general converter; its PDF route is covered in 3 ways to turn PDFs into LLM-ready Markdown (and why local wins).

What makes Markdown the best format for RAG ingestion?

Headings survive, tables stay tables, lists stay lists. Those are the units a chunker and an embedder work with. The server converts DOCX, DOC, RTF, ODT, XLSX, XLS, ODS, CSV, TSV, EPUB, MOBI, TXT, CHM and 20+ more formats through one tool, convert_to_markdown. Word-family sources convert most cleanly because their structure is explicit in the file. Spreadsheet sheets become Markdown tables.

Three options matter for an index:

Option Setting for ingestion Reason
images "file", or "skip" for prose only The default base64 puts a multi-megabyte data URI into the middle of a chunk
frontMatter true YAML with title, author, format and page count becomes chunk metadata, so an answer can cite a document title and not a file name
flavor "github" (default) or "commonmark" Match the parser that reads the result

Example session (abridged)

You:   Before converting anything: what is the license status of the markdown server?
Agent: get_license_status -> mode: licensed, engine GroupDocs.Markdown 26.3, server 26.9.0.
You:   Convert every PDF in this folder to Markdown, images as files, with front matter.
Agent: convert_to_markdown (images: "file", frontMatter: true) for each PDF.
       Saved one .md per document, with extracted images alongside.
You:   For each converted file, give me the source page count and the Markdown length.
Agent: get_document_info for each source, then the length of each saved .md.
       One 300-page file produced only 4 KB of Markdown: likely a scan.

How do I check the corpus before I index it?

Treat the run as untrusted until you have looked at it. One prompt catches most failures:

For each converted file, tell me the source page count and the resulting Markdown length.

get_document_info returns fileName, fileFormat, pageCount, title, author and isEncrypted without converting, so the agent can compare the page count to what it wrote. A 300-page manual that became 4 KB of Markdown is a scan or a truncated run.

Scanned PDFs are a hard limit. They have no text layer, and this server has no OCR step; handle those files separately.

Honest limits: the evaluation cap

Without a license the server runs in evaluation mode, where only the first 3 pages are processed, with an evaluation watermark on each. The output is valid Markdown and nothing reports that the rest was skipped, so an index built this way looks complete and is almost empty. Make get_license_status the first call of every ingestion run.

FAQ

What is the best format for RAG ingestion?

Markdown is a strong choice because chunkers and embedders handle headings, lists and tables well. Use images: "file" and frontMatter: true so chunks carry no base64 and keep their source metadata.

Can I convert Word to Markdown for an LLM without uploading the documents?

Yes. GroupDocs.Markdown.Mcp runs as a local process over stdio, and the agent reads the files from the folder you configured.

Does it turn a PDF into Markdown for a vector database?

It converts PDFs that have a text layer, with headings, lists and tables recovered where the source allows. Scanned PDFs have no text to convert, and there is no OCR step.

Go deeper