To migrate documentation to Markdown, give an AI agent the archive folder and let it run convert_to_markdown once per file, in reviewed batches. With GroupDocs.Markdown.Mcp this happens locally in Claude Code, Cursor, GitHub Copilot or another MCP client, so an archive of internal procedures is never uploaded to a conversion service:
Convert everything in the archive folder to Markdown with front matter and images as files. Report the source format, page count, and output size for each.
The step-by-step version with config and troubleshooting is in the documentation: How to migrate legacy documentation to Markdown with AI.
Why is a legacy docs migration a job for an agent?
The conversion is mechanical and the surrounding work is repetitive: list the files, convert each, measure the result, find the failures, repeat. A person doing that for hundreds of files makes the same decision hundreds of times. An agent makes it consistently and reports each outcome, while the engine, not the model, produces the Markdown. The saved Markdown is not a paraphrase of your procedures.
This is a runbook. Each step has one prompt, and the table after it separates what the agent decides from what the engine does.
Runbook: how do I migrate a Word archive to Markdown?
-
Check the license before anything else.
Before converting anything: what is the license status of the markdown server?
get_license_statusreturnsmode,licensed,consumptionand the server and engine versions. Evaluation mode processes 3 pages per document, so an unlicensed run produces plausible stubs. -
Inventory the archive.
List every file in the archive folder with its format, page count and whether it is encrypted.
The Markdown server has no folder-listing tool, so the agent lists the folder with its client’s own file access or works from the file names you give it.
Then
get_document_inforeturnsfileFormat,pageCount,title,authorandisEncryptedfor each file without converting it. Encrypted files need thepasswordparameter. -
Run a pilot of about twenty files.
Convert the first twenty documents to Markdown with front matter and images as files. Report format, pages and output size for each.
-
Find what did not convert.
List the PDFs whose conversion produced less than 2 KB of Markdown for more than 5 pages.
Those are almost always scans. There is no OCR step in this server, so they need another tool or stay as PDFs indexed by metadata.
-
Fix the pattern, then run the rest in batches. Twenty files per round, with a report each time, keeps the run reviewable and keeps metered usage predictable.
-
Review the output as you would a pull request. The next sections list what to look at.
| Decision | Agent | Engine |
|---|---|---|
| Which files to convert, in what order and batch size | Yes | No |
Which options to pass (images, frontMatter, flavor) |
Yes, from your prompt | No |
| Turning headings, lists and tables into Markdown | No | Yes |
Saving the .md and image files |
No | Yes |
| Spotting a 300-page file that became 4 KB | Yes, from the report | No |
| Renaming, folder layout, extra front-matter keys | Yes | No |
Which formats migrate well?
- DOCX, DOC, RTF, ODT. Structure is explicit in the source, so headings and tables come across cleanly.
- EPUB and MOBI. Chapter structure maps onto Markdown headings.
- CHM. Old help files, often the only copy of a procedure.
- XLSX. Worksheets become tables, which suits parameter lists.
What should the review pass check?
Ask the agent to do the first sweep, then read a sample yourself:
For every converted file, list the images it references and tell me which image files are missing.
Show me the five outputs with the smallest size relative to their page count.
Images written with images: "file" sit next to the .md and are referenced by relative path, which is what a docs repository expects. A .md that points to pictures that were never written renders as broken images. Then open the messiest source and its output side by side before you apply the pattern to the rest. Commit the reviewed result with your normal git workflow; the server writes files and does not touch your repository.
Example session (abridged)
You: Convert the first twenty documents in /archive to Markdown with front matter
and images as files. Report format, pages and output size for each.
Agent: convert_to_markdown (frontMatter: true, images: "file") x 20
ops-manual-2014.docx DOCX 88 pages converted
install-guide.chm CHM 41 pages converted
supplier-forms-2009.pdf PDF 12 pages 1 KB of Markdown: probably a scan
Nineteen files look consistent; one scan needs separate handling.
Honest limits
Without a license, evaluation mode processes only the first 3 pages of each document, with an evaluation watermark on each. A migration run in that state yields hundreds of short files that look right, and the gap is found weeks later when someone searches for a procedure that is missing. Migration needs a license, so run get_license_status before the first batch. Scanned PDFs do not convert.
FAQ
Can I convert hundreds of Word documents to Markdown at once?
Yes, but not in one prompt. Convert in batches of about twenty and read the report after each, so a wrong option is fixed before it is applied to the whole archive.
How do I migrate documentation into git with an AI agent?
The server writes .md files and image files into the output folder. You review them and commit them with your usual git workflow.
Does the docs-as-code migration upload my documents?
No. GroupDocs.Markdown.Mcp is a local process, and the archive is read from disk. Only what the agent writes in the conversation reaches the model provider.
Go deeper
- Documentation, canonical how-to: How to migrate legacy documentation to Markdown with AI
- Documentation hub: GroupDocs.Markdown MCP Server
- Start here: Your RAG pipeline starts with Markdown — keep that step local
- Related: 3 ways to pull just the section you need into Markdown with MCP
- Related: From documents to a static site: Hugo, Docusaurus or MkDocs via an AI agent
- On-premise and security model: 3 architectures for AI document processing, and the one that keeps files inside your network
- Questions: GroupDocs Markdown forum
- Source: GroupDocs.Markdown.Mcp on GitHub