To migrate documentation to Markdown, give an AI agent the archive folder and let it run convert_to_markdown once per file, in reviewed batches. With GroupDocs.Markdown.Mcp this happens locally in Claude Code, Cursor, GitHub Copilot or another MCP client, so an archive of internal procedures is never uploaded to a conversion service:

Convert everything in the archive folder to Markdown with front matter and images as files. Report the source format, page count, and output size for each.

The step-by-step version with config and troubleshooting is in the documentation: How to migrate legacy documentation to Markdown with AI.

Why is a legacy docs migration a job for an agent?

The conversion is mechanical and the surrounding work is repetitive: list the files, convert each, measure the result, find the failures, repeat. A person doing that for hundreds of files makes the same decision hundreds of times. An agent makes it consistently and reports each outcome, while the engine, not the model, produces the Markdown. The saved Markdown is not a paraphrase of your procedures.

This is a runbook. Each step has one prompt, and the table after it separates what the agent decides from what the engine does.

Runbook: how do I migrate a Word archive to Markdown?

  1. Check the license before anything else.

    Before converting anything: what is the license status of the markdown server?

    get_license_status returns mode, licensed, consumption and the server and engine versions. Evaluation mode processes 3 pages per document, so an unlicensed run produces plausible stubs.

  2. Inventory the archive.

    List every file in the archive folder with its format, page count and whether it is encrypted.

    The Markdown server has no folder-listing tool, so the agent lists the folder with its client’s own file access or works from the file names you give it.

    Then get_document_info returns fileFormat, pageCount, title, author and isEncrypted for each file without converting it. Encrypted files need the password parameter.

  3. Run a pilot of about twenty files.

    Convert the first twenty documents to Markdown with front matter and images as files. Report format, pages and output size for each.

  4. Find what did not convert.

    List the PDFs whose conversion produced less than 2 KB of Markdown for more than 5 pages.

    Those are almost always scans. There is no OCR step in this server, so they need another tool or stay as PDFs indexed by metadata.

  5. Fix the pattern, then run the rest in batches. Twenty files per round, with a report each time, keeps the run reviewable and keeps metered usage predictable.

  6. Review the output as you would a pull request. The next sections list what to look at.

Decision Agent Engine
Which files to convert, in what order and batch size Yes No
Which options to pass (images, frontMatter, flavor) Yes, from your prompt No
Turning headings, lists and tables into Markdown No Yes
Saving the .md and image files No Yes
Spotting a 300-page file that became 4 KB Yes, from the report No
Renaming, folder layout, extra front-matter keys Yes No

Which formats migrate well?

  • DOCX, DOC, RTF, ODT. Structure is explicit in the source, so headings and tables come across cleanly.
  • EPUB and MOBI. Chapter structure maps onto Markdown headings.
  • CHM. Old help files, often the only copy of a procedure.
  • XLSX. Worksheets become tables, which suits parameter lists.

What should the review pass check?

Ask the agent to do the first sweep, then read a sample yourself:

For every converted file, list the images it references and tell me which image files are missing.

Show me the five outputs with the smallest size relative to their page count.

Images written with images: "file" sit next to the .md and are referenced by relative path, which is what a docs repository expects. A .md that points to pictures that were never written renders as broken images. Then open the messiest source and its output side by side before you apply the pattern to the rest. Commit the reviewed result with your normal git workflow; the server writes files and does not touch your repository.

Example session (abridged)

You:   Convert the first twenty documents in /archive to Markdown with front matter
       and images as files. Report format, pages and output size for each.
Agent: convert_to_markdown (frontMatter: true, images: "file") x 20
       ops-manual-2014.docx   DOCX   88 pages   converted
       install-guide.chm      CHM    41 pages   converted
       supplier-forms-2009.pdf PDF   12 pages   1 KB of Markdown: probably a scan
       Nineteen files look consistent; one scan needs separate handling.

Honest limits

Without a license, evaluation mode processes only the first 3 pages of each document, with an evaluation watermark on each. A migration run in that state yields hundreds of short files that look right, and the gap is found weeks later when someone searches for a procedure that is missing. Migration needs a license, so run get_license_status before the first batch. Scanned PDFs do not convert.

FAQ

Can I convert hundreds of Word documents to Markdown at once?

Yes, but not in one prompt. Convert in batches of about twenty and read the report after each, so a wrong option is fixed before it is applied to the whole archive.

How do I migrate documentation into git with an AI agent?

The server writes .md files and image files into the output folder. You review them and commit them with your usual git workflow.

Does the docs-as-code migration upload my documents?

No. GroupDocs.Markdown.Mcp is a local process, and the archive is read from disk. Only what the agent writes in the conversation reaches the model provider.

Go deeper