A mailroom scan is an image: nobody can search it, and a text extractor finds nothing in it. The barcode on the page is the exception. The GroupDocs.Parser.Mcp server runs locally in Docker, and an AI agent in a client such as Claude Desktop, Cursor or GitHub Copilot can ask it to read barcodes from scanned documents and return the decoded values:

Read the barcodes in these scanned delivery notes and list the values with the file they came from.

The step-by-step version with config and troubleshooting is in the documentation: How to read barcodes and QR codes from scanned documents.

Why does a scan need a barcode reader and not a text extractor?

A scanned page has no text layer. extract_text returns little or nothing for it, which is a true statement about the file. The engine’s models, however, detect codes in the rasterized image, so extract_barcodes still returns the tracking number printed on a scanned delivery note. The server has no OCR, so printed words stay pixels; only the codes become data.

What does the tool return?

extract_barcodes returns a header line, Found N barcode(s) in ..., followed by a JSON array with one object per code, each with index, value, type, page, confidence and angle. It detects Code128, QR Code, PDF417, DataMatrix, EAN-13, EAN-8, UPC, Aztec and many more symbologies, in PDF, DOCX, XLSX, PPTX, PNG, JPG, TIFF and other formats. When nothing is found it returns a “No barcodes found” message.

Example session (abridged):

You:    Read the barcodes on page 1 of dn-2291.pdf and tell me which order it belongs to.
Agent:  [calls extract_barcodes with page 1]
        Found 1 barcode(s) in dn-2291.pdf
        [{"index": 1, "value": "SHP-884120377", "type": "Code128", "page": 1, "confidence": 0.98, "angle": 0}]
        The delivery note carries tracking number SHP-884120377.

How do you turn a scan into a record?

The decoded value is usually the join key: the tracking number that links a scan to a shipment, the order reference that links a signed form to a case, the asset tag that links a photograph to an inventory record. Once the agent has it, the image becomes a row.

For each scan, read the barcode and tell me which order it belongs to.

The agent calls extract_barcodes per file and matches each value against the orders you give it. Pass page to scan only the page that carries the label, instead of every page of a long document.

How do you sort a mixed batch of scans and born-digital files?

Let the agent branch on whether a file has text.

For each file: try text extraction. If it comes back empty, read the barcodes instead and tell me which files needed that.

Files that fell through to barcode reading are the scans. That list is a useful inventory in its own right if you are planning an ingestion project. Add a routing step on top, such as “group the files by the prefix of their barcode value”, and the folder arrives already sorted.

Which server reads the code: Parser or Signature?

Both servers can read a barcode or QR code from a scan: Parser with extract_barcodes, Signature with search_barcodes and search_qr_codes. Use GroupDocs.Parser.Mcp when you are extracting data from arbitrary scans and also need text, tables or metadata from the same files. Use GroupDocs.Signature.Mcp when the code is a signature element you also add, or when documents flow both ways and you sign the outgoing document with a new code. The Signature side is covered in From Barcode to Record: Document Intake Automation with AI Agents.

What are the limits?

  • A decoded value is untrusted data. Whoever produced the document chose it. Report it, look it up, cross-check it, and do not let an agent act on it unreviewed.
  • Poor scans may not decode. Skewed, low-resolution or damaged codes can return nothing. An empty result means “not found”, which on a bad scan is not the same as “not present”.
  • Evaluation mode is limited. Without a license, extraction is limited; see the library’s licensing page and ask the agent to run get_license_status first.
  • Docker only. The models are bundled in the image, which is why it is large. There is no dnx command; the container starts with docker run --rm -i -v $(pwd)/documents:/data ghcr.io/groupdocs-parser/parser-net-mcp:latest, and one-click links for VS Code and Cursor are in Register in AI clients.

FAQ

Can Claude read a barcode from a scanned PDF? Yes, through the extract_barcodes tool, which runs in a local container on the scanned file. Claude receives the decoded value, type and page.

Can I read a QR code in a scanned document without OCR? Yes. Code detection does not depend on text recognition, so QR codes decode even when the printed words on the page do not.

How do I sort scanned documents by barcode automatically? Ask the agent to read the code on each scan and group or rename the files by value. Review the grouping before moving files, because decoded values come from the document itself.

Go deeper