Introduction
A colleague sends two revisions of a contract and asks you to diff them. You drop both into your comparison service, the result comes back, and everything looks normal. What you did not see is that one of those documents carried a linked image pointing at a URL, and your server contacted that host the moment the file was opened. Nothing about the output tells you it happened.
This is not a defect - it is what loading a document faithfully means. An OOXML file can
reference an image that lives on a web server rather than inside the package, and both
Word and any library that loads the document properly resolve that reference.
GroupDocs.Comparison for .NET exposes two properties on LoadOptions that let you decide
whether it does: SkipExternalResources and WhitelistedResources.
Between them they give three configurations, and this article compares all three - the permissive default, blocking everything, and blocking everything except named references. By the end you will know which to pick for a given document source, and the two mistakes that make these settings look as though they do not work.
💡 Full working example: block-external-resources-on-document-load-dotnet - a runnable console project that serves the referenced images itself and logs every request, so you can watch each setting take effect.
Where External References Hide
Before choosing a setting, it is worth knowing what you are choosing about. A .docx
carries external references in two distinct places, and they are easy to miss because
neither is visible in the document text.
The first is a relationship in word/_rels/document.xml.rels carrying
TargetMode="External" and an absolute URL. The image shows up in the body as a drawing
that points at the relationship by ID, so the URL itself never appears near the content it
affects.
The second is an INCLUDEPICTURE field code in the document body, holding its URL inside
a field instruction. Word resolves it when the page renders; a comparison library resolves
it when the document loads.
Both mechanisms respect the two load options discussed below, which matters because a document can use either or both. A reference you spotted in the relationships file is not proof there is no second one in a field code.
Approach 1: The Default - References Resolved
SkipExternalResources defaults to false, so a document loaded without configuration
has its remote references resolved:
LoadOptions loadOptions = new LoadOptions
{
SkipExternalResources = false
};
using (Comparer comparer = new Comparer(sourcePath, loadOptions))
{
comparer.Add(targetPath, loadOptions);
comparer.Compare(outputPath);
}
This gives the highest fidelity: the compared documents contain everything they reference, exactly as Word would render them. For documents your own application or templates produced, where every reference URL points at infrastructure you operate, it is the right choice - and a missing linked image could make the comparison actively misleading.
The cost is that every reference is contacted, whoever put it there. There is also a timing cost that has nothing to do with trust: a reference URL that no longer resolves makes loading wait out the full connection attempt, on every single comparison.
Approach 2: Block Every External Resource
One property turns remote reference resolution off for that document:
LoadOptions loadOptions = new LoadOptions
{
SkipExternalResources = true
};
using (Comparer comparer = new Comparer(sourcePath, loadOptions))
{
comparer.Add(targetPath, loadOptions);
comparer.Compare(outputPath);
}
No request is issued. The referenced images are absent from the result, and - this is the part worth being clear about - nothing else changes. The setting governs what gets loaded, not how differences are found, so textual and structural changes between the two documents are detected exactly as before. The one thing you lose is the ability to detect a change inside a referenced image, which was never loaded.
This is the configuration to treat as your baseline for documents you did not create: user uploads in a web application, files received by email, anything compared on a build agent where an outbound request is rarely intended. It is all or nothing, though - a linked image you actually wanted is blocked along with the rest, and the result simply lacks it without announcing the fact.
Approach 3: Block Everything Except Named References
The third configuration is the one that rewards a close reading.
WhitelistedResources takes a List<string> and is consulted only when
SkipExternalResources is true:
LoadOptions loadOptions = new LoadOptions
{
SkipExternalResources = true,
WhitelistedResources = new List<string> { "includepicture-field.png" }
};
using (Comparer comparer = new Comparer(sourcePath, loadOptions))
{
comparer.Add(targetPath, loadOptions);
comparer.Compare(outputPath);
}
The entries are URL fragments, not file names. Each is matched against the reference
URL, and a match anywhere in it admits that resource. That is what makes the whitelist
portable: "includepicture-field.png" admits the image whatever scheme, host and path
precede it, so the same list works in development and production without rewriting.
The same property cuts the other way. A short or generic fragment - logo.png, or worse,
.png - can match references you never intended to allow. Pick a fragment specific enough
to identify the one resource you meant.
In the reference sample, this configuration fetches the whitelisted image and leaves a second referenced image, which no entry covers, blocked. The request log shows three requests where the permissive default produced five, and names only the whitelisted file.
Which Configuration Should You Use?
Match the setting to where the document came from. Documents your own application or
templates generated can keep the default, because every reference URL points at
infrastructure you already operate. Anything arriving from outside - user uploads, email
attachments, third-party files - warrants SkipExternalResources = true. Add a narrow
WhitelistedResources fragment only when one trusted reference genuinely has to resolve.
Comparing the Three
| Concern | Default | Skip all | Skip + whitelist |
|---|---|---|---|
| Properties to set | 0 | 1 | 2 |
| Outbound requests | all references | none | whitelisted only |
| Per-reference control | no | no | yes |
| Dead URL costs load time | yes | no | whitelisted only |
| Best for | documents you produced | documents from anywhere else | trusted templates among untrusted content |
The decision follows document provenance rather than performance. Documents your systems generated can keep the default. Documents from outside warrant blocking. Whitelist at the point where one specific reference genuinely has to resolve - a corporate template pulling its header image from an internal URL, say, among reports whose authors pasted in images from wherever they liked.
The Two Mistakes
Both of these produce the same symptom: you set the option, and it appears to do nothing.
A whitelist without the switch. WhitelistedResources is consulted only when
SkipExternalResources is true. Set on its own, it does nothing at all - there is no
blocking for it to make an exception to. If a whitelist looks ignored, check this first.
Options on the source only. This is the subtler one. Load options describe how one
document is loaded. The Comparer constructor takes the options for the source; every
Add() call takes the options for that target:
using (Comparer comparer = new Comparer(sourcePath, loadOptions))
{
comparer.Add(targetPath, loadOptions);
comparer.Compare(outputPath);
}
Pass them to the constructor and forget the Add() call, and the source is protected
while every target still fetches its references. The comparison succeeds, the result looks
plausible, and half your documents are still reaching out to the network. Where source and
target need different handling, pass separate LoadOptions instances - that is precisely
why the API takes them per document.
Verifying It Actually Worked
A blocked resource leaves almost no trace. The output document is missing an image, which looks much like a document that never had one. Reading the result file is therefore a poor way to confirm the setting took effect.
Watch the serving side instead. The reference sample takes this approach deliberately: it starts a small HTTP listener on a free loopback port, writes its demo documents pointing at that port, and logs every request it receives, printing the count per comparison. Five requests, then zero, then three. A network trace against your real document sources gets you the same confidence.
Conclusion
Three configurations, one decision rule: let documents you generated keep the default, set
SkipExternalResources = true for everything else, and whitelist a narrow URL fragment
only where a specific trusted reference still needs to resolve.
Then check the two things that silently undo the work - a whitelist without
SkipExternalResources = true, and options passed to the Comparer constructor but not to
every Add() call - and verify from the serving side rather than the output file.
Additional Resources
- Configuring external resource loading in GroupDocs.Comparison for .NET - the use-case guide, with the decision matrix and FAQ
- block-external-resources-on-document-load-dotnet - the complete runnable sample, including the loopback image host
- Load password-protected documents -
LoadOptions.Password, the sibling setting with the same per-document scope rule - Load custom fonts - resolving non-standard fonts at load time with
LoadOptions.FontDirectories - GroupDocs.Comparison for .NET API reference - full details on
LoadOptionsand theComparerclass - Free support forum - questions about external resource handling and comparison behaviour