Why store sync duplicates product images, and how to stop it properly
The WooCommerce media library fills up on its own, filenames chain into -1, -2, -3 and the disk fills with them. Why the REST API uploads a new attachment on every sync, and what the check has to do so an image is linked rather than duplicated.
The symptom: a media library that fills itself
A store has three thousand products and five thousand images. The media library holds thirty thousand. The names repeat: product.jpg, product-1.jpg, product-2.jpg, product-3.jpg. The author of all of them is the user who once entered the connection details for the back-office system.
The upload date is every day, or every hour if the sync runs hourly. The images look identical because they are identical. The product gallery shows one image; the server stores twenty-four copies a day.
This is not a WordPress bug and not a server bug. It follows from how the REST API reads the images field, and no cleanup plugin removes that cause. Cleanup removes what has accumulated; the cause stays and uploads a fresh batch tomorrow.
Why images are uploaded rather than just linked
When a back-office system sends a product through the REST API with an images field containing an entry that has src and no id, WooCommerce does not read that entry as a reference to an existing image. It reads it as an instruction: download the file at this URL and add it to the media library.
The result is a new attachment. A new ID, a new database row, a new file on disk, new generated sizes. On every sync again, because an entry with src and no id always means the same thing.
From WooCommerce's point of view this is correct behaviour. The API has no way to guess which existing attachment you meant, so it does what you wrote. What is wrong is the sender's assumption that the receiver will recognise the image.
The cost to the server scales with how often you sync, not with how big the catalogue is. A store with five hundred products syncing hourly creates more files than a store with five thousand products syncing daily.
Why filenames chain into -1, -2, -3
WordPress does not allow two files with the same name in the same folder. When an uploaded file hits an existing name, the unique-filename function appends a counter: product-1.jpg, then product-2.jpg.
That counter is exactly what hides the problem. The product gallery shows one image, because the gallery shows the last one. The media library holds as many as there have been syncs. A chain of a hyphen and a rising number over the same name stem is therefore the most reliable sign that something is uploading images programmatically and without a check.
The counter also means comparing exact filenames is not enough. The check has to recognise that product-7.jpg is the same image as product.jpg, or on the eighth sync it will upload yet another one.
Watch out for generated sizes too. Alongside the original, WordPress writes files with the dimensions in the name, for example -300x300. Those are not duplicates and must not be treated as copies.
What happens when you leave images out of the payload
The second trap is the opposite one, and more expensive. The images field you send REPLACES the gallery. It does not add to it.
If the product in the store has five images and you send one, it will have one. If you send an empty array, it will have none. So the answer to duplication is not to simply stop sending images; that also strips out whatever an editor added by hand.
The correct answer is to send the whole gallery, with existing images listed by id and no src, and new ones by src. An entry with id and no src only links the existing attachment and uploads nothing.
A consequence of that rule is that the sender has to know the state of the gallery in the store before sending it. A sync that only writes and never reads cannot get this right.
Shopify behaves the same way, in different words
Shopify uses different naming and the same logic. A product update containing an images field replaces the set. An entry with src and no id creates a new image.
The difference is that Shopify keeps the original on its own content delivery network rather than in a folder you can browse over FTP. So the problem does not surface as a full disk partition but as a rising image count on the product and a slower product page.
Because the logic is the same, the check has to be the same. A store running both channels gains nothing from writing two solutions. What makes sense is one check that returns the same answer for both platforms, plus two thin translators into the shape each API expects.
How to tell sync apart from manual editing
Before you blame the sync, check. Three indicators are enough, and all three are visible in the media library.
Author. Attachments created by the sync have as their author the user whose credentials set up the connection. With manual uploads, the author is whoever uploaded the file.
Rhythm. Upload times at even intervals, for example every hour on the same minute, are a machine. A person does not upload at 3:00, 4:00 and 5:00.
Name. A chain of a hyphen and a rising number over the same name stem is a programmatic upload. An editor replacing an image usually names the file something else.
If all three line up, the cause is not the editor, and cleaning the media library is not the fix.
How the check is done correctly
The check has to answer one question: does an attachment for this image already exist in the media library. A reliable answer is possible without comparing file contents.
The basis is the filename from the image URL, normalised the same way WordPress normalises it on upload: no extension, no accents, lower case, punctuation turned into hyphens. The normalised key is what gets compared.
Then you strip the trace of duplication: the trailing counter on the name. product-7 and product are the same key; product-7-grey and product are not.
When the check finds an existing attachment, the gallery entry carries its id and no src. When it finds none, the entry carries src and a new attachment is created once rather than every hour.
The whole gallery is assembled before sending, because the images field replaces the set. Manually added images are preserved, existing ones are only linked, and a new file appears only when the image really is new.
Three traps in filename matching
The first is normalisation. WordPress cleans the name on upload: it strips accents and special characters. An image called Cevlji Rjavi.jpg with accents is not stored under that name in the media library. Comparing against the raw name from the URL therefore finds nothing and uploads a new file.
The second is generated sizes. Alongside the original there are files with the dimensions in the name. Those are not matching candidates and the check must never offer one as the product's existing image.
The third is an extension swap. Image-compression plugins create a variant in a newer format on upload. The name stem stays, the extension changes, so compare the stem without the extension rather than the whole filename.
And one trap that has nothing to do with names: parsing URLs with multi-byte characters. The standard URL and path parsing functions are not multi-byte safe, so build the key with expressions that operate on bytes guaranteed to be inside ASCII.
What to do about the duplicates already created
Fixing the cause stops new ones appearing; it does not remove the existing ones. Removal is data deletion, so it is not a job for a script running in passing.
A procedure that does not end in empty galleries: back up the uploads folder and the database first. Then build a list of attachments that match the duplication pattern and that no product, no content record and no gallery refers to. Review the list before deleting anything.
Deletion should go through WordPress's own attachment delete function, not by removing files in the folder. Deleting the file leaves the database row behind and the gallery shows an empty frame.
Work in batches and check a few products in the store between batches. Ten thousand deletions in one transaction is a bad idea even when the list is right.
How often the sync should run
Frequency is a decision about what genuinely has to be refreshed. Stock and price change several times a day and need frequent syncing. A product image changes when somebody replaces it.
So split the flows. The frequent sync carries stock and price. Images go only when the gallery really changed, which the system knows from the product's last-modified timestamp.
This is not an optimisation for speed. It is a reduction in surface area for mistakes: a flow that does not send images cannot duplicate them either.
When you do send images, the check should always run, with no option to switch it off. A setting that disables the check is the road back to the same full media library six months from now.
Comparison
| Entry in images[] | What WooCommerce does | Consequence |
|---|---|---|
| src without id | downloads the file from the URL and creates a new attachment | a new image on every sync |
| id without src | links the existing attachment | nothing is uploaded |
| id with name or alt | updates the existing attachment's metadata | no new file |
| one entry instead of the whole gallery | replaces the gallery with that single entry | manually added images are lost |
| empty array [] | deletes the product's images | the product ends up with no image |
See how Entexia checks whether an image already exists in the store before every sync, and links it instead of duplicating it.
Start free →