A catalogue manager’s fear is not that an import will fail. A failed import is obvious, can be corrected and restarted. What is frightening is an import that succeeds and wipes out six months’ work: carefully crafted descriptions replaced by two lines from the supplier, visuals selected by your teams replaced by generic packshots, and negotiated prices reverted to the list price. This article describes the safeguards that ensure a change to the product database is verifiable beforehand, and reversible afterwards.
Let’s clarify the scope, as three different stages are often confused. The retrieval of files — FTP, SFTP, EDI, collection scheduling — is covered in automating your supplier imports via FTP and EDI. The transformation of values — units, encodings, attribute mapping, duplicates — is covered in standardising data from twenty suppliers. Here, the file has arrived and the values are clean. What remains is the riskiest stage : writing to your database.
The dreaded scenario
One Tuesday morning, the import from the main supplier runs just as it has on the fifty previous occasions. On Wednesday, a sales representative reports that a product record displays a description spanning three lines instead of the text rewritten last winter. Upon investigation, the team discovers that thousands of product listings are affected, that the edited images have been replaced and that no one can say when this happened. The source file, meanwhile, has already been overwritten by the next day’s file.
Three causes crop up almost every time.
- An incomplete file mistaken for a complete one— the supplier sent an extract from a product range, but the import was configured to replace everything. Anything not included in the extract was deactivated or cleared.
- An empty column interpreted as an empty value— this is the most insidious cause. A cell with no content does not mean ‘the value is empty’; it means ‘this supplier does not fill in this field’. Without an explicit distinction between
absentandvide, an export where the ‘description’ column is no longer populated will delete all your descriptions. - The mapping was modified without being checked— someone added a mapping in a rush on a Friday evening to include a new column. The change was correct for the twenty lines tested but incorrect for the thirty thousand others.
What these three cases have in common is that none of them produced a technical error. The import ran exactly as instructed. That is why monitoring failures offers no protection here — you have to monitor the successes.
An import is a transaction, not a copy
The natural instinct is to view an import as a copy: I take the file, I populate the records with the values. It is this mental model that causes problems. An import is a transaction: a set of operations that are either applied in one go, or not at all, following verification. Four conditions must therefore be met before a single value is written to the database.
- The file structure is correct— the expected columns are present, in a readable format. A file missing a required column is not processed in a degraded mode; the process stops.
- The volume is plausible— a variation threshold compares the number of rows received with the last successful import. A drop by half suspends processing and requires human validation, as this is almost always a truncation of the transfer, not a mass delisting.
- The scope of writes is defined— this import is authorised to access specific suppliers, specific families and specific fields. Nothing beyond that. A supplier import that can write anywhere is a ticking time bomb.
- The previous state is preserved— before writing, the current version of the relevant records is archived. Without this snapshot, no rollback is possible, regardless of the quality of the rest of the system.
The dry run
The simulation applies all the rules to the received file and produces a report of what would happen, without actually writing anything to disk. This is the only time when making corrections takes just a few minutes rather than several days.
A useful simulation report fits on one screen and answers five questions.
| What the report shows | What you check | Warning signal |
|---|---|---|
| New entries | Number of new SKUs and their family | Mass creations on a flow that is supposed to be stable: the reconciliation has failed |
| Updates | Number of records affected and number of fields modified per record | An average number of modified fields significantly higher than usual |
| Deletions or deactivations | Outgoing references and reason | Any deletion not explicitly requested by the supplier |
| Fields affected | Breakdown of entries by attribute | A field that this supplier is not supposed to populate appears in the list |
| Before and after sample | Around ten actual lines, with the old and new values side by side | An enriched value replaced by a raw value from the supplier |
The before-and-after sample is the section that teams read the least, yet it catches the most errors. Sort it by number of modified fields in descending order: the rows most affected appear at the top, and any mapping anomaly is immediately apparent. A simulation that simply states the total number of updates without specifying which ones is of no use.
Locked fields
A catalogue combines two types of data: that held by the supplier and that which you have generated. Confusing the two is the real cause of overwrites. Locking involves declaring, once and for all, which is authoritative for what.
| Field | Authority | Behaviour on import |
|---|---|---|
| Manufacturer reference, barcode, technical specifications | Supplier or manufacturer | Free overwriting |
| Purchase price, availability, packaging | Supplier, most recent source | Free overwriting, with history of the previous value |
| Product description, optimised title, selected images | Your teams | Locked — incoming values are treated as proposals, not as entries |
| Selling price, internal categorisation, marketplace attributes | Your teams | Locked without exception |
Locking is defined at three levels of granularity, to be combined as appropriate.
- By field— the default setting, applicable to the entire catalogue: no one can overwrite the long description via an import.
- Byproductfamily— useful when the rule depends on the product type. For consumables, the supplier description is sufficient; for a reworked premium range, it is locked.
- By data source— the most granular: the field remains editable via an import, but only if the existing value also originated from an import. As soon as a human has edited the field, it locks automatically. This is the setting that requires the least maintenance, provided that each value carries its source and date.
A lock must never cause incoming data to disappear: the value from the supplier remains stored alongside it and can be adopted with a single click. The lock protects against automatic writing; it does not prevent human decision-making.
Versioning and rollback
Even with a simulation being run and fields locked, an import will eventually go through. It is therefore necessary to be able to roll back without having to request a new file from the supplier. To this end, each execution becomes an identified object: date, supplier, source file retained, in their current version, operator, associated simulation report. Retaining the source file is just as important as retaining the result — it is the only way to replay an import when it is discovered that the rule, and not the data, was incorrect.
Rollback is then organised into three levels of granularity, from the broadest to the most precise.
- The entire import— the whole batch is rolled back and the repository returns to its state from the previous day. This is the solution for configuration errors discovered within the last hour.
- A supplier or a scope— we roll back what a feed has written without affecting other imports that have taken place since. Essential when several suppliers are feeding the same records.
- A product or a field— the most common scenario in practice: a record flagged by a sales representative, restored to its previous version without affecting the rest of the import.
A clarification to avoid a pitfall: undoing an import must never delete the records it created if they have since been ordered or published. Undoing the action restores values and statuses; it does not destroy business objects.
The error log
A technical log that simply displays a parser exception is of no use to anyone in a catalogue team. A useful log is written for the person who will correct the issue, i.e. a business administrator, not a developer. One line per anomaly, and on each line four pieces of information: where (file line and product reference), what (the field and the value received), why (the rule breached, in plain English), and what was done (value rejected, line ignored, value replaced by the default, field locked and therefore not written to).
Three features distinguish a reviewed log from an ignored one: it can be filtered by supplier, by field and by type of anomaly; it can be exported to be sent back to the supplier as is; and it can be re-run once the corrections have been made. A manager can then deal with these anomalies themselves rather than opening a ticket. This is the approach adopted in the import and data quality modules of Pixee PIM: simulation before writing, per-row log, history per import and rules stored per supplier.
What to monitor afterwards
The safeguards described above protect the day’s import. What remains is to detect issues that deteriorate slowly, over weeks, without ever triggering an error. Three indicators are sufficient, recorded at each run and compared with the history of the same data flow.
- The volume of fields modified per import— a stable data stream writes a comparable number of fields from one week to the next. A sudden doubling signals a change in format on the supplier’s side or a mapping that has shifted, well before anyone complains about it.
- The rejection rate— particularly useful when it drifts. A rate that rises gradually over the course of a month indicates a deteriorating supplier export. A rate that drops abruptly to zero is just as suspicious: a validation rule may have been disabled.
- Silent drift— the hardest to spot: the number of locked fields that receive a suggestion, the proportion of records whose description reverts to the supplier’s, and the completeness score by family. None of these metrics trigger a technical alert, but their trend reveals the erosion of your data enrichment efforts.
Frequently asked questions
Should a field be locked or left in suggestion mode?
Both approaches are used in combination. Locking prevents automatic updates, whilst suggestions retain the incoming value and present it for approval. The wrong setting is one that locks the field whilst discarding the data: you will never know that the supplier had updated their technical record.
How long should the import history be kept?
Long enough to cover the time between an error occurring and its detection, which is rarely immediate: a description anomaly is usually flagged by a sales representative or a customer, not by an alert. Several full import cycles is a good benchmark — the cost of storage is incommensurable with that of re-enrichment.
Does the simulation slow down the import cycle?
It adds a step, but not necessarily a manual intervention. For a new data feed or one whose format has just changed, a human review of the report is essential. For a stable data feed, it runs automatically and only requires validation if a threshold is exceeded — abnormal volume, unexpected field, or deletions detected.
Can these safeguards be tested without involving the entire catalogue?
Yes, and that’s the right way to proceed: one supplier, one product family, one full cycle monitored over a few weeks before scaling up. Pixee PIM’s 30-day free trial — 2,000 products and 2 suppliers, with no credit card required — is sufficient to validate a locking configuration and view your first simulation reports. Details of volumes per offer can be found on thepricingpage.
Import without overwriting your enrichment work
Dry run before writing, per-field locking, per-import history and targeted rollback — try it on one supplier during the 30-day free trial.
See the plans