How to keep a degraded batch out of the catalog is a different problem from a single bad field. Why to trust a catalog made with AI covers trust in content: where each datum comes from, authority control, fields with a fixed rule, scoring per record. At the scale of thousands of documents another risk appears: something fails in the run and degrades the whole batch before anyone notices.
A batch of a thousand records missing links to their source files still looks fine —titles, dates, authors— until someone tries to open the document from the catalog and the link is gone. By then the batch is loaded. Collect does not keep a backup of the batch before cataloging: if a bad batch is delivered, there is no automatic undo. That is why the defences sit on not delivering: stop in the open.
Check cloud sources before starting
When an institution catalogs from Google Drive or another cloud source, that source supplies the link to the original document on each record. If Collect starts without having connected a cloud source that is enabled, it produces records that look complete but lack that link; a later run can overwrite ones that had it.
At start-up Collect checks that each enabled cloud source actually connected. If one did not, it stops and names which failed. Processing without them is not incomplete work to fix later: it is silently dropping links.
You can continue on purpose, with that source down. Then the error is logged and affected records are flagged for review. That is a decision, not the default.
A local folder does not trigger this check: it is not the case that loses the Drive link.
Stop the degraded batch before delivery
The second control sits at the end, when the batch is packed for the catalog —the hand-off in from record to catalog. Collect compares what is about to go out with what came in. If there are fewer links to the original than at the start, or a record arrives without the form type the catalog needs, the load is rejected: the process ends in error, not with a footnote nobody will read.
It does not repair the batch. It stops it so it is not loaded into the ILS.
The operator can force delivery and accept the result. The scope is batch integrity against the input, not the accuracy of every field. A batch can pass the gate and still contain a mistyped date; that is content review.
Limits
There are two controls in the process: at start-up (cloud sources) and when the batch is delivered. They cover links to the original and that each record carries its type. They do not check every field. There is no automatic backup before cataloging: a later run overwrites what was there. The defence is not to load the degraded batch, and that running with a downed source is an explicit choice.
If your institution processes large batches and you want to see how these controls behave with your sources, write to us at info@janium.com; we can review which ones apply in your case.