Blog

Articles on automated cataloging, catalog formats and use cases.

From a sound recording to a catalog record

How Janium Collect catalogs collections with audio —sound recordings, interviews, oral minutes— transcribing the voice and producing a record in Dublin Core, MARC or ISAD-G, with the option to segment by intervention into linked child records.

From loose documents to Fonds and Series: generating the ISAD-G hierarchy automatically

How to automatically generate the ISAD-G hierarchy —Fonds, Section, Series— while cataloging: the exporter builds the tree from the classification scheme, raises to the upper levels what most of the children share, and adds date range and extent.

What to demand from an AI cataloging tool

Almost any AI cataloging tool looks good in a demo. These are the questions that reveal whether its records are reliable in your catalog, beyond that first impression: where each datum comes from, against which authorities it verifies, what it computes by rule, whether it is correct against the standard, whether it evaluates each record, whether it fits your collection, whether it respects your sovereignty and whether it is honest about its limits. They work for evaluating any vendor.

Colegio de La Salle: from a four-field Excel to a library system in the cloud

The library of a bilingual school in Bogotá kept its collection in a spreadsheet with four columns per book. How that list became MARC21 records with Collect, and a system that today catalogs, circulates and publishes its catalog online.

The Diocesan Library of Bilbao: cataloging a vinyl collection from a photograph of the record

An old collection of sound recordings —vinyl records of religious music in Spanish, Basque and Latin— is hard to catalog because its information lives on the sleeve, not in a database. How the library cataloged it with JaniumCollect from photographs.

Generic isn't enough: why a catalog of record needs more than an LLM that fills fields

A generic LLM already produces complete, plausible records; enriching is the value of AI in cataloging. What a catalog of record needs on top of that is knowing where each datum comes from: the source that backs each one, deterministic computation where there's a rule, per-record evaluation and flagging, the logic of the standard, multi-language. Along with the limits of enrichment from the model's knowledge.

Why trust a catalog made with AI

How a catalog made with AI becomes trustworthy: not by denying that the model enriches, but by knowing where each piece of data comes from. It rests on the source that backs each one, verification against authorities, deterministic computation of rule-based fields, per-record evaluation and a quality gate that stops degraded batches, with the limits made explicit.

How to describe an archive in ISAD-G with AI

What sets archival description apart from bibliographic description, why an archive needs ISAD-G, and how ISAD-G description is produced with AI —by provenance and hierarchy, not item by isolated item—.

Glossary: the AI and cataloging terms in this blog

Short definitions of the terms that cross this series: from AI (language model or LLM, frontier model, on-premise, air-gapped, anonymization, OCR, ASR) and from cataloging (MARC21, Dublin Core, ISAD-G, CDWA, authority control). For readers coming from libraries and archives who run into AI jargon, or the other way around.

Cataloging with AI without the material leaving the institution

Some collections cannot leave the institution — data residency, policy or confidentiality. How Collect catalogs with AI on the institution's own infrastructure, with a local model and even air-gapped; the hybrid mode with anonymization for the hard cases; and why the rare thing isn't sovereign AI or cataloging on their own, but the two together. With the limits of each mode.