← Blog

Glossary: the AI and cataloging terms in this blog

  • janiumcollect
  • ia
  • referencia

This series crosses two worlds: cataloging —libraries, archives, museums— and artificial intelligence. Each brings its own vocabulary, and not everyone knows both. This is a short reference to the terms that come up most. If you come from a collection, the AI section will help you; if you come from AI, the cataloging one.

Artificial intelligence

Language model (LLM)

The kind of artificial intelligence behind tools like ChatGPT: a program trained on enormous amounts of text that, from what it reads, identifies, summarizes and completes information. “LLM” stands for large language model. In this blog it is what reads the document and proposes the record.

Frontier model

The most capable and recent language models, run by large providers and reachable over the internet (in the cloud). They perform better on the hard cases, but processing with them means sending them the material. They contrast with a local model, which runs on the institution’s infrastructure.

Anonymization and tokens

Replacing the personal data in a text with neutral markers —the tokens— before sending it to an external service, and restoring them on return. It reduces the exposure of personal data, though it does not guarantee that none remains.

OCR — optical character recognition

Converting the image of a scanned document into text a machine can read. It is the prior step when the material is a scan with no selectable text.

ASR — automatic speech recognition

Converting speech into text. In audio and video it is what produces the transcript from which the record is extracted. “ASR” stands for automatic speech recognition.

Infrastructure and sovereignty

On-premise

That processing runs on the institution’s own servers, not on a third-party service. The material does not leave.

Air-gapped

An isolated network, with no internet connection. The highest degree of sovereignty: the system works without communicating with the outside.

Data sovereignty and data residency

The principle that certain information cannot leave the institution —or the country— that holds it, by law, by policy or by the type of content. Developed in Cataloging with AI without the material leaving the institution.

Cataloging standards

MARC21 / UNIMARC

The standard format for bibliographic records in libraries. It structures the description into coded fields and subfields.

Dublin Core

A simple set of fields to describe any resource in an interoperable way. Less detailed than MARC, easier to exchange.

ISAD-G

The international standard for archival description. It describes the material in a hierarchy —fonds, section, series, item— where provenance and context matter, not just the isolated piece.

CDWA

The standard for describing works of art and cultural objects, with the Getty vocabulary. It distinguishes the work from its reproduction.

Authority control

The mechanism that ensures the same person, entity or subject is always recorded in the same form, checking it against reference lists —VIAF, ISNI, the Getty vocabularies, or the institution’s own catalog—. It keeps the same author from entering the index under three different spellings.


Is there a term you would like to see here? Write to us at info@janium.com.