How to generate a MARC21 or UNIMARC record with AI is decided by the standard the institution already uses, not by the model. A library cataloger does not choose the record format for taste: the institution’s cataloging standard, the system where the catalog lives and, often, the tradition of the region fix it. The same book can be described in MARC21 or in UNIMARC: two different field schemes for the same work. Janium Collect produces the one the institution already catalogs in and delivers it in the wrapper the receiver accepts.
MARC21: the library standard
MARC21 is the bibliographic record format most widely used in libraries worldwide. It organizes the description into numbered fields —100 for the main author, 245 for the title, 260 or 264 for publication, 300 for physical description— and, within each field, into lettered subfields. That granularity is what lets a catalog tell an author from a title and sort, search and share records among systems that speak MARC.
Many libraries that catalog in MARC21 apply the RDA rules (Resource Description and Access), the cataloging standard that succeeds AACR2. RDA is not a record format: it is the set of rules with which MARC21 is filled. It defines, among other things, the punctuation that separates and closes subfields —the spaces, hyphens, periods and slashes a cataloger recognizes at once in a well-formed record— and the content, media and carrier type fields. That punctuation is not decided by the model: it is applied deterministically in post-processing, after the model proposes the content. A fixed rule is better kept by a rule, not by asking a model that can get it right almost always but not always.
UNIMARC: the alternative in some regions
UNIMARC is a bibliographic record format with the same vocation as MARC21 —describing the resource with fields and subfields— but with its own field scheme and a different history. It was born as a universal exchange format driven by IFLA, and several national libraries and networks, especially in parts of Europe, adopted it as their local standard. For an institution that catalogs in UNIMARC, delivering MARC21 would not be compatible with its catalog; it needs records in its own scheme.
Producing MARC21 or UNIMARC takes the same work: choosing the right field, splitting the subfields cleanly and respecting the indicators. Collect leans on a language model that can hold that structure, in either format.
Two major standards coexist because they answer to real cataloging traditions, and each institution already knows which one it works in. Collect does not impose either.
Interchange wrappers
The description format is MARC21 or UNIMARC. How that record travels to the catalog is another decision. The exporter delivers the wrapper that fits:
- MarcEdit text (
.mrk): the MARC21 or UNIMARC record written as readable text, one field per line (=245 10$aTitle…), the form used by the tool of the same name. It is for reviewing and correcting by hand, or for loading with MarcEdit. The record remains MARC21 or UNIMARC. - ISO 2709 (
.mrc): binary MARC, when the receiving system expects the canonical form. - MARCXML (
.xml): the same record in XML, the Library of Congress schema, when the receiver expects XML rather than the binary form or MarcEdit text. - FLAT: Janium’s native interchange, when the destination is a Janium catalog.
A MARC21 record and a UNIMARC record can leave in any of those wrappers, according to what the receiver accepts. Before delivering the batch, Collect checks its integrity and stops it if it detects degradation, so damaged material is not loaded into the catalog. How that file reaches the ILS —load batch, SFTP, folder— is in Integrating AI cataloging with your catalog.
Archival description, by provenance and hierarchy, is another format: How to describe an archive in ISAD-G.
Limits
Collect does not decide for the institution whether the catalog is MARC21 or UNIMARC. That choice is fixed by the cataloging standard and the system where the catalog lives; Collect adjusts to it. It also does not guarantee that RDA punctuation covers every edge case of the standard: it applies the punctuation rules in post-processing to the content the model proposes, and the particular cases a deterministic rule does not contemplate are left for the cataloger to review.
The model completes the record with what the source did not bring when it can sustain it, and marks what it inferred: an inferred datum is not presented as verified, and what cannot be sustained is not filled in to satisfy the format.
To continue the conversation
If your library catalogs in MARC21 or in UNIMARC, and you have material waiting for description, we can review how your standard fits what Collect produces and in which wrapper you load into the catalog today. Write to us at info@janium.com.