A university library does not describe everything it receives under the same standard, because not everything ends up in the same place. The new titles acquisitions buys go into the catalog and are described in MARC21. A final-year project is deposited in the institutional repository, described in qualified Dublin Core and, at many universities, never gets a catalog record at all. A report or a yearbook produced by the institution itself, when the university archives hold it, is described by fonds and series in ISAD-G.
When a university asks Janium Collect to help it catalog, what it asks for is not a generic record: it is the record for the destination it already runs. Collect reads a row from a list or the PDF of a document and produces the record for that schema. A cataloger reviews it before loading.
The bibliographic catalog
Acquisitions lists are the most direct case. The vendor’s file usually carries the author, the title, the ISBN and the year already, and each value is moved into the field the library assigned: a cell that says 2024 still says 2024 in the record. When the row is only enough to identify the book — at times it is little more than an ISBN and a shortened title — Collect completes the subjects and a summary, and marks as added whatever was not written down in the source. That path, from the spreadsheet to the record, is covered in Cataloging from lists and spreadsheets.
The author’s name is checked first against the library’s own authority file and then against VIAF and ISNI, so that the same person does not end up in the index under three different spellings; the problem and how it is handled are in The same author spelled three ways. The call number comes from the scheme the library already applies, whether UDC, LC or Dewey, and the Cutter number, when that scheme calls for one, is taken from the table rather than estimated, as explained in The Cutter number and the ISAD-G level.
The output is MARC21, or UNIMARC when that is the institution’s standard, serialized as a flat file (FLAT), ISO 2709 or MarcEdit depending on the wrapper the catalog accepts. Collect does not write inside Alma or any other ILS: it hands over a batch and the library loads it by whatever means it already has configured, whether a load from the interface, SFTP or a watched folder. Delivery to the catalog is described in From record to catalog.
The institutional repository
The repository works on a different logic. A final-year project or a master’s thesis arrives as a PDF, and its title page usually carries the title, the author, the year and often the supervisor’s name. Collect reads the title page, the preliminaries and the body of the document. The title and the author are carried over as they stand. The supervisor, when the title page states the name, comes in as a contributor. The abstract in the preliminaries feeds the description, and a subject the model recognizes in the body is added to the record marked as added, kept apart from the keywords the work itself already declared.
What the title page does not say is not inferred. The program and the degree
are not deduced from the university’s name. The identifier the repository
assigns on deposit and the URL of the PDF do not yet exist when the record is
produced, so those fields come out empty and the deposit fills them. The
university’s own rule sets the license and the embargo, and Collect proposes
neither. If the repository expects the supervisor or the degree in elements of
its own, such as dc.contributor.advisor or thesis.degree, that installation
is configured to emit them: the qualified Dublin Core that comes out of the box
does not include them. Where the schema falls short of a catalog is covered in
When to use Dublin Core and where it falls short.
The volume of the academic year sits in the final-year projects and the master’s theses: there are dozens or hundreds each year and they arrive undescribed. Doctoral theses are far fewer and usually arrive already described with some care. Articles, working papers and the institution’s own publications are handled like the projects: the record comes from the PDF.
Beside the PDF there is also a markdown copy with the abstract, the chapters and the text of the document, and the description and the search that sit next to the record come from there; the process is in From any format to a catalog record. Querying the repository over that text, if the university wants it later, is a separate layer: Access points, full text and RAG.
When the same work goes to both places
That a project is deposited in the repository and has no catalog record is ordinary practice at many universities, not a gap in the workflow. When the library does want the work in both places, the same PDF is read once per schema and two batches come out. Each process emits one schema, so if the document goes to a single destination only that record is produced.
A thesis that does enter the catalog carries, when the document states them, the author and added entries (100 and 700), the title (245), the publication (264), the physical description (300), the dissertation note (502), the summary (520) and the subjects (650 and 653). Fields the document does not state come out empty. The 856 waits for an address: as long as the repository has not assigned the handle, the field stays without a value. The name of the degree-granting university, when the title page does not state it, can come from a value the institution set in advance for all of its student work, rather than from a reading of the document.
The rest of the campus collections
A report, a yearbook, course material or a special collection that has already been digitized and has no record is handled like the student project: the PDF yields the record in the schema that collection uses. If the holding is the institutional archive, the unit is described by fonds and series in ISAD-G, which is a hierarchical description rather than a standalone card; Describe an archive, not a catalog goes into that difference.
In a work written in another language, the vocabulary of the description is translated, while the title, the names, the dates and the identifiers stay as they are in the original. The full criterion is in Cards in another language.
When the PDF contains personal data — and on a campus the minutes, the student files and some of the theses do — the processing can run on the institution’s own servers: Processing documents without exposing personal data.
Before loading
The cataloger receives the record with its provenance in view: what came from the row or the title page, what the model added, and what stayed empty because the document did not support it. An empty field is a pending task that someone settles with the missing source, which is why it is preferred over a plausible value nobody declared. The batch is loaded into the repository or the catalog once that review is good enough.
To continue the conversation
If the catalog and the repository at your university live under different criteria, and you want to see what record comes out of an acquisitions list or of a deposit of student work from the last academic year, write to us at info@janium.com. We are interested in how the vendor’s lists arrive and in what state the projects reach the deposit.