Anyone coming from the library into the archive tends to bring a habit that does not work: describing each piece as if it stood on its own. A book does stand on its own. Its record describes it in full —author, title, edition, subject— and that record is enough to find it and use it, no matter what sits on the shelf next to it. An archival document is not like that. A loose memo, with no way to know who produced it, which file it belonged to or what came before it, says little. Its meaning comes as much from its content as from its place within a whole.
That difference is what separates bibliographic description from archival description, and it is the reason ISAD-G exists —the General International Standard Archival Description, from the International Council on Archives—. This text is about what it is and why an archive needs it. The mechanism Collect uses to generate the upper levels is left for another post; here what matters is the principle.
Provenance and original order
Two ideas hold up all of archival practice.
The first is the principle of provenance: the documents of a single producer —an agency, a company, a person— are kept together and not mixed with those of another, even when they speak to the same subject. A sale contract from Notary A and another from Notary B do not go together in a “sales” folder; each belongs to the fonds of its notary. What groups them is not the subject, it is the origin.
The second is respect for original order: the archive keeps the organization the producer gave to its documents, instead of reordering them by external criteria. That order is evidence. It tells how the office worked, which procedure followed which, how matters were grouped. Reordering alphabetically or by date may look tidier, but it erases information the archive was meant to preserve.
Everything else follows from these two ideas. An archive is not a collection of independent pieces cataloged one by one; it is a structure that is described as a whole.
From the whole to the part
That is why ISAD-G describes in levels, from the larger whole down to the smallest unit:
Fonds → Section → Series → File → Item
The fonds is everything produced by a single body or person. Within it, the sections correspond to divisions —areas, functions—; the series group documents generated by a single activity (correspondence, minutes, payrolls); the file brings together the documents of a single matter; and at the bottom of the hierarchy is the item, the concrete piece.
Description runs from the top down, and each level inherits from the one above. What is said in the fonds —who produced it, its history, the conditions of access— is not repeated in each file or in each document; it is described once, at the level where it belongs, and applies to everything hanging below it. ISAD-G calls this the rule of non-repetition. To describe an archive as if it were a catalog of flat records —each document with its producer, its history and its conditions repeated— not only takes more work: it loses the structure that made the whole legible.
What Collect produces, and what it needs to do it
Janium Collect reads the documents of a holding and produces records in the format the institution uses. For archives, that format is ISAD-G, with its levels and its logic of provenance.
Collect describes in a particular order. First the items —what is in plain sight in each document— and, from them, its hierarchical exporter generates the upper levels: it recognizes which series, section and fonds each item belongs to, builds the tree and raises to the level where they belong the data the items share. The result is a description that runs from the fonds to the piece, built from the piece upward.
The hierarchy is not invented, but neither does it require that provenance come declared in advance. Collect builds it from the signals it finds in the material. When they are explicit —a classification, a folder structure, data that say which series and fonds each item belongs to—, it builds the tree directly. When there are none, it leans on patterns in the content itself: if it recognizes that a document is an invoice and identifies its issuer, it infers which section it corresponds to; by grouping those inferences it proposes an organization that was not given. That proposal is a starting point for an archivist to review and correct, not a verdict.
The limit is where there is no signal left to read. A set of loose scans, with no classification, no structure and no recognizable patterns in their content, gives nothing from which to deduce a hierarchy, and forcing one would be inventing it. There the organization is prior human work, and Collect describes what already has an order. That is also the criterion for when it contributes on the hierarchical side: when there is provenance to respect or clues to start from, not when none remains.
To continue the conversation
ISAD-G is not one more format among MARC and Dublin Core; it is the way to describe when what matters is where a document comes from and which others it goes with. If your institution has a fonds to describe and you want to see how level generation behaves with your own provenance structure —or to discuss where its limits are with your concrete material—, write to us at info@janium.com; we are interested in understanding how your archive is organized today.