← Blog

Cataloging a collection video with AI: from the recording to linked records

  • JaniumCollect
  • AI

An audiovisual collection poses a problem a book does not: the content is not written down anywhere. A thirty-minute newsreel or a recording of a session does not come with a text that describes what happens inside. To catalog them you have to watch, listen and note what they are about —and an hour of material may touch ten different subjects, each useful to someone looking only for that fragment. Describing the whole video with a single record leaves out almost everything it contains.

Janium Collect treats the video with the same destination as a document: a catalog record. It reads what is said and what is seen, and a language model proposes the metadata. By default one record comes out for the whole video. If the material warrants it —a newsreel with several items, a session with several points— you can also ask for a record per scene, linked to the first, in the same format as the rest of the collection: Dublin Core, ISAD-G or MARC21. The sound-only path is in From a sound recording to a catalog record.

One record, or one record per scene

Each scene is searchable on its own, and at the same time it is clear which video it belongs to. Anyone looking for a topic finds the fragment; anyone who wants the whole thing finds the record for the entire video.

What is heard and what is seen feed that description. Duration and file type are read from the file itself; the model does not estimate them.

Limits

The visual description does not replace viewing by someone who knows how to describe audiovisual material. It comes from representative images, not from the video frame by frame: a gesture or a brief shot may not be reflected.

The split into scenes is a starting point for navigating the material, not the editorial cut someone describing it would make. A long scene may contain two subjects; two different shots may treat the same one.

If there is no intelligible speech —music, atmosphere— the transcription adds little, and a description is not forced from it. What cannot be supported is left empty; it is not presented as verified. All of this reaches human review, from a draft already structured and linked, not from scratch.

If your institution has audiovisual fonds —newsreels, video recordings, recorded sessions— waiting for description, watching and annotating them one by one is exactly where this flow can save time, without taking the last word from the cataloger. If you want to talk about what your video collection is like and how you describe it today, write to us at info@janium.com.