Skip to main content

For libraries, special collections & public archives

Let a reader ask forwhat they cannot name.

The Main Reading Room of the Library of Congress, seen from the gallery above
Main Reading Room · Library of Congress

A catalogue answers the reader who already knows the vocabulary of the finding aid. Humy puts a conversation over the records your staff curates, and every quote it returns is copied out of your scan.

0
quotes the AI made up, across 1,323 checked
0
of your records used to train any model
1,000
pages per record, handwriting included

A catalogue rewards the reader who already knows the words. Everyone else leaves with nothing.

Undescribed holdings are, in the profession's own words, all but unknown to the scholars they were kept for. Dooley & Luce, OCLC Research, 2010 (survey of 169 institutions)

How a question becomes a citation
  1. 01

    Your staff choose the records

    Each one keeps its creator, date, repository and accession number. Restricted material simply stays out.

  2. 02

    Scans become searchable

    Pages with no text layer are transcribed one at a time, secretary hands and copperplate included.

  3. 03

    The reader asks in plain words

    The answer comes back with the passage, the record it sits in, and a route to the scan.

I have seen the misery of the poor in this parish, and I say plainly that the mill has made it.

Rev. T. Hollis, Letter to the Parish Vestry · 14 March 1843 · MS 4471, f.212 · illustrative

The AI never types a quote. It marks the passage; our system copies the words out of your document and discards anything that is not an exact match. A citation to a record that does not exist cannot reach the screen, because there is nothing to copy.

On invented sources

Ask any AI for a citationand it will invent a convincing one.

This is the first objection in every conversation we have with a repository, and it should be. Ours is not permitted to write one.

0
quotes that did not match the document
0
highlights on the wrong passage
0
quotes from outside the figure's own records

Humy internal benchmark, August 2026: one configuration, across the 1,323 quotes it produced. One citation in that run opened nothing, and we report it rather than round it away.

What that does not cover: the sentences around a quote are still written by the AI. In the same benchmark, roughly one answer in eight carried an invented date or reference in that surrounding text. The quote itself is always real. That gap is precisely why the record sits one click away. Read the benchmark.

Custody & control

A donor who signed in 1974was not contemplating a model.

So an AI request is a question about custody before it is a question about access, and it lands on whoever signed the agreement. Three things settle it: can an answer be traced to the item, does the attribution hold when someone checks, and can the institution make the use stop.

  1. 01

    Your material is never used to train a model

    Nothing you load enters training data, ours or a model provider's. This is the part that cannot be undone later: once archival material has trained a model, there is no pulling it back out. Your documents are read when a reader asks, and never learned from.

  2. 02

    You can require us to stop, and to delete

    Stop-use and deletion go in the agreement, not a support ticket. Take a simulation offline, withdraw a record, or end the arrangement and require removal, and the obligation is contractual. A promise that lives only on a web page is worth what it costs to edit one.

  3. 03

    Each item carries its own source details, not just the collection

    Every record is described by hand: creator, date, document type, holding repository, accession number, and whether it is first-hand or a later account. That description travels with each citation, down to the position on the page, so a reader can see what kind of evidence they are holding.

  4. 04

    Visibility is set record by record

    A source can stay private to the account that holds it, be shared inside your institution, or be published openly. Material whose rights are still being cleared can sit in the workspace unreleased and go live later. Nothing changes state on its own.

  5. 05

    Restricted material is kept out, not hidden

    Embargoed files, donor restrictions, personal data and culturally sensitive records stay out of the archive altogether, so the archive itself enforces the restriction. A prompt can be talked around. A document that was never loaded cannot be quoted.

  6. 06

    It does not rule on what the record means

    A simulation is a route into the papers. It does not rule on what they mean. Answers end at a document, and the authoritative reading is still the one in your finding aid. Where the record is contested, the documents should carry the contest.

  7. 07

    Reference work is not being replaced

    Everything a reader can reach is material your staff selected. We have no data on how this shifts a desk's workload. What we can say is structural: a question that arrives after the reader has already been routed to a box and a folder is a different question.

What we have built

Three collection builds,all of them with university libraries.

01National Louis University · Chicago

A founder built fromher own papers.

The library supplied 25 of Elizabeth Harrison's writings, photographs and documents, and we built the university's founder from those and nothing wider. Students now question Harrison about what she actually wrote, and can open each cited item beside its scan.

For a library the interesting part is what the build asked of the institution: a selection decision on every item, a description attached to each one, and a release decision before anything went live. That is the same workflow a reading-room deployment would run, at a size a single collections librarian could carry.

Also in progress
  1. 02

    University of Wisconsin–River Falls

    Wisconsin, United States

    Collection-based figures in development with the university.

  2. 03

    Private University of Education, Diocese of Linz

    Austria

    Ten licences under a one-year research collaboration, with collection-based figures in development.

Where we have not been yet

These are university-library builds. We have not run a deployment inside a state or national archive, and we will not claim a track record we do not have. The repository that goes first will shape how the tooling handles series-level description and restricted series.

The nearest thing to public scale we can point at comes from an exhibition, not a repository: an AI Hatshepsut built for South Florida PBS from more than 60 curated artifacts, which opened in October 2025, ran for eight months, and reached 60,000 visitors.

Starting a pilot

Name one collection,and three documents at its heart.

A first build works best scoped to one collection. Repository-wide scope can wait. Tell us which one you would start with, the three documents you would want a reader to reach first, and whether the subject is a person, an event, or a place, and we will scope the work against them on a call.

Collection builds are scoped with our team. There is no self-serve path, and pricing is quoted against the collection once we have seen it.

Frequently asked

What repositories ask before the first record goes in.

  • It gives a patron a way into the material by conversation, so keyword search stops being the only door. You curate a set of records from the collection, each carrying its own metadata — creator, date, document type, holding repository, accession number, and whether it is first-hand or a later account. A reader asks a question in their own words, and the answer comes back with the passage and a route to the scan it came from. Material that would otherwise need a reader to already know what to search for becomes reachable.

  • No, and it is not designed to. Everything a patron can reach is a set of records your staff selected, described, and released. What it changes is the volume of first-contact questions that never reach a reference desk at all, either because the reading room is closed or because the reader did not know the collection existed. Staff time moves toward the questions that genuinely need an archivist.

  • By never letting the AI write a quote at all. It marks which part of which record it means, and our system copies the words straight out of the document, throwing away anything that does not match letter for letter. Across the 1,323 quotes in our published Humy 2.6 test, none had been reworded, none highlighted the wrong passage, and none came from anything but that figure's own records. One citation out of the 1,323 opened no source, and we report that rather than round it away.

  • Saying so is the behaviour we test hardest. One of the five reasoning categories in our 2.6 benchmark is exactly this: recognising an unrecorded fact, a false premise, or an apocryphal quotation and declining to fill the gap. Grounding helped most there, moving from 24.6% of checks passed without an archive to 59.3% with one. A confident answer to a question the records do not support is the failure mode this whole design exists to catch.

  • No, and the distinction matters for a repository. The quotation is compiled from your document and cannot be fabricated. The sentences around it are model-generated, and in the 2.6 run roughly one answer in eight still contained an invented date or document reference. We therefore put the record one click away from every claim, so a reader can check rather than trust.

  • Up to 50 source records per figure and up to 1,000 pages per record: scanned books and periodicals, correspondence, catalogue and archive records, photographs, paintings, audio recordings, and 3D-scanned objects. Pages with no text layer are transcribed one at a time by Gemini 3.1 Pro under an archival transcription prompt, which makes copperplate and secretary hands searchable. You approve an estimated cost before any page is processed.

  • You do, record by record. Each source can stay private to your institution, be shared across it, or be published openly. Restricted, embargoed, or culturally sensitive material simply stays out of the set you release, and because a figure can only answer from records you released, exclusion is enforced by what the archive contains. No model is asked to keep a secret.

  • No. Access can run through a plain link or a QR code with no account, email, or password, which suits a catalogue page, a reading-room sign, or an exhibition panel. Where you want conversations attributable, named entry can be required on a particular link instead.

  • Yes. At National Louis University in Chicago we built a figure of the university's founder, Elizabeth Harrison, from 25 of her writings, photographs, and documents supplied by the university library. Collection-based builds are also running with the University of Wisconsin–River Falls and the Private University of Education, Diocese of Linz in Austria. These are scoped with our team rather than self-serve.

  • With one figure and a small, well-described set of records. The whole repository can wait. Tell us the collection you would start with and three documents that sit at the heart of it, at hello@humy.ai, and we will scope a pilot against it.

Something here we have not answered? Write to hello@humy.ai.