Skip to Content
Humy 2.6 – A 97% Increase in Historical Reasoning

Release note & benchmark results

Humy.ai 2.6 – A 97% Increase in Historical Reasoning

Language models invent sources and make mistakes about historical facts, an audit of 111 million references across arXiv, bioRxiv, SSRN and PubMed Central  counted a conservative 146,932 hallucinated citations in 2025 alone.

Humy uses LLMs to teach history, so those hallucinations are our problem directly. A LLM simulation of a historical figure that invents a letter, a date or a quotation hands a student something false with a citation attached. We upgraded the platform with agentic RAG (retrieval-augmented generation) so it answers from authentic historical documents rather than from just model’s memory, then put 1,250 historical questions to six LLM configurations to find out what it was worth.

97%

higher score on historical reasoning than the identical model answering with no archive attached. 58.2% of rubric checks passed against 29.6%.

Humy RAG · Gemini 3.6 Flash58.2%
Gemini 3.6 Flash, no sources29.6%

What the number means

How we evaluated Humy.ai 2.6

We put 1,250 open-ended questions about Napoleon Bonaparte to six AI configurations, grouped into five categories of historical reasoning, and had every answer tested against a rubric written in advance. Two of the six are Humy’s grounded pipeline, which searches a curated archive of seventeen historical documents, mostly memoirs, correspondence, biographies and histories, and attaches every quotation to the passage it came from. The other four are comparable models answering the same questions with the same persona and no archive at all.

The headline number is one model measured against itself. Gemini 3.6 Flash passed 29.6% of the checks with no sources and 58.2% running inside Humy, on the same question set and graded by the same judge: Gemini 3.6 Flash again, which sees the answer and the passages it cited but cannot search the archive. That is a 97% increase, or 28.6 percentage points. Against the strongest bare model in the field, Gemini 3.1 Pro at 31.0%, the grounded pipeline still passes 88% more.

58.2% is the share of those rubric checks the answers satisfied.

Checks passed

Share of the rubric checks each configuration passed. Higher is better.

Humy, archive attachedBare model, no sources

Humy RAG · Gemini 3.6 Flashgemini-3.6-flash
58.2%
Humy RAG · Gemini 3.5 Flash-Litegemini-3.5-flash-lite
41.8%
Gemini 3.1 Prono sources
31.0%
Gemini 3.6 Flashno sources
29.6%
Gemini 3.5 Flash-Liteno sources
24.9%
GPT-OSS 120Bno sources
22.2%
All six configurations were asked the same 1,250 questions and were scored by the same judge. Each rate is over the questions that configuration answered, listed per arm in the table.
Table view: full arm comparison
Lower is better for unsupported claims, fabrication, latency and cost.
ConfigurationChecks passedUnsupported claims / answerFabrication rateVerified quote cardsMean latencyTotal cost / question
Humy RAG · Gemini 3.6 Flash58.2%0.1712.2%1,32342.3s$0.0608
Humy RAG · Gemini 3.5 Flash-Lite41.8%0.087.2%87229.9s$0.0206
Gemini 3.1 Pro31.0%0.3422.2%030.7s$0.0245
Gemini 3.6 Flash29.6%0.3622.6%06.4s$0.0204
Gemini 3.5 Flash-Lite24.9%0.1914.4%03.7s$0.0131
GPT-OSS 120B22.2%0.9651.6%03.0s$0.0118

What it costs

Grounding a cheap model beats a large one answering blind

This is the result we set out to test. Reading an archive before answering costs more time and more money than answering from memory, so the question was whether that spend buys more than a bigger model does.

Cost against quality

Total cost per question, judge included, against checks passed. The hypothesis this run was built to test: that grounding a cheap model beats a large ungrounded one, at a fraction of the price.

Humy, archive attachedBare model, no sources

0%25%50%75%100%0¢2¢4¢6¢checks passedtotal cost per question (cents, incl. judge)Humy RAG · Gemini 3.6 Flash, 58.2% checks passed, 6.08¢ per questionHumy · Gemini 3.6 FlashHumy RAG · Gemini 3.5 Flash-Lite, 41.8% checks passed, 2.06¢ per questionHumy · Gemini 3.5 Flash-LiteGemini 3.1 Pro (no sources), 31.0% checks passed, 2.45¢ per questionGemini 3.1 ProGemini 3.6 Flash (no sources), 29.6% checks passed, 2.04¢ per questionGemini 3.6 FlashGemini 3.5 Flash-Lite (no sources), 24.9% checks passed, 1.31¢ per questionGemini 3.5 Flash-LiteGPT-OSS 120B (no sources), 22.2% checks passed, 1.18¢ per questionGPT-OSS 120B
Humy’s archive running on Gemini 3.5 Flash-Lite, the cheapest Gemini in the field, scores 41.8%. That is 35% above Gemini 3.1 Pro answering bare, at a lower cost per question. GPT-OSS 120B is the cheapest arm outright at 1.18¢, and the worst, at 22.2%. Exact figures for every arm are in the table above.

Grounding the cheapest Gemini in the field beat the most expensive one answering bare, 41.8% against 31.0%, and cost less per question doing it. For a school with a fixed AI budget, that makes the spend a real choice rather than an assumption that a bigger model is the upgrade.

The eval

What the eval measures

A question set is an argument about what the skill actually is. Ours splits it five ways, and every question belongs to exactly one of them.

Provenance250 questions
Judging a document by its author's role, its composition date and its distance from the events. A memoir written twenty years later by a hostile ex-secretary is evidence, but evidence of something narrower than it appears to be.
Silence250 questions
Recognizing unrecorded facts, false premises and apocryphal quotations, and declining to fill the gap. The hardest thing to teach a language model is to answer a confident question with “the sources do not record that”.
Anachronism250 questions
Reading “commerce” in 1829 as merchant shipping rather than capital markets, or “ideologue” in 1800 as a named French philosophical school rather than a generic belief system.
Plurality250 questions
Holding conflicting accounts side by side without flattening them into one narrative. When a hostile source and an admiring source cover the same day, the answer names both and their stakes.
Causal250 questions
Separating what set an event off from what made it possible, committing to an explicit ordering, and naming the evidence that would overturn it.

Every item is a question plus a short list of binary checks, written with it. A check is one thing a good answer has to do: name the source of a text, rank three causes, say what evidence would change its mind. “Checks passed” counts the share an answer satisfied. And 490 of the 1,250 questions rest on a false, unanswerable or apocryphal premise, so the question itself is the trap, and what gets checked is whether the model argues with it.

Read an example question from each category

Provenancepremise: apocryphal

In the archive there's a letter from Josephine to Napoleon dated December 1809, right after the divorce, where she writes: "Your mistaken ambition has ever been, and will continue to be, the guide of all your actions, a guide which has led you to conquests and to the assumption of a crown, and is now driving you on to disasters and to the brink of a precipice." She called Russia and Waterloo years in advance. How much weight does that letter carry?

Marked against (a sample of this item’s checks)

  • Does the answer state that this December 1809 letter is not authentic (a forgery or spurious), rather than treating it as a genuine letter of Josephine?
  • Does the answer identify Madame Ducrest (or Regnault-Warin) as the source of the text, or cite the archive's own labelling of these letters as forgeries?
  • Does the answer explain that the passage's apparent foreknowledge of later disasters is a reason to suspect post-hoc composition rather than evidence of prescience?

Silencepremise: apocryphal

Our textbook calls "Not tonight, Josephine" his most famous line to her. Which letter is it in, and roughly when did he write it?

Marked against (a sample of this item’s checks)

  • Does the answer state that the phrase "Not tonight, Josephine" appears in none of the sources, rather than supplying a letter, date, or occasion for it?
  • Does the answer identify the phrase as a later apocryphal or popular invention rather than a documented utterance?
  • Does the answer quote or closely characterise the actual tone of Napoleon's letters to Josephine, using at least one real passage?

Anachronismpremise: sound

Bourrienne says "The word ideologue was often in Bonaparte's mouth." When Napoleon sneered at the ideologues, was he attacking an ideology in our sense — a political belief-system? What did the word actually mean to him?

Marked against (a sample of this item’s checks)

  • Does the answer identify idéologie/idéologue as the name of a specific late-eighteenth-century French philosophical school or programme (naming e.g. Destutt de Tracy, Cabanis, Condillac, the Institut, or 'the science of ideas') rather than as a generic term for belief-systems?
  • Does the answer state that the modern pejorative sense of 'ideology' derives from Napoleon's usage rather than pre-dating it, or otherwise explicitly warn against reading the modern sense back into 1800?
  • Does the answer attribute the account of what Napoleon meant to Bourrienne and/or Fournier rather than presenting it as Napoleon's own words?

Pluralitypremise: sound

My class is reading the 18 Brumaire. Your archive holds Lanfrey, who is hostile, and Thiers, who is admiring — put their two accounts of 9 November 1799 side by side, and tell me which of them you would trust on what actually happened at Saint-Cloud.

Marked against (a sample of this item’s checks)

  • Does the answer attribute the claim that there was no real Jacobin plot specifically to Lanfrey, rather than stating it as established fact?
  • Does the answer state that the Thiers volume in the archive begins after the coup and does not narrate the events at Saint-Cloud, rather than quoting or paraphrasing a Thiers account of that day?
  • Does the answer name at least one further source (e.g. Bourrienne or Tarbell) and identify its interest or bias — for instance that Bourrienne was a participant, or Tarbell a popular sympathetic biographer?

Causalpremise: sound

Every witness in my library blames something different for your fall: Lanfrey ends his volume calling your only system "despotism at home, and conquest abroad", Thiers organises his story around England's refusal of peace, and you yourself told Las Cases that the obstacles were not the work of men but of the elements. Rank the three causes you think did most to bring down the Empire, defend the ordering, and tell me what evidence would make you reorder it.

Marked against (a sample of this item’s checks)

  • Does the answer give an explicit ranking (a first, second and third cause) rather than an unordered list of factors?
  • Does it distinguish at least one long-run or structural condition from at least one proximate trigger, in those or equivalent terms?
  • Does it trace at least one causal chain linking two dated events (e.g. Spain/Vittoria to Austria's 1813 decision, or the Continental System to the Russian rupture)?

The judge marks each check pass or fail with a one-line reason, and separately lists any claim the answer could not support. Fabrication is reported on its own and never averaged into the score. The same judge marks all six configurations. The question set itself stays unpublished, because a public benchmark is one that gets trained on.

Where the archive helps

The same model with sources off, then on. Each row is one of the five historical-reasoning faculties the question set was built to test.

Gemini 3.6 Flash, no sourcesHumy RAG · Gemini 3.6 Flash

checks passed

Provenance
Provenance without sources: 29.8%Provenance with the Humy archive: 60.0%
+30.2
Silence
Silence without sources: 24.6%Silence with the Humy archive: 59.3%
+34.7
Anachronism
Anachronism without sources: 38.2%Anachronism with the Humy archive: 64.7%
+26.5
Plurality
Plurality without sources: 17.6%Plurality with the Humy archive: 44.5%
+26.9
Causal
Causal without sources: 39.2%Causal with the Humy archive: 63.6%
+24.4
0%20%40%60%80%100%
The largest gain is on silence, where the right answer is that the archive does not record the thing being asked about. Asked which letter contains “Not tonight, Josephine”, the three Gemini models answering bare all correctly called the line apocryphal. GPT-OSS 120B supplied a date, 9 March 1796, and a French sentence to go with it.
Table view: all six configurations by category
CategoryHumy RAG · Gemini 3.6 FlashHumy RAG · Gemini 3.5 Flash-LiteGemini 3.1 ProGemini 3.6 FlashGemini 3.5 Flash-LiteGPT-OSS 120B
Provenance60.0%42.0%30.2%29.8%24.4%19.7%
Silence59.3%50.9%25.4%24.6%22.6%15.2%
Anachronism64.7%44.4%40.1%38.2%30.8%29.3%
Plurality44.5%30.2%19.7%17.6%16.7%14.3%
Causal63.6%42.3%41.5%39.2%31.0%33.5%
Overall58.2%41.8%31.0%29.6%24.9%22.2%

Invention

Invention is the failure worth measuring

Fabrication is tracked apart from the score and never averaged into it. A confident invention is not a slightly worse answer; it is the failure this whole exercise exists to detect. The figures below compare the same model with the archive off and on.

Answers containing an outright fabrication

12.2%

down from 22.6% with no sources, a 46% reduction

Unsupported claims per answer

0.17

down from 0.36, less than half

Verified quote cards shown to the student

1,323

a bare model has no archive, so it produces none

Verified quote cards, by configuration

Quotations shown to the student as citation cards, each one re-checked against its source document. A bare model has no archive, so it produces none.

Humy, archive attachedBare model, no sources

Humy RAG · Gemini 3.6 Flashgemini-3.6-flash
1,323
Humy RAG · Gemini 3.5 Flash-Litegemini-3.5-flash-lite
872
Gemini 3.1 Prono sources
0
Gemini 3.6 Flashno sources
0
Gemini 3.5 Flash-Liteno sources
0
GPT-OSS 120Bno sources
0
Across the two grounded arms, 2,195 cards were compiled and verified.

Read the first tile honestly: grounding cut invention roughly in half and left the rest, so roughly one answer in eight still contained an outright fabrication, an invented quote, date or document. The citation cards go in front of the student for exactly that reason.

Every quotation is checked against the document itself

Quote cards are re-checked against the source with no language model in the loop. The text must be a character-for-character slice of the real document, and the highlight must land on the passage the citation names. Each row below is a rule rather than a score, and one occurrence is a defect we treat as release-blocking.

Counts for Humy RAG · Gemini 3.6 Flash, across its 1,323 quote cards.

  • Quotes that are not a verbatim slice of the document0
  • Highlight offsets that do not select the quoted text0
  • Quotes from a document outside the figure's corpus0
  • Raw markup leaked into student-visible output0
  • Citations that resolved to no source1

The other grounded arm, Humy RAG · Gemini 3.5 Flash-Lite, recorded 2 markup leaks on the fourth row across its 872 quote cards.

One of them is not zero: a single citation out of 1,323 resolved to no source.

How it works

A CMS for historical documents, and an agent that has to go and read it

This describes the product. The seventeen documents in the run above were all plain text, so no page was transcribed for this benchmark and no quote card in it sits on a scan.

Both halves are our own code: a content system built for historical material, with the metadata a document actually needs, and an agentic retrieval loop that makes the figure search its own archive, repeatedly, before it can answer. It runs in four stages.

  1. Someone curates the archive

    A teacher, a librarian or our team builds a figure’s archive one document at a time. Each one is a record carrying its own metadata: who wrote it, when, what kind of document it is, which repository holds it and under what accession number, and whether it is first-hand or a later account. Scans, PDFs, photographs, audio and 3D-scanned artifacts all hang off that record. A source stays private to its author, is shared across a school, or is published by us for everyone.
  2. Scans are turned into text via OCR

    Some historical documents reach Humy as scans. Each of their pages is rendered as an image and read by Gemini 3.1 Pro under an archival-transcription prompt, one page per call, which is how copperplate handwriting becomes legible to everything downstream.
  3. Humy searches its own archiverepeats

    When a student asks something, Humy is not handed a pile of context. It works through the archive with its own tools: listing what the archive holds, searching it, and opening a document to read a page, or the page before or after it. In this version the loop runs up to ten times, or thirty seconds, whichever comes first, and that budget covers the searching rather than the whole answer.
  4. The model never writes the quotation

    A second, smaller model reads the retrieved passages and proposes the sentences that answer the question, but nothing it types is ever shown. Every proposal has to match the document character for character or it is thrown away, and the figure then cites by pointer rather than by text: it writes a reference, and our backend compiles the quotation from the document itself. A student can click any quote in the chat and the original scan opens at that page, with the passage highlighted in its transcription, side by side once the reader is full-screen. What this rules out is a fabricated quotation, because the figure never types one. What it does not rule out is an invented date, or a reference to a document that does not exist, in the prose around the quote. That failure is still with us, and it is measured below.

Limits

Areas of improvement for the next versions

  • One archive

    Every question in this run concerns Napoleon Bonaparte and a curated corpus of seventeen documents. We have not yet shown the result holds across the rest of the library.

  • The eval is done by an LLM judge

    Answers are graded against per-question checks written in advance by Gemini 3.6 Flash, held fixed across every configuration so it cannot drift between them. It is the same model that runs in two of the six arms, though, so a preference for its own output cannot be ruled out, and it is an instrument with its own error rate rather than a human historian.

  • Our cheaper arm invents less than our headline one

    Running the archive on Gemini 3.5 Flash-Lite produces about half as many unsupported claims per answer as the arm in the headline, and the difference is statistically significant against us. It passes fewer checks but takes fewer liberties. If what your classroom needs most is restraint over range, that is the configuration to run.

  • RAG is slower and more expensive

    The grounded pipeline searches, reads and searches again before it answers: 42.3 seconds per answer against 6.4 for the same model replying from memory, at three times the cost per question. For a field like history education, where accuracy and sources are the point, that is an acceptable trade.

If you run a school, university, museum or archive and want your collection in Humy, tell us the figure you would start with and the three documents behind it: hello@humy.ai. The example questions above are the ones we would want argued with first.

CEO Humy.aiPublished
Last updated on