Skip to the text
Scholaris

What we measured, and under what conditions

Reading times and costs, folio accuracy, search quality, invented citations and latency, as they came out of the benchmark of 6 October 2026, and what those figures do not say.

Reviewed on This page as Markdown

Conditions

Everything was measured on 6 October 2026, from a Mac in Tenerife over a home connection, against the cloud APIs (Gemini 3.5 Flash-Lite and 3.8 Flash, Gemini Transcribe, Gemini Embedding 2 at 1536 dimensions, Jev, Crossref and OpenAlex) with no local GPU. It was not measured on Cloudflare, which should be more stable; nor was the latency of "ask". Dollar figures are provider costs, not prices.

Reading a document

DocumentBefore (first Scholaris)NowCost
Video interview, 54 min3 h 39 min45 s (30 s without speaker attribution)$0.41
Video interview (Cortázar), 2 h 2 minthe 83-min ones took 5 h 43 min70 s$0.94
The Discarded Image, 245 pp., with old OCRnot recordedmedian 85 s over 16 runs (37 to 175 s)$0.64
El casamiento en la muerte, 43 scanned pp., 17th centurynot recorded27 s$0.20
Attention Is All You Need, 15 digital pp.not recorded12 s$0.04
El perseguidor, 37 digital pp.not recorded15 s$0.04

Times are end to end, conversion included. The first pages are searchable earlier: in under a tenth of a second for a digital PDF, in about 3 to 6 seconds for a scan and in about 3 to 6 seconds for a video. The slow runs of the 245-page book came from the API or the network, not the code. Very short media (one minute) now take a little longer to be ready than before: from 5 to 6 seconds it went to between 8 and 10.

Economy mode cut the cost of the Casamiento from $0.206 to $0.111 (46 % less), in exchange for taking 6.9 minutes instead of 19.7 seconds. Transcription costs about $0.005 a minute.

Reading well

  • Character error rate (CER) on seventeenth-century drama: 0.007 with Gemini 3.8 Flash, against 0.073 for the first Scholaris's reader. Measured on a single hand-transcribed page.
  • The Discarded Image: the first Scholaris left 195 pages empty (74,000 characters in all); now there are 358,000 characters and 7 empty pages.
  • Printed folio: 64 of 64 exact on the pages checked by eye (45 of 45 with the number visible). In The Discarded Image, 229 folios read agree with the PDF's labels on 229 of 232 pages.
  • Speakers: attribution right in 20 of 21 turns and in 23 of 23 in two interviews (25 random turns each, annotated by reading the text).
  • Records: 55 of 55 checks right across 9 documents after checking against external sources (30 of 55 before).

Searching well

Over 183 queries and 8,879 relevance judgments. The judgments are "silver": two models and an arbiter made them, not people; they agreed 82 % exactly and 99 % within one point, and of 30 checked by hand, the author agreed with 29.

SystemnDCG@10Recall@20MRR
First Scholaris0.6950.6180.925
Lexical only0.564
Visual only0.199
Semantic only0.8040.8600.931
Hybrid, not reranked0.8180.8840.942
Current (hybrid, reranked)0.8890.9040.979

By query type (nDCG@10 of the current system): cross-language 0.897; literal 0.880; audio and video 0.882; old spelling 0.814; visual 0.950 (only 4 queries).

Old spelling: on Lope de Vega's El casamiento en la muerte, with 30 queries, average recall per query went from 32 % to 100 % with the modernised-spelling layer; on modern documents, 250 of 250 searches returned exactly the same.

Latency: median of about 650 ms from home, no caches (about 360 ms for the query vector and about 300 ms for the judge); without reranking, about 350 ms with nDCG 0.818. In the measured runs, median and 95th percentile were 687 and 856 ms.

Citing well

Over 41 claims (30 supported and 11 not): 95.1 % precision, 96.7 % recall, zero invented citations, zero citations on unsupported claims and the verification verdict right 100 % of the time. Each claim took 998 ms and cost $0.0054.

What these figures do not say

  • They come from a home connection, not Cloudflare, and do not include the latency of "ask".
  • 1,594 relevance judgments with an arbiter or a disagreement are still awaiting human review.
  • The CER was computed on a single page.
  • The benchmark is small and was made by the person who makes Scholaris. If you want to repeat it with your documents, write to us.