Skip to the text
Scholaris

Which formats go in, and how each citation is anchored

Digital and scanned PDFs, photos, EPUB, Word, slides, spreadsheets, audio, video, web pages, YouTube and podcasts. For each, how it is read and which anchor it leaves: printed folio, second, slide, rows or paragraph.

Reviewed on This page as Markdown

How a file goes in

Everything first goes through the imprenta (the press), Scholaris's converter. It recognises the type from the first bytes of the file (then from its MIME type and, last, its extension) and converts it in your own browser whenever it can. What the browser cannot do (decode some images, extract video frames, convert a Keynote deck) is done by a server container with ffmpeg and LibreOffice.

Then whatever has to be read with eyes (scanned pages, photos, slides) is read by a vision model, four pages per request, and whatever has to be heard is transcribed by a speech model. Which provider does what is in Privacy.

Formats and anchors

What you uploadExtensionsHow it is readAnchor and how it is cited
Digital PDFpdfThe text layer, with its lines and blocks; headers and footers set apart; the PDF's own page labels, outline and metadata are usedPrinted page: "p. 23", "pp. 23-24", "p. xiv"
Scanned PDF, or one with old OCRpdfVision: each page as an image. If the PDF carries a poor old OCR layer, it is read againPrinted page, or the physical one in brackets: "p. [12]"
Photos of a bookjpg, png, webp, gif, tiff, heic, bmp, avifSorted by name in numeric order (IMG_2 before IMG_10), straightened and croppedPrinted page of each photo
Text documentsdocx, odt, rtf, html, md, txtBlocks (headings, paragraphs, lists, quotations, tables, notes) with their heading pathSection and paragraph: "Introduction, para. 4"
EPUBepubReading order, table of contents and, above all, the print page list when the EPUB has one"p. 23" with a page list; otherwise "para. N"
Slidespptx, odp, keyText and speaker notes of every slide; ODP and Keynote are converted whole with LibreOffice"slide 7"
Spreadsheetsxlsx, xls, ods, csv, tsvTables in chunks of 50 rows, with the header repeated in each chunk"Sheet1, rows 2-51"
Audiomp3, wav, m4a, ogg, opus, flac, aacWord-by-word transcription with who is speaking; citable stretches of 30 to 60 seconds that end at a sentence boundary"12:04", "1:02:03", "12:04-12:40"
Videomp4, mov, webm, mkv, aviAudio as above; plus frames at every scene change (or every 20 seconds), which are searchable tooSame as audio
A web pageURLA dated copy is stored and the article is split into blocksSection and paragraph, with the access date
YouTubeURLNot downloaded: Gemini watches the video from its address, ten minutes at a timeThe second
Vimeo, podcasts, audio linksURLVimeo, the smallest progressive file; a podcast, the audio in its RSSThe second
SPDF and .scholaris packagesspdf, scholarisNothing is read again: text, anchors and vectors are copiedThe anchors they carried

The printed folio

The number that matters in a footnote is the one printed on paper, not the page's position in the file. A book may start in roman numerals, skip unnumbered plates or have two pages per image. For each page Scholaris gathers several pieces of evidence:

  1. the PDF's own page labels, if any;
  2. candidate numbers from the header, the footer and the text edges, plus what the vision reader saw;
  3. the longest coherent sequence of those readings;
  4. for doubtful cases, a judge (Jev, by TypeSafe) choosing between candidates;
  5. and, for pages with no visible number, interpolation within each stretch.

It recognises roman numerals (and the page where arabic numbering starts), foliation by leaves ("23r", "23v"), two-up scans, unnumbered plates and years that are not folios. Every page records where its number came from (read, inferred, from the EPUB, or none) and how confident it is. If a book prints no number at all, none is invented: the physical position is cited in brackets.

In the benchmark, all 64 folios checked by eye came out exact (see Performance).

Original spelling is kept

The reader transcribes faithfully: it does not modernise "dixo", "assi" or "muger", keeps u/v, i/j/y and ç, does not fix misprints and keeps running heads, catchwords and signatures out of the body. The one exception is the long s (ſ), written as "s". So that those spellings can still be found, there is a separate layer, explained in Search.

Size limits

PlanLargest fileStorage
Free200 MB1 GB
Pro4 GB100 GB
Home version16 GByour disk

Through API v1 in the cloud a single request takes files up to 95 MB; for more, the app and the Python SDK upload in parts. There is no maximum duration for audio and video beyond the file size.

What does not work well yet

  • Old Word .doc files have no reader of their own: convert them to .docx first.
  • Vimeo only works when the video offers a downloadable file; otherwise, download it and upload it.
  • Web anchors are paragraph anchors: if the page changes later, the dated copy we stored is the one that counts.