Which formats go in, and how each citation is anchored
Digital and scanned PDFs, photos, EPUB, Word, slides, spreadsheets, audio, video, web pages, YouTube and podcasts. For each, how it is read and which anchor it leaves: printed folio, second, slide, rows or paragraph.
Reviewed on This page as Markdown
How a file goes in
Everything first goes through the imprenta (the press), Scholaris's converter. It recognises the type from the first bytes of the file (then from its MIME type and, last, its extension) and converts it in your own browser whenever it can. What the browser cannot do (decode some images, extract video frames, convert a Keynote deck) is done by a server container with ffmpeg and LibreOffice.
Then whatever has to be read with eyes (scanned pages, photos, slides) is read by a vision model, four pages per request, and whatever has to be heard is transcribed by a speech model. Which provider does what is in Privacy.
Formats and anchors
| What you upload | Extensions | How it is read | Anchor and how it is cited |
|---|---|---|---|
| Digital PDF | The text layer, with its lines and blocks; headers and footers set apart; the PDF's own page labels, outline and metadata are used | Printed page: "p. 23", "pp. 23-24", "p. xiv" | |
| Scanned PDF, or one with old OCR | Vision: each page as an image. If the PDF carries a poor old OCR layer, it is read again | Printed page, or the physical one in brackets: "p. [12]" | |
| Photos of a book | jpg, png, webp, gif, tiff, heic, bmp, avif | Sorted by name in numeric order (IMG_2 before IMG_10), straightened and cropped | Printed page of each photo |
| Text documents | docx, odt, rtf, html, md, txt | Blocks (headings, paragraphs, lists, quotations, tables, notes) with their heading path | Section and paragraph: "Introduction, para. 4" |
| EPUB | epub | Reading order, table of contents and, above all, the print page list when the EPUB has one | "p. 23" with a page list; otherwise "para. N" |
| Slides | pptx, odp, key | Text and speaker notes of every slide; ODP and Keynote are converted whole with LibreOffice | "slide 7" |
| Spreadsheets | xlsx, xls, ods, csv, tsv | Tables in chunks of 50 rows, with the header repeated in each chunk | "Sheet1, rows 2-51" |
| Audio | mp3, wav, m4a, ogg, opus, flac, aac | Word-by-word transcription with who is speaking; citable stretches of 30 to 60 seconds that end at a sentence boundary | "12:04", "1:02:03", "12:04-12:40" |
| Video | mp4, mov, webm, mkv, avi | Audio as above; plus frames at every scene change (or every 20 seconds), which are searchable too | Same as audio |
| A web page | URL | A dated copy is stored and the article is split into blocks | Section and paragraph, with the access date |
| YouTube | URL | Not downloaded: Gemini watches the video from its address, ten minutes at a time | The second |
| Vimeo, podcasts, audio links | URL | Vimeo, the smallest progressive file; a podcast, the audio in its RSS | The second |
| SPDF and .scholaris packages | spdf, scholaris | Nothing is read again: text, anchors and vectors are copied | The anchors they carried |
The printed folio
The number that matters in a footnote is the one printed on paper, not the page's position in the file. A book may start in roman numerals, skip unnumbered plates or have two pages per image. For each page Scholaris gathers several pieces of evidence:
- the PDF's own page labels, if any;
- candidate numbers from the header, the footer and the text edges, plus what the vision reader saw;
- the longest coherent sequence of those readings;
- for doubtful cases, a judge (Jev, by TypeSafe) choosing between candidates;
- and, for pages with no visible number, interpolation within each stretch.
It recognises roman numerals (and the page where arabic numbering starts), foliation by leaves ("23r", "23v"), two-up scans, unnumbered plates and years that are not folios. Every page records where its number came from (read, inferred, from the EPUB, or none) and how confident it is. If a book prints no number at all, none is invented: the physical position is cited in brackets.
In the benchmark, all 64 folios checked by eye came out exact (see Performance).
Original spelling is kept
The reader transcribes faithfully: it does not modernise "dixo", "assi" or "muger", keeps u/v, i/j/y and ç, does not fix misprints and keeps running heads, catchwords and signatures out of the body. The one exception is the long s (ſ), written as "s". So that those spellings can still be found, there is a separate layer, explained in Search.
Size limits
| Plan | Largest file | Storage |
|---|---|---|
| Free | 200 MB | 1 GB |
| Pro | 4 GB | 100 GB |
| Home version | 16 GB | your disk |
Through API v1 in the cloud a single request takes files up to 95 MB; for more, the app and the Python SDK upload in parts. There is no maximum duration for audio and video beyond the file size.
What does not work well yet
- Old Word .doc files have no reader of their own: convert them to .docx first.
- Vimeo only works when the video offers a downloadable file; otherwise, download it and upload it.
- Web anchors are paragraph anchors: if the page changes later, the dated copy we stored is the one that counts.