foliodoc v0.1 · Python 3.10+

Documents in, structure out.

foliodoc converts PDFs, scans, images and Word files into Markdown, JSON or plain text with headings, lists, tables and reading order intact. It reads the exact text a PDF already contains, runs OCR only on pages that need it, and rebuilds the layout geometrically. No GPU, no layout model.

Windows · Linux · macOS 0.08 s per digital page OCR on-device Pure Python API
terminal
pip install "foliodoc[ocr]"
title
paragraph
heading L1
list_item ×2
table 4×3
paragraph
Acme Report · Internal · p. 4
furniture, dropped
# Quarterly Operations Review

Revenue grew 12.5% to $1.23M in Q3,
driven by the new Pune warehouse.

## 1 Logistics

- On-time delivery rose to 96.4%
- Returns fell to 1.8%

| Region | Orders | Late |
|---|---|---|
| North | 4,210 | 88 |
| West  | 3,905 | 61 |
| South | 2,774 | 40 |

What foliodoc sees on a page (left) and the Markdown it writes (right). The running footer is recognised and left out.

How it works

Four stages, every page

  1. Read the text layerpdfium gives every character with its box, font size and weight. Exact, and about 10 ms per page.
  2. OCR only where neededPages with no text layer or a broken font encoding are straightened and sent to OCR. Lines the engine skipped are re-read from tight crops.
  3. Rebuild the layoutReading order by XY-cut, tables from ruling lines or aligned columns, headings from size, weight and numbering, lists, hyphenation.
  4. ExportMarkdown, plain text, JSON, or the Python objects directly, with page numbers and bounding boxes.
Exact digital textNo OCR on PDFs that already carry text, so no recognition errors there.
TablesRuled grids, rules-only tables and borderless tables. Tables split across pages or columns are stitched back together.
Reading orderMulti-column pages read column by column. Tables and figures stay whole.
Headers and footersRunning headers, footers and page numbers are marked furniture and kept out of the text.
Rotated and spaced textRotated labels read in their own direction; letter-spaced headings keep their words.
Bounded memoryPages stream through a small worker pool, so a 1,000-page scan never holds more than a few page images.
Install

One package, an OCR engine per platform

Digital PDFs and Word files need nothing beyond the base install. Scans and images need an OCR engine: foliodoc[ocr] installs the right ones for your platform, and ocr="auto" uses the first one available, in the order below. The source is on GitHub.

EnginePlatformsInstallNotes
Apple VisionmacOSpip install "foliodoc[apple]"Neural Engine, fastest and most accurate on a Mac.
RapidOCRWindows, Linux, macOSpip install "foliodoc[rapidocr]"PaddleOCR v5 models on ONNX Runtime, CPU only. The portable default.
TesseractWindows, Linux, macOSpip install "foliodoc[tesseract]"Also needs the tesseract binary on your PATH.
terminal
pip install foliodoc              # digital PDFs and Word files only
pip install "foliodoc[ocr]"       # + OCR: Apple Vision on macOS, RapidOCR everywhere
Quickstart

Convert a file in three lines

python
from foliodoc import convert

doc = convert("report.pdf")      # or scan.png, photo.jpg, fax.tiff, contract.docx
print(doc.to_markdown())

A Document is plain Python objects, so you can also walk it directly:

python
for block in doc.blocks:
    if block.type == "heading":
        print("  " * (block.level - 1) + block.text)

for table in doc.tables:           # each table is a list of rows of cell strings
    header, *rows = table
    print(header, len(rows), "rows")

print(doc.timings)                 # {'total_s': 0.04, 'per_page_s': 0.02, ...}
API reference

Everything you import

convert(source, **options) → Document

Converts one file. source is a path or raw bytes (PDF, PNG, JPEG, TIFF including multi-page, BMP, GIF, WebP, DOCX). Options are the same as Converter's.

Converter(**options)

Reuse one converter for many files: the OCR engine loads once. Call .convert(source) on it.

OptionDefaultMeaning
ocr"auto""auto", "apple", "rapidocr", "tesseract", "none", or the name of an engine you registered.
dpi300Resolution for pages that need OCR. Coarse scans are upsampled to it; very large ones are reduced.
force_ocrFalseIgnore the PDF text layer and OCR every page.
trust_ocr_layerTrueAccept an existing invisible OCR layer in scanned PDFs instead of re-reading them.
languagesNoneOCR languages for Apple Vision, e.g. ["en-US", "zh-Hans"].
workersNoneConcurrent OCR pages. By default whatever the engine scales to.

Document

MemberReturns
to_markdown()Markdown with # headings, - lists and pipe tables.
to_text()Plain text in reading order; table rows become lines.
to_json(**kw) / to_dict()Every page and block with positions. kw goes to json.dumps.
blocksAll body blocks in reading order, headers and footers excluded.
tablesEach table as a list of rows of cell strings.
pagesPage objects: number, width, height, ocr (whether OCR was used), blocks.
timingsSeconds spent, total and per page.

Block

FieldMeaning
typetitle, heading, paragraph, list_item, table, figure or furniture
textThe block's text (empty for tables and figures).
levelHeading level, 1 = top. From font size, or from numbering such as 2.3.
cellsFor tables: rows × columns of strings.
bbox(x0, y0, x1, y1), top-left origin, in PDF points (pixels for image input).
page1-based page number.
sourcenative (read from the PDF) or ocr.
Command line

Batch conversion from a shell

terminal
foliodoc report.pdf                                 # Markdown to stdout
foliodoc *.pdf scans/*.png -o out/ --to md          # one out/<name>.md per file
foliodoc invoice.pdf --to json --timings            # JSON, and timings on stderr
foliodoc fax.tiff --ocr rapidocr --dpi 300
foliodoc old_scan.pdf --force-ocr                   # ignore a bad embedded text layer
Examples

Common tasks

Tables to pandas

python
import pandas as pd
from foliodoc import convert

doc = convert("annual_report.pdf")
frames = [pd.DataFrame(rows, columns=header) for header, *rows in doc.tables]

A folder of mixed files, one converter

python
from pathlib import Path
from foliodoc import Converter

conv = Converter(ocr="rapidocr")          # models load once
for path in Path("inbox").iterdir():
    doc = conv.convert(path)
    Path("out", path.stem + ".md").write_text(doc.to_markdown())

Chunks for search or RAG, with page citations

python
chunks, section = [], ""
for b in convert("handbook.pdf").blocks:
    if b.type in ("title", "heading"):
        section = b.text
    elif b.type in ("paragraph", "list_item"):
        chunks.append({"section": section, "page": b.page, "text": b.text})

Scans in other languages

python
doc = convert("brochure_zh.png", ocr="apple", languages=["zh-Hans", "en-US"])

Text only, never OCR

python
doc = convert("contract.pdf", ocr="none")   # fastest; scanned pages come back empty
if any(p.ocr for p in doc.pages): ...           # with OCR on, see which pages needed it

Plug in your own OCR engine

An engine is a function from a PIL image to text lines with pixel boxes. Register it once and select it by name.

python
from foliodoc import Segment, convert, register_ocr_engine

@register_ocr_engine("my-ocr")
def make_engine():
    client = MyOcrClient()                       # load models once
    def run(image):
        return [Segment(line.text, *line.box, size=line.height * 0.78, conf=line.score)
                for line in client.read(image)]
    return run

doc = convert("scan.png", ocr="my-ocr")
Benchmarks

7,822 public documents against docling

Four public datasets with ground truth: 4,999 DocLayNet digital PDF pages, 1,651 OmniDocBench page images, 973 SROIE receipts and 199 FUNSD forms. Both tools ran one after the other on an Apple M2 with default settings, which on macOS means Apple Vision OCR for both. Docling version 2.63.

Where foliodoc wins

  • Speed: ~20× on digital PDFs, 5–10× on scans.
  • Text of digital PDFs, forms and receipts.
  • Receipt totals found: 97% vs 90%.
  • Robustness: 1 failure vs 214.

Where docling wins

  • Layout labels on real documents: tables, headings, lists.
  • Table structure on real pages.
  • Text on book, paper and slide images (OmniDocBench).
  • Removing running headers and footers.
Dataset · metricfoliodocdoclingBetter
DocLayNet · text token F10.9390.931foliodoc
DocLayNet · seconds per page (mean)0.081.66foliodoc
DocLayNet · table / heading / list detection F10.35 / 0.41 / 0.360.86 / 0.86 / 0.87docling
DocLayNet · headers/footers leaking into text ↓32%7%docling
OmniDocBench English · text edit distance ↓ (703 pages both read)0.2980.229docling
OmniDocBench English · table TEDS0.3430.493docling
OmniDocBench · seconds per page0.997.94foliodoc
SROIE · token F1 (840 receipts both read)0.7020.640foliodoc
SROIE · receipt total found97.1%89.5%foliodoc
FUNSD · token F10.8990.856foliodoc
Failed documents (of 7,822)1214foliodoc

Windows and Linux configuration

Both tools with RapidOCR, the OCR they use off macOS. This ran on a subset (all FUNSD forms, the SROIE test split, 493 English and mixed-language OmniDocBench pages). DocLayNet needs no OCR, so its numbers above hold on every OS.

Dataset · metricfoliodocdoclingBetter
FUNSD · token F10.8100.844docling
FUNSD · seconds per form6.714.18docling
SROIE · token F10.6010.620docling
SROIE · receipt total found97.4%93.4%foliodoc
SROIE · seconds per receipt4.665.71foliodoc
OmniDocBench · text edit distance ↓0.3240.296docling
OmniDocBench · table TEDS0.4610.697docling
OmniDocBench · seconds per page8.5611.81foliodoc
  • Every system is scored on exactly the same documents. A document a tool failed on counts as empty output, except in rows marked "both read".
  • Docling's layout model was trained on DocLayNet. 102 DocLayNet pages whose own ground-truth text is garbled are excluded for both tools.
  • Docling's 214 failures are large images it rejects after internal upscaling (Pillow's 179-megapixel safety limit).
  • Text edit distance is normalized Levenshtein distance of the page text in reading order (0 is perfect). TEDS is the standard tree-edit similarity for tables (1 is perfect).
  • Reproduce with the scripts in bench/: real_prep.py, real_run.py, real_score.py.
Limits

Know before you rely on it

  • Layout on complex real documents is foliodoc's weak point. If you need precise table structure or section labels from reports and papers, docling's trained models do better today.
  • OCR quality depends on the engine: best with Apple Vision on macOS; RapidOCR on Windows and Linux runs at roughly 3–8 s per scanned page on a laptop CPU.
  • Formulas are returned as plain text, not LaTeX. Figures are marked but not described. Handwriting is only as good as the OCR engine.
  • Non-Latin scripts need the OCR language set explicitly (languages=[...] with Apple Vision).