Four stages, every page
- Read the text layerpdfium gives every character with its box, font size and weight. Exact, and about 10 ms per page.
- OCR only where neededPages with no text layer or a broken font encoding are straightened and sent to OCR. Lines the engine skipped are re-read from tight crops.
- Rebuild the layoutReading order by XY-cut, tables from ruling lines or aligned columns, headings from size, weight and numbering, lists, hyphenation.
- ExportMarkdown, plain text, JSON, or the Python objects directly, with page numbers and bounding boxes.
furniture and kept out of the text.One package, an OCR engine per platform
Digital PDFs and Word files need nothing beyond the base install. Scans and images need an OCR engine: foliodoc[ocr] installs the right ones for your platform, and ocr="auto" uses the first one available, in the order below. The source is on GitHub.
| Engine | Platforms | Install | Notes |
|---|---|---|---|
| Apple Vision | macOS | pip install "foliodoc[apple]" | Neural Engine, fastest and most accurate on a Mac. |
| RapidOCR | Windows, Linux, macOS | pip install "foliodoc[rapidocr]" | PaddleOCR v5 models on ONNX Runtime, CPU only. The portable default. |
| Tesseract | Windows, Linux, macOS | pip install "foliodoc[tesseract]" | Also needs the tesseract binary on your PATH. |
pip install foliodoc # digital PDFs and Word files only pip install "foliodoc[ocr]" # + OCR: Apple Vision on macOS, RapidOCR everywhere
Convert a file in three lines
from foliodoc import convert doc = convert("report.pdf") # or scan.png, photo.jpg, fax.tiff, contract.docx print(doc.to_markdown())
A Document is plain Python objects, so you can also walk it directly:
for block in doc.blocks: if block.type == "heading": print(" " * (block.level - 1) + block.text) for table in doc.tables: # each table is a list of rows of cell strings header, *rows = table print(header, len(rows), "rows") print(doc.timings) # {'total_s': 0.04, 'per_page_s': 0.02, ...}
Everything you import
convert(source, **options) → Document
Converts one file. source is a path or raw bytes (PDF, PNG, JPEG, TIFF including multi-page, BMP, GIF, WebP, DOCX). Options are the same as Converter's.
Converter(**options)
Reuse one converter for many files: the OCR engine loads once. Call .convert(source) on it.
| Option | Default | Meaning |
|---|---|---|
ocr | "auto" | "auto", "apple", "rapidocr", "tesseract", "none", or the name of an engine you registered. |
dpi | 300 | Resolution for pages that need OCR. Coarse scans are upsampled to it; very large ones are reduced. |
force_ocr | False | Ignore the PDF text layer and OCR every page. |
trust_ocr_layer | True | Accept an existing invisible OCR layer in scanned PDFs instead of re-reading them. |
languages | None | OCR languages for Apple Vision, e.g. ["en-US", "zh-Hans"]. |
workers | None | Concurrent OCR pages. By default whatever the engine scales to. |
Document
| Member | Returns |
|---|---|
to_markdown() | Markdown with # headings, - lists and pipe tables. |
to_text() | Plain text in reading order; table rows become lines. |
to_json(**kw) / to_dict() | Every page and block with positions. kw goes to json.dumps. |
blocks | All body blocks in reading order, headers and footers excluded. |
tables | Each table as a list of rows of cell strings. |
pages | Page objects: number, width, height, ocr (whether OCR was used), blocks. |
timings | Seconds spent, total and per page. |
Block
| Field | Meaning |
|---|---|
type | title, heading, paragraph, list_item, table, figure or furniture |
text | The block's text (empty for tables and figures). |
level | Heading level, 1 = top. From font size, or from numbering such as 2.3. |
cells | For tables: rows × columns of strings. |
bbox | (x0, y0, x1, y1), top-left origin, in PDF points (pixels for image input). |
page | 1-based page number. |
source | native (read from the PDF) or ocr. |
Batch conversion from a shell
foliodoc report.pdf # Markdown to stdout foliodoc *.pdf scans/*.png -o out/ --to md # one out/<name>.md per file foliodoc invoice.pdf --to json --timings # JSON, and timings on stderr foliodoc fax.tiff --ocr rapidocr --dpi 300 foliodoc old_scan.pdf --force-ocr # ignore a bad embedded text layer
Common tasks
Tables to pandas
import pandas as pd from foliodoc import convert doc = convert("annual_report.pdf") frames = [pd.DataFrame(rows, columns=header) for header, *rows in doc.tables]
A folder of mixed files, one converter
from pathlib import Path from foliodoc import Converter conv = Converter(ocr="rapidocr") # models load once for path in Path("inbox").iterdir(): doc = conv.convert(path) Path("out", path.stem + ".md").write_text(doc.to_markdown())
Chunks for search or RAG, with page citations
chunks, section = [], "" for b in convert("handbook.pdf").blocks: if b.type in ("title", "heading"): section = b.text elif b.type in ("paragraph", "list_item"): chunks.append({"section": section, "page": b.page, "text": b.text})
Scans in other languages
doc = convert("brochure_zh.png", ocr="apple", languages=["zh-Hans", "en-US"])
Text only, never OCR
doc = convert("contract.pdf", ocr="none") # fastest; scanned pages come back empty if any(p.ocr for p in doc.pages): ... # with OCR on, see which pages needed it
Plug in your own OCR engine
An engine is a function from a PIL image to text lines with pixel boxes. Register it once and select it by name.
from foliodoc import Segment, convert, register_ocr_engine @register_ocr_engine("my-ocr") def make_engine(): client = MyOcrClient() # load models once def run(image): return [Segment(line.text, *line.box, size=line.height * 0.78, conf=line.score) for line in client.read(image)] return run doc = convert("scan.png", ocr="my-ocr")
7,822 public documents against docling
Four public datasets with ground truth: 4,999 DocLayNet digital PDF pages, 1,651 OmniDocBench page images, 973 SROIE receipts and 199 FUNSD forms. Both tools ran one after the other on an Apple M2 with default settings, which on macOS means Apple Vision OCR for both. Docling version 2.63.
Where foliodoc wins
- Speed: ~20× on digital PDFs, 5–10× on scans.
- Text of digital PDFs, forms and receipts.
- Receipt totals found: 97% vs 90%.
- Robustness: 1 failure vs 214.
Where docling wins
- Layout labels on real documents: tables, headings, lists.
- Table structure on real pages.
- Text on book, paper and slide images (OmniDocBench).
- Removing running headers and footers.
| Dataset · metric | foliodoc | docling | Better |
|---|---|---|---|
| DocLayNet · text token F1 | 0.939 | 0.931 | foliodoc |
| DocLayNet · seconds per page (mean) | 0.08 | 1.66 | foliodoc |
| DocLayNet · table / heading / list detection F1 | 0.35 / 0.41 / 0.36 | 0.86 / 0.86 / 0.87 | docling |
| DocLayNet · headers/footers leaking into text ↓ | 32% | 7% | docling |
| OmniDocBench English · text edit distance ↓ (703 pages both read) | 0.298 | 0.229 | docling |
| OmniDocBench English · table TEDS | 0.343 | 0.493 | docling |
| OmniDocBench · seconds per page | 0.99 | 7.94 | foliodoc |
| SROIE · token F1 (840 receipts both read) | 0.702 | 0.640 | foliodoc |
| SROIE · receipt total found | 97.1% | 89.5% | foliodoc |
| FUNSD · token F1 | 0.899 | 0.856 | foliodoc |
| Failed documents (of 7,822) | 1 | 214 | foliodoc |
Windows and Linux configuration
Both tools with RapidOCR, the OCR they use off macOS. This ran on a subset (all FUNSD forms, the SROIE test split, 493 English and mixed-language OmniDocBench pages). DocLayNet needs no OCR, so its numbers above hold on every OS.
| Dataset · metric | foliodoc | docling | Better |
|---|---|---|---|
| FUNSD · token F1 | 0.810 | 0.844 | docling |
| FUNSD · seconds per form | 6.71 | 4.18 | docling |
| SROIE · token F1 | 0.601 | 0.620 | docling |
| SROIE · receipt total found | 97.4% | 93.4% | foliodoc |
| SROIE · seconds per receipt | 4.66 | 5.71 | foliodoc |
| OmniDocBench · text edit distance ↓ | 0.324 | 0.296 | docling |
| OmniDocBench · table TEDS | 0.461 | 0.697 | docling |
| OmniDocBench · seconds per page | 8.56 | 11.81 | foliodoc |
- Every system is scored on exactly the same documents. A document a tool failed on counts as empty output, except in rows marked "both read".
- Docling's layout model was trained on DocLayNet. 102 DocLayNet pages whose own ground-truth text is garbled are excluded for both tools.
- Docling's 214 failures are large images it rejects after internal upscaling (Pillow's 179-megapixel safety limit).
- Text edit distance is normalized Levenshtein distance of the page text in reading order (0 is perfect). TEDS is the standard tree-edit similarity for tables (1 is perfect).
- Reproduce with the scripts in
bench/:real_prep.py,real_run.py,real_score.py.
Know before you rely on it
- Layout on complex real documents is foliodoc's weak point. If you need precise table structure or section labels from reports and papers, docling's trained models do better today.
- OCR quality depends on the engine: best with Apple Vision on macOS; RapidOCR on Windows and Linux runs at roughly 3–8 s per scanned page on a laptop CPU.
- Formulas are returned as plain text, not LaTeX. Figures are marked but not described. Handwriting is only as good as the OCR engine.
- Non-Latin scripts need the OCR language set explicitly (
languages=[...]with Apple Vision).