When PyMuPDF Can’t See the Table: Parse PDFs for RAG with Azure Layout

Enterprise Document Intelligence [Vol.1 #5bis] – The same relational tables. Native table cells. OCR for scanned pages and images. Captions and headings without regex.

Kezhan Shi
Jun 12, 2026
16 min read

by Amsterdam City Archives, via Unsplash.

This article is a parsing companion in Enterprise Document Intelligence, the series that builds an enterprise RAG system from four bricks. Article 5 (document parsing) built the parser with PyMuPDF (fitz). This companion keeps the same goal and the same relational tables, and swaps the engine for Azure Layout (the prebuilt-layout model), a richer package that recovers what fitz cannot. That gap is where we start.

1. Where fitz is blind

Four cases. In each one, fitz misses and Azure works.

1.1. Tables: fitz returns flat words, Azure returns cells

A contract table has rows and columns. The label “Renewal fee” sits in column 1, the value 500 sits in column 2. Fitz reads the page top to bottom and emits one line per text segment. The four cells of a row come back as four loose words. Sometimes the cells from the row below get mixed in if the y-coordinates are close. The chunker downstream sees a soup of words. The row-and-column structure that makes a table a table is gone.

Azure’s prebuilt-layout model detects each table as a structured object. result.tables is a list of tables, each with cells indexed by (row_index, column_index). The header row is flagged (cell.kind == "columnHeader"). The cell content is the cell text, exactly as the author typed it. We flatten the table into markdown rows so it lives inside line_df like any other content.

1.2. Images: fitz returns the bbox, Azure returns the text

Many PDFs have figures with text inside them. Architecture diagrams with box labels. Charts with axis ticks and legends. Signed seal stamps. Embedded screenshots of spreadsheets. Fitz returns each image as a bbox and the raw bytes. The text inside is invisible to the parser.

Azure’s OCR runs on every page, including the pixels inside figure regions. For each figure, we collect every Azure word whose bbox sits inside the figure region and join them as ocr_text. “Multi-Head Attention Concat Linear h” now lives in image_df.ocr_text for the figure on page 4 of the Attention paper.

1.3. Scanned pages: fitz returns nothing, Azure returns OCR

A 30-page native contract gets a 10-page scanned amendment glued at the end. Fitz reads the native pages and returns empty strings for the scanned ones. Azure runs OCR on every page regardless of source. Native pages and scanned pages come back through the same result.pages[i].lines path with the same shape.

1.4. Captions and headings: fitz uses regex, Azure has explicit roles

Azure’s paragraphs field has role labels: each paragraph in the result carries a tag like "figureCaption", "tableCaption", "title", or "sectionHeading" that tells us what kind of block it is, without any regex.

2. Same contract, richer data

One Azure call shared by every builder. The prebuilt-layout model returns native table cells (rows, columns, headers), OCR text for every page (native or scanned), figures with the text inside them, and paragraph roles (title, sectionHeading, figureCaption, tableCaption).

from azure.ai.documentintelligence import DocumentIntelligenceClient
from azure.ai.documentintelligence.models import AnalyzeDocumentRequest
from azure.core.credentials import AzureKeyCredential

client = DocumentIntelligenceClient(endpoint, AzureKeyCredential(key))

with open("contract.pdf", "rb") as f:
    poller = client.begin_analyze_document(
        "prebuilt-layout",
        AnalyzeDocumentRequest(bytes_source=f.read()),
    )

result = poller.result()   # tables, paragraph roles, OCR, reading order

3. What each table gains

3.1. line_df gains table-cell rows, image OCR, selection marks

3.2. image_df gains an ocr_text column

3.3. toc_df gets reconstructed from paragraph roles

3.4. object_registry gets caption-role detection

3.5. parsing_summary gains Azure-specific stats

3.6. page_df and cross_ref_df: unchanged

4. The parsing_method column: provenance for adaptive parsing

5. Cost and latency

Latency

Azure is not free. The order of magnitude is stable: fitz is free, Azure costs roughly a cent per page. The exact tier prices change with region and time: treat the numbers above as a calibration, not a contract.

6. When to call which

Default to fitz. Escalate to Azure when a specific signal says fitz is not enough.

7. Conclusion

Two engines, one contract: the same relational tables out, same downstream code regardless of which one ran.