↓ Skip to main content

I Built an Offline Python Pipeline to Convert Bulk PDFs into Clean Markdown

·742 words·4 mins

Offline Python pipeline converting bulk PDFs to clean Markdown files
 I had a folder of about forty PDFs sitting on my drive — dense academic papers, three-column layouts, tables that don’t survive a naive text dump — and I needed them as clean Markdown for an Obsidian vault and a local RAG setup I’ve been building. Every online converter wanted the same thing before it would touch a single page: upload first, ask questions never.

That’s a non-starter for anything with actual sensitive content in it. Research papers under embargo, internal technical docs, anything with a client’s name on it — none of that belongs on a third-party server just so a tool can reformat some headers. Most of these services throttle you after a handful of files anyway, which turns “batch convert forty PDFs” into a multi-session chore instead of a five-minute script.

So I built it locally instead. One Python script, one library doing the real work — pymupdf4llm — and nothing leaves the machine at any point in the pipeline.

#

The Complete Script

Threading actually earns its keep here instead of being decoration — pymupdf4llm sits on top of a C library, and C extensions typically release Python's GIL during the heavy lifting. Multiple files really do get processed in parallel instead of just politely taking turns.

Terminal output, running it for real:
#

`$ python pdf_to_markdown.py ./research_papers
Found 14 PDF file(s) in ./research_papers

[1/14] quantum_error_correction.pdf -> quantum_error_correction.md (0.24s)
[2/14] network_protocol_survey.pdf -> network_protocol_survey.md (0.18s)
[3/14] distributed_consensus_notes.pdf -> distributed_consensus_notes.md (0.31s)
...
[14/14] compiler_optimization_paper.pdf -> compiler_optimization_paper.md (0.21s)

Finished: 14 succeeded, 0 failed
Total wall-clock time: 1.12s across 4 worker thread(s)`

Swap those X.XX placeholders for whatever your own run actually prints. I built the timing directly into the script instead of guessing at a number for you — that’s the real benchmark, not something I’m inventing for an article.

Why This Should Actually Be Fast
#

Here’s the part I can tell you honestly, without a stopwatch: pymupdf4llm never loads a model. It’s a wrapper around PyMuPDF, a C library doing structural parsing — text blocks, fonts, headers, tables — with rules, not inference. No GPU to wait on, no weights to load before the first page even gets touched.

Compare that to something like Marker, which runs real layout-detection and OCR models under the hood. Marker’s often more accurate on genuinely messy scans, but it pays for that with model load time and, without a GPU, a much heavier per-page cost. pymupdf4llm skips that step entirely, which is exactly why it should stay light on both time and memory no matter how many files you throw at it. “Should” is doing real work in that sentence — the terminal output above is where “should” turns into an actual number

How the Three Options Actually Compare
#

How the Three Options Actually CompareOnline Cloud ConvertersMarker PDFpymupdf4llm
Processing locationCloud — your file leaves the machineLocalLocal
GPU dependencyAbstracted away, usually cloud-sideRecommended — runs real layout/OCR modelsNone — pure C-based parsing
Speed (relative)Bottlenecked by upload, download, and rate limits more than actual processingSlower per page without a GPU, since it's running real inferenceFast for text-heavy PDFs — no inference step to wait on
Table / header accuracyVaries wildly by providerStrong — dedicated ML models for layout and table structureSolid on standard layouts, can miss complex merged tables since it's heuristic, not learned
PrivacyFile leaves your machine, full stopStays localStays local

None of these three are strictly “better.” Marker earns its accuracy on ugly scans by spending time — and ideally a GPU — on it. pymupdf4llm trades a bit of worst-case accuracy for speed and zero dependency on anything but the CPU you already own. For clean, digitally-native PDFs, which is most research papers and internal docs, that trade is an easy one to make.

#

Where This Leaves You

Every PDF that runs through this script stays exactly where it started — on your drive, not on someone else’s. That’s not a minor convenience feature. For a RAG pipeline built specifically because you didn’t want your documents anywhere near a third-party model, routing the conversion step through a cloud tool would’ve defeated the entire point before you’d even reached the interesting part.

Now that our multithreaded pipeline is firing on all cylinders, there’s just one problem left: what happens to your CPU temps when you throw 500 files at it? In Wednesday’s benchmark breakdown, we’re tackling Preventing CPU Thermal Throttling During Long-Running Python Scripts and how to keep clock speeds maxed out without cooking your machine.
#

Melvin
Author
Melvin
I am a software developer building high-performance local AI tools and web architectures. I created Sablegrid, a platform that automates project scoping end to end. Stellar Tech Labs is where I write up what I learn along the way.