pdf2image: A simple port of Python's pdf2image library to render PDFs into images
Converting PDF to PNG with Python (without pdf2image) - Stack Overflow
Converting PDF to Image without non-python dependencies - Stack Overflow
Simple example to convert pdf to image using pdf2image library in Windows environment. This example shows both jpg or jpeg and png format output from image in pdf.
PyMuPDF supports pdf to image rasterization without requiring any external dependencies.
Sample code to do a basic pdf to png transformation:
import fitz # PyMuPDF, imported as fitz for backward compatibility reasons
file_path = "my_file.pdf"
doc = fitz.open(file_path) # open document
for i, page in enumerate(doc):
pix = page.get_pixmap() # render page to an image
pix.save(f"page_{i}.png")
Here is a snippet that generates PNG images of arbitrary resolution (dpi):
# note: pymupdf can be imported as fitz
# for backward compatibility (use `import pymupdf` in new code)
import fitz
file_path = "my_file.pdf"
dpi = 300 # choose desired dpi here
zoom = dpi / 72 # zoom factor, standard: 72 dpi
magnify = fitz.Matrix(zoom, zoom) # magnifies in x, resp. y direction
doc = fitz.open(fname) # open document
for page in doc:
pix = page.get_pixmap(matrix=magnify) # render page to an image
pix.save(f"page-{page.number}.png")
Generates PNG files name page-0.png, page-1.png, ...
By choosing dpi < 72 thumbnail page images would be created.
I wasn't able to find a solution, apparently one needs a PDF renderer no matter what. The most lightweight solution is https://pymupdf.readthedocs.io/en/latest/intro.html. It is still a python binding for a PDF renderer (https://www.mupdf.com/), but you can install it, including its dependency, by:
pip install PyMuPDF
No need to install poppler or imagemagick.
Then you can convert a pdf to images as follows:
import fitz
doc = fitz.open(stream=your_pdf_file_stream, filetype="pdf")
for idx, page in enumerate(doc):
pix = page.get_pixmap(dpi=600)
the_page_bytes=pix.pil_tobytes(format="PNG")
with open("page-%s.png"%idx, "wb") as outf:
outf.write(the_page_bytes)
Unfortunately, mupdf has a copyleft license, so keep that in mind.
Actually, it took me a while to handle this, but I think it worth it. You need to do all steps carefully to make it work.
- Install pdf2image with
pip install pdf2image. - Get poppler windows binaries.
- Create a new directory like
myproject. - Create a script
converter.pyinsidemyprojectand add below code. - Create another directory inside
myprojectand name itpoppler. - Copy all files in the binary folder of downloaded poppler into
popplerdirectory. Try to testpdfimages.exeif it is working. - Use
pyinstaller converter.py -F --add-data "./poppler/*;./poppler" --noupx - Your executable is now ready. Run it like
converter.exe myfile.pdf. Results would be created inside theoutputdirectory next to the executable. - Now your standalone PDF2IMAGE converter app is ready!
converter.py:
import sys
import os
from pdf2image import convert_from_path
def current_path(dir_path):
if hasattr(sys, '_MEIPASS'):
return os.path.join(sys._MEIPASS, dir_path)
return os.path.join(".", dir_path)
if __name__ == "__main__":
if len(sys.argv) < 2:
print("PASS your PDF file: \"converter.exe myfile.pdf\"")
input()
sys.exit(0)
os.environ["PATH"] += os.pathsep + \
os.pathsep.join([current_path("poppler")])
if not os.path.isdir("./output"):
os.makedirs("output")
images = convert_from_path(sys.argv[-1], 500)
for image, i in zip(images, (range(len(images)))):
image.save('./output/out{}.png'.format(i), 'PNG')
PS: If you like it, you can add a GUI and add more settings for pdf2images.