# pdf-to-docx-semi-manual-convertor semi manual convertion tool from pdf to docx # PDF → DOCX OCR Editor A Windows-first desktop application for manually defining PDF page regions as text, images, or tables and converting them into visually positioned DOCX content. ## Features - Python 3.11+ - QtPy + PySide6 - PyMuPDF PDF rendering - Local Tesseract OCR - OpenCV preprocessing - Pillow image processing - python-docx DOCX generation - Normalized page coordinates - Text/image/table selection - Green text selections - Yellow image selections - Red table selections - OCR in background threads - Image extraction in background threads - Table recognition in background threads - DOCX generation in background threads - Mixed portrait/landscape pages - Empty/skipped pages retained - Project JSON persistence - Font selection - Tesseract configuration - Logging - High-DPI support - Mouse-wheel zoom - Page navigation - OCR text editing - Table image fallback ## Installation Python 3.11 or newer is recommended. Create a virtual environment: py -3.11 -m venv .venv Activate it: .venv\Scripts\activate Upgrade pip: python -m pip install --upgrade pip Install dependencies: pip install -r requirements.txt ## Tesseract OCR Tesseract is an external executable and is not installed by pip. Recommended Windows location: C:\Program Files\Tesseract-OCR\tesseract.exe The application automatically checks common Windows installation locations. If Tesseract is installed elsewhere: 1. Start the application. 2. Open OCR Settings. 3. Select the Tesseract executable. 4. Set the OCR language. 5. Save the settings. The default OCR language is: eng Additional languages can be installed into the Tesseract tessdata directory. ## Running Start the application: python main.py A PDF file chooser opens automatically. ## Workflow 1. Open a PDF. 2. Select Text, Image, or Table. 3. Draw a rectangle on the PDF page. 4. The corresponding background worker processes the region. 5. The right-side preview updates. 6. Move or resize selections if necessary. 7. Double-click text to edit OCR. 8. Navigate to the next page. 9. Generate the DOCX. ## Coordinate system All selections are stored as normalized values: x y width height where all values are between 0 and 1. For example: x = 0.20 y = 0.30 width = 0.40 height = 0.10 means the object starts 20% from the left, 30% from the top, is 40% of the page width and 10% of the page height. The PDF viewer can therefore zoom or resize without changing the logical document layout. ## Page handling Every source PDF page becomes a DOCX section. If a page contains no selections, the section is still generated and remains an empty page. Mixed page dimensions and orientations are supported. ## DOCX positioning The renderer uses low-level Office Open XML techniques. Text: VML text boxes Images: DrawingML floating anchors Tables: w:tblpPr floating tables This avoids simply appending content as normal flowing paragraphs. ## Project files Projects use: filename.pdfocrproject They contain: - PDF path - DOCX path - current page - page sizes - selections - OCR text - extracted images - table data - font settings - OCR settings ## Testing Run: pytest -q ## Windows packaging Run: build_windows.bat The result is: dist\PDFToDOCXOCR\PDFToDOCXOCR.exe Tesseract is deliberately not embedded by the Python package because it is a separate native OCR application. For a completely self-contained installer, distribute the Tesseract runtime alongside the application and point the application to that executable, or add it to the PyInstaller spec as a native dependency. ## Logs Logs are stored in the operating-system application-data directory. Windows: %APPDATA%\PDFToDOCXOCR\logs\application.log ## Troubleshooting ### "Tesseract not found" Install Tesseract and configure the executable under OCR Settings. ### OCR produces poor text Try: - increasing render DPI - enabling preprocessing - selecting a tighter region - installing the correct Tesseract language - selecting the correct language code ### Table becomes an image The table recognizer currently targets bordered/grid-style tables. Tables without clear ruling lines can be difficult to segment reliably. When recognition fails, the application intentionally preserves the selected region as an image rather than destroying the document. ### DOCX looks slightly different from PDF Word and PDF have different layout engines. The renderer uses absolute positioning to minimize differences, but font metrics, line wrapping, VML rendering, and Word's pagination engine can introduce small variations. Microsoft Word is the primary target for output verification. ## Security PDF files are processed locally. The application does not upload PDFs. It does not execute embedded PDF JavaScript or macros. Tesseract runs locally. ## Distribution Before commercial distribution, review the licenses of: - PyMuPDF - Qt/PySide6 - python-docx - Pillow - OpenCV - Tesseract In particular, PyMuPDF is available under AGPL or commercial licensing, so choose the appropriate licensing path for your product. ## Recommended production deployment For enterprise deployment, additionally: - code-sign the Windows executable - bundle a tested Tesseract release - pin all dependencies - run automated DOCX rendering regression tests - test output in Microsoft Word - test output in LibreOffice separately - maintain a PDF corpus covering invoices, reports, forms and tables - retain application logs for diagnostics