PyMuPDF is a high performance Python library for data extraction, analysis, conversion & manipulation of PDF (and other) documents.
-
Updated
Sep 15, 2026 - Python
PyMuPDF is a high performance Python library for data extraction, analysis, conversion & manipulation of PDF (and other) documents.
⚡ From finding text to search and replace, from sorting to beautifying text and more 🎨
Diff Match Patch is a high-performance library in multiple languages that manipulates plain text.
Intuitive find & replace CLI (sed alternative)
fastNLP: A Modularized and Extensible NLP Framework. Currently still in incubation.
Open-source text humanization pipeline with every intermediate step published. Two LLM rewrites at temp 1.3, then two hops across different NMT engines. Four documented methodologies you can read, modify, and run locally.
Python library for creating PEG parsers
Text Classification Algorithms: A Survey
A fast and convenient fuzzy matcher library for rust
Persian NLP Toolkit
The most accurate natural language detection library for Go, suitable for short text and mixed-language text
A fast implementation of Aho-Corasick in Rust.
Program to convert lines of text into a tree structure.
Thai natural language processing in Python
High performance and CommonMark compliant HTML to Markdown converter. Maintained by the Kreuzberg team. Kreuzberg is a fast, polyglot document intelligence engine with a Rust core. It extracts structured data from 98+ document formats using streaming parsers and built-in OCR.
Text Normalization & Inverse Text Normalization
All-in-one text de-duplication
A simple Python module for parsing human names into their individual components
To associate your repository with the text-processing topic, visit your repo's landing page and select "manage topics."