Click or drag to resize
Pdf Library for .NET

PdfToText Class

Extracts text from an existing PDF document and searches it for text, using the cross-platform native PDF engine.
Inheritance Hierarchy
SystemObject
  SelectPdf.UniversalPdfTool
    SelectPdf.UniversalPdfToText

Namespace: SelectPdf.Universal
Assembly: SelectPdf.Universal (in SelectPdf.Universal.dll) Version: 26.4
Syntax
public class PdfToText : PdfTool

The PdfToText type exposes the following members.

Constructors
 NameDescription
Public methodPdfToTextInitializes a new instance of the PdfToText class.
Top
Properties
 NameDescription
Public propertyClipText Do not return hidden text from the PDF document.
Public propertyDocumentInformation PDF document information. Populated after the first operation that returns metadata (e.g. GetPageCount, GetInfo).
(Inherited from PdfTool)
Public propertyEndPageNumber The 1-based page number where the current operation ends. The default value is 0, which means the operation runs to the last page.
(Inherited from PdfTool)
Public propertyHtmlCharset The charset written into the <meta http-equiv="Content-Type"> tag of the HTML produced by GetHtml / SaveHtml(String). The default is UTF-8.
Public propertyLayout The layout used when extracting text: Original preserves the original on-page layout (columns / spacing; the default), or Reading returns the text in reading order.
Public propertyMarkPageBreaks A flag indicating if a page break mark is inserted into the extracted text at the end of each page. The default value is False.
Public propertyPageBreakMark The character inserted into the extracted text at the end of each page when MarkPageBreaks is enabled. This is the form feed character (U+000C).
Public propertyStartPageNumber The 1-based page number where the current operation starts. The default value is 1 (first page).
(Inherited from PdfTool)
Public propertyTimeoutTimeout in seconds for the current operation. Default 600 seconds.
(Inherited from PdfTool)
Public propertyUserPasswordThe user password used to open the PDF document for reading. Default null (no password).
(Inherited from PdfTool)
Top
Methods
 NameDescription
Public methodCloseReleases the loaded document.
Public methodExtractText Extracts the text inside a rectangle (in PDF points, page origin top-left) on a single page (page numbers start at 1).
Public methodGetHtml Returns the document text wrapped in a minimal HTML document — the document metadata as <meta> tags in the head, and the extracted text inside a <pre> block. The result contains only text; no images are extracted.
Public methodGetInfoGets the document information (title, author, page count, etc.).
(Inherited from PdfTool)
Public methodGetPageCountGets the number of pages in the loaded PDF document.
(Inherited from PdfTool)
Public methodCode exampleGetTextReturns all the text in the document (respecting StartPageNumber / EndPageNumber).
Public methodGetText(Int32)Returns all the text on a single page (page numbers start at 1).
Public methodLoad(Byte)Loads a PDF document from a byte array.
(Inherited from PdfTool)
Public methodLoad(PdfDocument)Loads an in-memory PdfDocument for reading.
(Inherited from PdfTool)
Public methodLoad(Stream)Loads a PDF document from a stream.
(Inherited from PdfTool)
Public methodCode exampleLoad(String)Loads a PDF document from a file.
(Inherited from PdfTool)
Public methodLoad(Byte, String)Loads a password-protected PDF document from a byte array.
(Inherited from PdfTool)
Public methodLoad(PdfDocument, String)Loads an in-memory PdfDocument. The password is used to open the document.
(Inherited from PdfTool)
Public methodLoad(Stream, String)Loads a password-protected PDF document from a stream.
(Inherited from PdfTool)
Public methodLoad(String, String)Loads a password-protected PDF document from a file.
(Inherited from PdfTool)
Public methodSaveHtml(String)Writes the document text wrapped in HTML (see GetHtml) to a file (UTF-8).
Public methodSaveHtml(String, Encoding)Writes the document text wrapped in HTML (see GetHtml) to a file with the given encoding.
Public methodSaveText(String)Writes all the document's text to a file (UTF-8).
Public methodSaveText(String, Encoding)Writes all the document's text to a file with the given encoding.
Public methodCode exampleSearch(String) Searches the document for the given text (case-insensitive substring match) and returns the bounding rectangle of every occurrence.
Public methodSearch(String, Boolean, Boolean) Searches the document for the given text and returns the bounding rectangle of every occurrence.
Top
Remarks
Load a document with one of the Load overloads (see PdfTool), then call GetText / ExtractText(Int32, Double, Double, Double, Double) / Search(String). Page numbers on this class are 1-based — note this differs from PdfPageCollection, whose indexer is zero-based, so a PageNumber maps to document.Pages[position.PageNumber - 1].
Example
PdfToText tool = new PdfToText();
tool.Load("document.pdf");
string allText = tool.GetText();
TextPosition[] hits = tool.Search("invoice");
tool.Close();
See Also