Pdf | |
The PdfToText type exposes the following members.
| Name | Description | |
|---|---|---|
| ClipText | Do not return hidden text from the PDF document. | |
| DocumentInformation |
PDF document information. Populated after the first operation that returns metadata (e.g.
GetPageCount, GetInfo).
(Inherited from PdfTool) | |
| EndPageNumber |
The 1-based page number where the current operation ends. The default value is 0, which means
the operation runs to the last page.
(Inherited from PdfTool) | |
| HtmlCharset | The charset written into the <meta http-equiv="Content-Type"> tag of the HTML produced by GetHtml / SaveHtml(String). The default is UTF-8. | |
| Layout | The layout used when extracting text: Original preserves the original on-page layout (columns / spacing; the default), or Reading returns the text in reading order. | |
| MarkPageBreaks | A flag indicating if a page break mark is inserted into the extracted text at the end of each page. The default value is False. | |
| PageBreakMark | The character inserted into the extracted text at the end of each page when MarkPageBreaks is enabled. This is the form feed character (U+000C). | |
| StartPageNumber |
The 1-based page number where the current operation starts. The default value is 1 (first page).
(Inherited from PdfTool) | |
| Timeout | Timeout in seconds for the current operation. Default 600 seconds. (Inherited from PdfTool) | |
| UserPassword | The user password used to open the PDF document for reading. Default null (no password). (Inherited from PdfTool) |
| Name | Description | |
|---|---|---|
| Close | Releases the loaded document. | |
| ExtractText | Extracts the text inside a rectangle (in PDF points, page origin top-left) on a single page (page numbers start at 1). | |
| GetHtml | Returns the document text wrapped in a minimal HTML document — the document metadata as <meta> tags in the head, and the extracted text inside a <pre> block. The result contains only text; no images are extracted. | |
| GetInfo | Gets the document information (title, author, page count, etc.). (Inherited from PdfTool) | |
| GetPageCount | Gets the number of pages in the loaded PDF document. (Inherited from PdfTool) | |
| GetText | Returns all the text in the document (respecting StartPageNumber / EndPageNumber). | |
| GetText(Int32) | Returns all the text on a single page (page numbers start at 1). | |
| Load(Byte) | Loads a PDF document from a byte array. (Inherited from PdfTool) | |
| Load(PdfDocument) | Loads an in-memory PdfDocument for reading. (Inherited from PdfTool) | |
| Load(Stream) | Loads a PDF document from a stream. (Inherited from PdfTool) | |
| Load(String) | Loads a PDF document from a file. (Inherited from PdfTool) | |
| Load(Byte, String) | Loads a password-protected PDF document from a byte array. (Inherited from PdfTool) | |
| Load(PdfDocument, String) | Loads an in-memory PdfDocument. The password is used to open the document. (Inherited from PdfTool) | |
| Load(Stream, String) | Loads a password-protected PDF document from a stream. (Inherited from PdfTool) | |
| Load(String, String) | Loads a password-protected PDF document from a file. (Inherited from PdfTool) | |
| SaveHtml(String) | Writes the document text wrapped in HTML (see GetHtml) to a file (UTF-8). | |
| SaveHtml(String, Encoding) | Writes the document text wrapped in HTML (see GetHtml) to a file with the given encoding. | |
| SaveText(String) | Writes all the document's text to a file (UTF-8). | |
| SaveText(String, Encoding) | Writes all the document's text to a file with the given encoding. | |
| Search(String) | Searches the document for the given text (case-insensitive substring match) and returns the bounding rectangle of every occurrence. | |
| Search(String, Boolean, Boolean) | Searches the document for the given text and returns the bounding rectangle of every occurrence. |