Convert Word Documents to Pdf | |
Using the SelectPdf Word to Pdf Converter for .NET, a Word document can be converted to pdf with a few lines of code. The converter reads the formats Microsoft Word writes - docx, docm, dotx, dotm, doc, dot - plus rtf, plain text, WordML, and html and markdown through its own importers. The conversion runs in-process: no engine is launched.
The Word to Pdf converter is available only in the full commercial SelectPdf library. It is not available in the free SelectPdf community edition. |
The WordToPdf class exposes ConvertFile(String), ConvertStream(Stream) and ConvertBytes(Byte), each with an overload that names the WordFormat to read. Without one, the format is taken from the file extension, or detected from the content for a stream or a byte array (docx family, doc, rtf, WordML). Plain text, html and markdown have no signature and must be named when given as a stream.
The result is a PdfDocument, exactly as the html converter returns one: it can be modified further, saved to a file, a stream or a byte array, and must be closed. The Options, Header and Footer properties mirror the html converter's, so a document setup written for HtmlToPdf applies here unchanged - page size and margins, header and footer bands, PDF/A and PDF/X standards, security, viewer preferences, document properties, compression, page ranges and single-page mode. The options are reviewed in Options Review.
A Word document has pages of its own, so PageSetup decides what the page size options mean:
Document (the default) keeps every page the size and orientation the Word document gives it, mixed sizes included. The page size and orientation options are ignored; the margins and the header and footer bands are applied around the rendered page, which is scaled uniformly to fit inside them. With zero margins and no bands the page is stamped 1:1 - nothing is reflowed and no page break moves.
Options reflows the text onto the page size, orientation and margins of the options, replacing the Word document's own margins. This is what "convert this docx to A4 landscape with one-inch margins" means.
This sample converts a Word document, keeping its own pages, under a header band and page numbers, and saves it as PDF/A-2b.
WordToPdf converter = new WordToPdf(); // Document keeps the Word page sizes; // Options would reflow onto PdfPageSize / margins converter.Options.PageSetup = WordPageSetup.Document; converter.Options.PdfStandard = PdfStandard.PdfA2B; // a text header band and a page-number footer converter.Options.DisplayHeader = true; converter.Header.Height = 40; converter.Header.Add(new PdfTextSection(0, 12, 500, "Quarterly report", PdfFont.CreateStandardFont(PdfStandardFont.HelveticaBold, 11), new PdfColor(44, 62, 80))); converter.Options.DisplayFooter = true; converter.Footer.Height = 30; converter.Footer.Add(new PdfTextSection(0, 10, 500, "Page {page_number} of {total_pages}", PdfFont.CreateStandardFont(PdfStandardFont.Helvetica, 9), new PdfColor(80, 80, 80))); // the format is taken from the extension PdfDocument doc = converter.ConvertFile("report.docx"); // the outline was built from the document's headings int chapters = doc.Bookmarks.Count; doc.Save("report.pdf"); doc.Close();
Every font the document uses is embedded in the pdf as a subset (or complete, with EmbedCompleteFonts). A font family the host does not have is served from the fonts the library bundles - the Liberation faces, metric-compatible with Arial, Times New Roman and Courier New - so the same document converts on a Windows server, a Linux container with no fonts installed, or a Mac. Scripts those faces do not cover (CJK, Arabic, Indic) need fonts installed on the host.
Every PDF/A level and PDF/X-1a are produced the way the html converter produces them, through PdfStandard.
Accessible (tagged) output is declared the way it is on the html converter: Tagged builds the logical structure - the Word document's headings, paragraphs, lists, tables with their header cells, pictures, links, footnotes and text boxes - and AccessibilityStandard declares PDF/UA-1 or PDF/UA-2 (which is written as PDF 2.0). A declared document needs a Title and a Language. PdfStandard.PdfA3A turns tagging on by itself, and combines with PDF/UA-1; PDF/A-4 is the archiving level that combines with PDF/UA-2. Running headers and footers are marked as pagination artifacts, and the fonts of a header or footer band are embedded on a declared document, PDF/UA requiring it.
A picture without alternative text in the Word document is given the placeholder Image; give every picture a description in Word for a document that is accessible in substance and not only in form. A validator such as veraPDF is the way to check the produced document at the level it declares.