Skip to main content

Search OfficeIMO

Enter a topic, API type, or PowerShell command.

Documentation Search all OfficeIMO/

API Reference

Class

PdfToWordOptions

Namespace OfficeIMO.Word.Pdf
Assembly OfficeIMO.Word.Pdf
Modifiers sealed

Options for importing PDF content as editable Word objects or rendered page images.

Inheritance

  • Object
  • PdfToWordOptions

Remarks

The default editable mode reconstructs content through the first-party logical PDF reader. It preserves supported metadata, page breaks, headings, paragraphs, list items, logical tables, and source placeholders. VisualPages instead embeds managed-rendered page appearances at their physical page sizes; image quality depends on the rendering resolution and the PDF renderer's supported content.

Usage

This type appears in these public API surfaces even when no hand-authored example is attached directly to the page.

Constructors

Methods

public PdfToWordOptions Clone() #
Returns: PdfToWordOptions

Creates a reusable copy of this option set.

public static PdfToWordOptions CreateTablesOnly() #
Returns: PdfToWordOptions

Creates an import profile that reconstructs only detected PDF tables.

public static PdfToWordOptions CreateVisualPages() #
Returns: PdfToWordOptions

Creates an appearance-preserving profile using rendered page images.

Properties

public PdfWordImportMode Mode { get; set; } #

Import strategy. Editable reconstruction remains the default.

public Double Dpi { get; set; } #

Resolution of rendered page images in VisualPages mode.

public Int32 MaxPages { get; set; } #

Maximum pages rendered in VisualPages mode.

public Int64 MaxPixelsPerPage { get; set; } #

Maximum pixels for each rendered page in VisualPages mode.

public Int64 MaxOutputBytesPerPage { get; set; } #

Maximum encoded bytes for each rendered page in VisualPages mode.

public Int64 MaxTotalOutputBytes { get; set; } #

Maximum aggregate encoded page-image bytes in VisualPages mode.

public Int64 MaxImageTextOverlapComparisons { get; set; } #

Maximum image versus text-span comparisons during editable import.

public PdfReadOptions ReadOptions { get; set; } #

Canonical semantic-read settings used when importing an opened PdfDocument. Null uses Default. This setting is ignored when the source is already a PdfDocumentReadResult.

public Boolean IncludeMetadata { get; set; } #

Whether PDF Info dictionary metadata should be copied into Word built-in properties.

public Boolean PreservePageBreaks { get; set; } #

Whether source PDF page transitions should be represented by Word page breaks.

public Boolean PreserveSourcePageSize { get; set; } #

Whether editable output should create page sections with the physical width and height of each source PDF page. This avoids silently converting A4, landscape, or mixed-size PDFs to Word's default Letter page size.

public Double EditablePageMarginPoints { get; set; } #

Uniform page margin, in points, used by source-sized editable sections. The narrow default leaves enough flow space for dense business documents while retaining a printable margin.

public Boolean PreserveCompactSourceSpacing { get; set; } #

Whether imported PDF paragraphs and table cells should suppress Word's style-level before/after spacing. PDF text already carries explicit line placement, so compact spacing avoids inflating dense business documents.

public Boolean IncludeEmptyPages { get; set; } #

Whether empty PDF pages should produce an empty Word paragraph when page breaks are preserved.

public Boolean ImportHeadings { get; set; } #

Whether logical heading lines should be imported as Word heading paragraphs.

public Boolean ImportParagraphs { get; set; } #

Whether grouped logical paragraphs should be imported as Word paragraphs.

public Boolean UseSharedPageReadingOrder { get; set; } #

Whether semantic content should use the shared crop-, rotation-, and column-aware logical reading order. Disable only when preserving the legacy top-to-bottom page sort is required.

public Boolean ImportLists { get; set; } #

Whether logical list items should be imported as editable Word list items.

public Boolean ImportTables { get; set; } #

Whether logical PDF tables should be imported as editable Word tables.

public String BookmarkPrefix { get; set; } #

Prefix used for generated Word bookmarks that represent imported PDF pages and named destinations.

public ISet<String> AllowedHyperlinkUriSchemes { get; } #

Absolute URI schemes allowed when creating active Word hyperlinks from PDF URI annotations.

public Boolean ImportImages { get; set; } #

Whether complete PDF image-file payloads should be embedded as native Word images.

public Boolean PreserveImagePlacementSize { get; set; } #

Whether embedded images should use detected PDF placement size when available.

public Boolean PreserveImagePlacementPosition { get; set; } #

Whether axis-aligned images on unrotated pages should retain their source page position as floating Word images. Unsupported rotations and transforms fall back to inline editable placement.

public Boolean IncludeImagePlaceholders { get; set; } #

Whether image resources should be represented by editable placeholder paragraphs when native image embedding is unavailable or disabled.

public Boolean IncludeFormFieldPlaceholders { get; set; } #

Whether AcroForm widgets should be represented by editable placeholder paragraphs.

public Int32 MaxTableRows { get; set; } #

Maximum body rows to import per detected table. Values less than or equal to zero import all rows.

public WordTableStyle TableStyle { get; set; } #

Word table style applied to imported PDF tables.

public Boolean RepeatHeaderRows { get; set; } #

When true, tables with inferred column headers repeat the first row at the top of each Word page.

public Boolean FitTablesToPageWidth { get; set; } #

When true, imported tables use the full content width while preserving detected PDF column proportions.

public Boolean AlignNumericColumns { get; set; } #

When true, body cells in inferred numeric PDF columns are right-aligned in the generated Word tables.

public String EmptyDocumentMessage { get; set; } #

Paragraph text written when no supported PDF content is detected, keeping the produced document meaningful.