Core Management
Core Management
The Document class is the central API for loading Word documents and converting them to other formats. This page covers format conversion workflows, save-options configuration, and text extraction.
Loading and Saving
Load a document with Document(), then use a dedicated Writer class to convert it to the target output format. Supported inputs: DOCX, DOC, RTF, TXT, Markdown. Supported outputs: PDF, Markdown, TXT.
import aspose.words_foss as aw
from aspose.words_foss.md_writer import LdmMarkdownWriter
from aspose.words_foss.pdf_writer import LdmPdfWriter
doc = aw.Document("input.docx")
LdmMarkdownWriter().write(doc, "output.md")
LdmPdfWriter().write(doc, "output.pdf")
with open("output.txt", "w", encoding="utf-8") as f:
f.write(doc.text)Reuse the same loaded Document with multiple Writer classes to produce multiple output formats without reloading.
PDF Export with PdfSaveOptions
For default PDF output, use LdmPdfWriter() with no options. For fine-grained control, pass a PdfSaveOptions object to its constructor:
import aspose.words_foss as aw
from aspose.words_foss.pdf_writer import LdmPdfWriter
from aspose.words_foss.saving import PdfSaveOptions
doc = aw.Document("input.docx")
# Default PDF export
LdmPdfWriter().write(doc, "default.pdf")
# Customized PDF export with save options
LdmPdfWriter(PdfSaveOptions()).write(doc, "custom.pdf")PdfSaveOptions defines properties for JPEG image quality, PDF standards compliance, font embedding, and other settings.
Note:
PdfSaveOptionsproperties (compliance, JPEG quality, font embedding mode, image compression, etc.) are defined for API forward-compatibility but are not yet consumed by the PDF writer. Setting them currently has no effect on output.
Markdown Export with MarkdownSaveOptions
For default Markdown output, use LdmMarkdownWriter() with no options. Pass a MarkdownSaveOptions object to its constructor when you need to control formatting behavior:
import aspose.words_foss as aw
from aspose.words_foss.md_writer import LdmMarkdownWriter
from aspose.words_foss.saving import MarkdownSaveOptions
doc = aw.Document("input.docx")
# Default Markdown export
LdmMarkdownWriter().write(doc, "default.md")
# Customized Markdown export with save options
LdmMarkdownWriter(MarkdownSaveOptions()).write(doc, "with_options.md")MarkdownSaveOptions supports controlling underline formatting preservation in the output.
Note: Only
export_underline_formattingis currently applied during Markdown export. OtherMarkdownSaveOptionsproperties (table_content_alignment,list_export_mode,export_images_as_base64,images_folder,images_folder_alias) are defined for API forward-compatibility but are not yet consumed by the Markdown writer.
Text Extraction
Extract plain text from any loaded document with the text property:
import aspose.words_foss as aw
doc = aw.Document("input.docx")
text = doc.textFor text file output, write the extracted string to a file:
with open("output.txt", "w", encoding="utf-8") as f:
f.write(doc.text)Common Issues
| Issue | Cause | Fix |
|---|---|---|
ModuleNotFoundError | Package not installed | Reinstall the package (see Installation) |
Empty text from Document.text | Input file is empty or corrupted | Verify the input file opens correctly in a word processor |
| PDF output missing images | Image format not supported by the converter | Use a DOCX input with standard embedded images |
API Reference Summary
| Class / Method | Description |
|---|---|
Document | Load Word documents from DOCX, DOC, RTF, TXT, or Markdown |
LdmPdfWriter().write(doc, path) | Save to PDF |
LdmMarkdownWriter().write(doc, path) | Save to Markdown |
Document.text | Extract plain text content |
SaveFormat | Constants: PDF, MARKDOWN, TEXT |
PdfSaveOptions | PDF export configuration (properties are forward-compatibility stubs; not yet applied) |
MarkdownSaveOptions | Configure underline formatting export |