DocumentReaderで軽量ドキュメントモデルを構築する

DocumentReaderで軽量ドキュメントモデルを構築する

DocumentReader を使用した軽量ドキュメントモデルの構築

DocumentReader は Aspose.Words FOSS における DOCX ファイル読み取りのエントリーポイントです。ファイルパス、I/O ストリーム、または生バイト列のいずれかを入力として受け取り、ドキュメントを light_document_model.Document (LDM) として公開します — 検査および変換が容易な構造化された Python オブジェクトツリーです。


前提条件

要件詳細
Python3.9 以降
Packageaspose-words-foss (MIT-licensed)
InputA .docx ファイル、ストリーム、または bytes
pip install aspose-words-foss

ファイルパスからのロード

DocumentReader.load_file(filepath) はパスで DOCX ファイルを読み取ります。その後、to_light_document() を呼び出して LDM 表現を取得します。

from aspose.words_foss import DocumentReader

reader = DocumentReader()
reader.load_file("data/report.docx")
doc = reader.to_light_document()

print("Page count:", doc.page_count)
print("Sections:", len(doc.sections))

ストリームからのロード

DocumentReader.load_stream(stream) は read() メソッドを持つ任意のファイルライクオブジェクトを受け付けます。これは HTTP のレスポンスや ZIP アーカイブ、またはメモリ内バッファからのデータに便利です。

from aspose.words_foss import DocumentReader

reader = DocumentReader()
with open("data/report.docx", "rb") as fh:
    reader.load_stream(fh)

doc = reader.to_light_document()
print("Paragraphs:", len(doc.all_paragraphs))

バイトからの読み込み

DocumentReader.load_bytes(data) は bytes または bytearray オブジェクトを受け付けます。DOCX コンテンツがすでにメモリ上にある場合(例: データベースから取得したり、プログラムで構築したりした場合)に使用してください。

from aspose.words_foss import DocumentReader

with open("data/report.docx", "rb") as fh:
    data = fh.read()

reader = DocumentReader()
reader.load_bytes(data)
doc = reader.to_light_document()
print("Text preview:", doc.text[:200])

LDM 構造のナビゲーション

to_light_document() を呼び出した後、返される Document オブジェクトは完全なドキュメントツリーを公開します:

from aspose.words_foss import DocumentReader

reader = DocumentReader()
reader.load_file("data/report.docx")
doc = reader.to_light_document()

# Sections
for section in doc.sections:
    print("Section paragraphs:", len(section.paragraphs))

# All paragraphs (flat list across all sections)
for para in doc.all_paragraphs:
    if para.text.strip():
        print(" -", para.text[:80])

# Headings
for heading in doc.headings(max_level=2):
    print("H:", heading.text)

全文ドキュメントテキストの読み取り

Document.text はドキュメントの完全なプレーンテキストコンテンツを返します。すべてのセクションの段落テキストを連結して:

reader = DocumentReader()
reader.load_file("data/report.docx")
doc = reader.to_light_document()
print(doc.text[:500])

ヒントとベストプラクティス

  • 常に load_file()、load_stream()、または load_bytes() を to_light_document() を呼び出す前に呼び出してください — 読み込みが行われていないリーダーで to_light_document() を呼び出すと、エラーが発生するか空のドキュメントが返される可能性があります。
  • メモリ上でドキュメントのライフサイクルを管理する場合は load_bytes() を使用してください。ファイルハンドルを開く必要がなくなります。
  • すべての段落をフラットに反復処理する場合は doc.all_paragraphs を使用し、セクションごとに処理する必要がある場合は doc.sections を使用してください。
  • DocumentReader は LdmBuilderMixin を実装しており、to_light_document() がミックスインメソッドとして利用可能であることを意味します — 同じインターフェースはライブラリ内の他のリーダークラスでも共有されています。

一般的な問題

問題原因修正
読み込み後に doc.sections が空ですファイルは有効な DOCX ではありませんファイルが Word で開くことを確認し、パスワードで保護されていないか確認してください
doc.all_paragraphs が空です文書に本文テキストがありません(例:ヘッダー/フッターのみ)代わりに doc.header_paragraphs と doc.footer_paragraphs を確認してください
AttributeError が to_light_document() にあります変換前にロードメソッドが呼び出されていませんまず、load_file()、load_stream()、またはload_bytes()のいずれかを呼び出してください

FAQ

LDM とは何ですか?

軽量ドキュメントモデル (light_document_model.Document) は、抽象化された、Python ネイティブのオブジェクトツリーで、Word ドキュメントを表します。ドキュメントロジックを DOCX バイナリ形式から分離し、プログラムによる容易な操作を可能にするよう設計されています。

LDM を変更して DOCX に書き戻すことはできますか?

はい。変更した後は doc オブジェクトを、そこに渡します LdmDocxWriter.write(doc, path) 更新されたDOCXファイルを生成します。以下をご覧ください LdmDocxWriter ガイド 詳細については.

LdmBuilderMixin は何を追加しますか?

LdmBuilderMixin は to_light_document() メソッドを提供します。DocumentReader はそれを継承するため、常にリーダーインスタンス経由でこのメソッドにアクセスします。


API Reference の概要

クラス / メソッド説明
DocumentReaderDOCX入力を読み取り、LDM表現を構築します
DocumentReader.load_file(filepath)パスでDOCXファイルを読み込みます
DocumentReader.load_stream(stream)ファイルのようなストリームから読み込みます
DocumentReader.load_bytes(data)bytes オブジェクトから読み込みます
DocumentReader.to_light_document()ロードされたドキュメントを LDM Document として返します
LdmBuilderMixinto_light_document() を提供するミックスイン
Document.sectionsドキュメント内のセクションの一覧
Document.all_paragraphsすべての段落のフラットなリスト
Document.text文書の完全なプレーンテキストコンテンツ
Document.page_countページ数
Document.headings(max_level)max_level までの見出し段落の一覧

参照

 日本語