DocumentReader로 가벼운 문서 모델 구축

DocumentReader로 가벼운 문서 모델 구축

DocumentReader으로 경량 문서 모델 구축

DocumentReader은(는) Aspose.Words FOSS에서 DOCX 파일을 읽기 위한 진입점입니다. 파일 경로, I/O 스트림, 혹은 원시 바이트를 입력으로 받아들인 뒤, 문서를 light_document_model.Document(LDM)으로 노출합니다 — 검사하고 변환하기 쉬운 구조화된 Python 객체 트리입니다.


전제 조건

요구사항세부사항
Python3.9 이상
Packageaspose-words-foss (MIT-라이선스)
Input하나의 .docx 파일, 스트림 또는 bytes
pip install aspose-words-foss

파일 경로에서 로드하기

DocumentReader.load_file(filepath)은(는) 경로를 통해 DOCX 파일을 읽습니다. 이후 to_light_document()을(를) 호출하여 LDM 표현을 얻으세요.

from aspose.words_foss import DocumentReader

reader = DocumentReader()
reader.load_file("data/report.docx")
doc = reader.to_light_document()

print("Page count:", doc.page_count)
print("Sections:", len(doc.sections))

스트림에서 로드하기

DocumentReader.load_stream(stream)은(는) read() 메서드를 가진 파일과 유사한 객체를 받아들입니다. 이는 HTTP 응답, ZIP 아카이브, 또는 메모리 내 버퍼에서 오는 데이터에 유용합니다.

from aspose.words_foss import DocumentReader

reader = DocumentReader()
with open("data/report.docx", "rb") as fh:
    reader.load_stream(fh)

doc = reader.to_light_document()
print("Paragraphs:", len(doc.all_paragraphs))

바이트에서 로드하기

DocumentReader.load_bytes(data)은(는) bytes 또는 bytearray 객체를 받아들입니다. DOCX 내용이 이미 메모리에 있을 때(예: 데이터베이스에서 가져오거나 프로그래밍 방식으로 구성한 경우) 사용하세요.

from aspose.words_foss import DocumentReader

with open("data/report.docx", "rb") as fh:
    data = fh.read()

reader = DocumentReader()
reader.load_bytes(data)
doc = reader.to_light_document()
print("Text preview:", doc.text[:200])

LDM 구조 탐색하기

to_light_document()을(를) 호출한 후 반환된 Document 객체는 전체 문서 트리를 노출합니다:

from aspose.words_foss import DocumentReader

reader = DocumentReader()
reader.load_file("data/report.docx")
doc = reader.to_light_document()

# Sections
for section in doc.sections:
    print("Section paragraphs:", len(section.paragraphs))

# All paragraphs (flat list across all sections)
for para in doc.all_paragraphs:
    if para.text.strip():
        print(" -", para.text[:80])

# Headings
for heading in doc.headings(max_level=2):
    print("H:", heading.text)

전체 문서 텍스트 읽기

Document.text은(는) 모든 섹션에 걸친 모든 단락 텍스트를 연결하여 문서의 전체 순수 텍스트 내용을 반환합니다:

reader = DocumentReader()
reader.load_file("data/report.docx")
doc = reader.to_light_document()
print(doc.text[:500])

팁 및 모범 사례

  • 항상 load_file(), load_stream(), 또는 load_bytes()를 to_light_document() 호출하기 전에 호출하십시오 — 로드되지 않은 리더에서 to_light_document()를 호출하면 오류가 발생하거나 빈 문서를 반환할 수 있습니다.
  • 문서 수명 주기를 메모리 내에서 제어할 때는 load_bytes()를 사용하십시오; 파일 핸들을 여는 것을 방지합니다.
  • 전체 단락을 평면적으로 반복하려면 doc.all_paragraphs를 사용하고, 섹션별로 작업해야 할 때는 doc.sections를 사용하십시오.
  • DocumentReader는 LdmBuilderMixin을 구현하므로 to_light_document()를 믹스인 메서드로 사용할 수 있습니다 — 동일한 인터페이스가 라이브러리의 다른 리더 클래스와 공유됩니다.

일반적인 문제

문제원인수정
로드 후 doc.sections가 비어 있습니다파일이 유효한 DOCX가 아닙니다파일이 Word에서 열리는지 확인하고, 암호로 보호되지 않았는지 확인하십시오
doc.all_paragraphs가 비어 있습니다문서에 본문 텍스트가 없습니다 (예: 머리글/바닥글만 존재)대신 doc.header_paragraphs와 doc.footer_paragraphs를 확인하십시오
AttributeError에 to_light_document()변환 전에 로드 메서드가 호출되지 않았습니다먼저 load_file(), load_stream(), 또는 load_bytes() 중 하나를 호출하십시오

FAQ

LDM이란 무엇입니까?

라이트 문서 모델(light_document_model.Document)은 추상화된 Python 네이티브 객체 트리로, Word 문서를 나타냅니다. 이는 문서 논리를 DOCX 바이너리 형식과 분리하며, 프로그래밍을 통한 손쉬운 조작을 위해 설계되었습니다.

LDM을 수정하고 다시 DOCX에 기록할 수 있나요?

예. 수정한 후 doc 객체를, 다음에 전달하십시오 LdmDocxWriter.write(doc, path) 업데이트된 DOCX 파일을 생성합니다. 자세한 내용은 LdmDocxWriter 가이드 자세한 내용은.

LdmBuilderMixin은(는) 무엇을 추가하나요?

LdmBuilderMixin은(는) to_light_document() 메서드를 제공합니다. DocumentReader는 이를 상속받기 때문에 언제나 리더 인스턴스를 통해 이 메서드에 접근합니다.


API Reference 요약

클래스 / 메서드설명
DocumentReaderDOCX 입력을 읽고 LDM 표현을 구축합니다
DocumentReader.load_file(filepath)경로를 통해 DOCX 파일을 로드합니다
DocumentReader.load_stream(stream)파일과 유사한 스트림에서 로드합니다
DocumentReader.load_bytes(data)bytes 객체에서 로드합니다
DocumentReader.to_light_document()로드된 문서를 LDM Document 형태로 반환합니다
LdmBuilderMixinto_light_document() 를 제공하는 Mixin
Document.sections문서의 섹션 목록
Document.all_paragraphs모든 단락의 평탄한 목록
Document.text문서의 전체 순수 텍스트 내용
Document.page_count페이지 수
Document.headings(max_level)max_level까지의 제목 단락 목록

참조

 한국어