DocumentReader로 가벼운 문서 모델 구축
DocumentReader으로 경량 문서 모델 구축
DocumentReader은(는) Aspose.Words FOSS에서 DOCX 파일을 읽기 위한 진입점입니다. 파일 경로, I/O 스트림, 혹은 원시 바이트를 입력으로 받아들인 뒤, 문서를 light_document_model.Document(LDM)으로 노출합니다 — 검사하고 변환하기 쉬운 구조화된 Python 객체 트리입니다.
전제 조건
| 요구사항 | 세부사항 |
|---|---|
| Python | 3.9 이상 |
| Package | aspose-words-foss (MIT-라이선스) |
| Input | 하나의 .docx 파일, 스트림 또는 bytes |
pip install aspose-words-foss파일 경로에서 로드하기
DocumentReader.load_file(filepath)은(는) 경로를 통해 DOCX 파일을 읽습니다. 이후 to_light_document()을(를) 호출하여 LDM 표현을 얻으세요.
from aspose.words_foss import DocumentReader
reader = DocumentReader()
reader.load_file("data/report.docx")
doc = reader.to_light_document()
print("Page count:", doc.page_count)
print("Sections:", len(doc.sections))스트림에서 로드하기
DocumentReader.load_stream(stream)은(는) read() 메서드를 가진 파일과 유사한 객체를 받아들입니다. 이는 HTTP 응답, ZIP 아카이브, 또는 메모리 내 버퍼에서 오는 데이터에 유용합니다.
from aspose.words_foss import DocumentReader
reader = DocumentReader()
with open("data/report.docx", "rb") as fh:
reader.load_stream(fh)
doc = reader.to_light_document()
print("Paragraphs:", len(doc.all_paragraphs))바이트에서 로드하기
DocumentReader.load_bytes(data)은(는) bytes 또는 bytearray 객체를 받아들입니다. DOCX 내용이 이미 메모리에 있을 때(예: 데이터베이스에서 가져오거나 프로그래밍 방식으로 구성한 경우) 사용하세요.
from aspose.words_foss import DocumentReader
with open("data/report.docx", "rb") as fh:
data = fh.read()
reader = DocumentReader()
reader.load_bytes(data)
doc = reader.to_light_document()
print("Text preview:", doc.text[:200])LDM 구조 탐색하기
to_light_document()을(를) 호출한 후 반환된 Document 객체는 전체 문서 트리를 노출합니다:
from aspose.words_foss import DocumentReader
reader = DocumentReader()
reader.load_file("data/report.docx")
doc = reader.to_light_document()
# Sections
for section in doc.sections:
print("Section paragraphs:", len(section.paragraphs))
# All paragraphs (flat list across all sections)
for para in doc.all_paragraphs:
if para.text.strip():
print(" -", para.text[:80])
# Headings
for heading in doc.headings(max_level=2):
print("H:", heading.text)전체 문서 텍스트 읽기
Document.text은(는) 모든 섹션에 걸친 모든 단락 텍스트를 연결하여 문서의 전체 순수 텍스트 내용을 반환합니다:
reader = DocumentReader()
reader.load_file("data/report.docx")
doc = reader.to_light_document()
print(doc.text[:500])팁 및 모범 사례
- 항상
load_file(),load_stream(), 또는load_bytes()를to_light_document()호출하기 전에 호출하십시오 — 로드되지 않은 리더에서to_light_document()를 호출하면 오류가 발생하거나 빈 문서를 반환할 수 있습니다. - 문서 수명 주기를 메모리 내에서 제어할 때는
load_bytes()를 사용하십시오; 파일 핸들을 여는 것을 방지합니다. - 전체 단락을 평면적으로 반복하려면
doc.all_paragraphs를 사용하고, 섹션별로 작업해야 할 때는doc.sections를 사용하십시오. DocumentReader는LdmBuilderMixin을 구현하므로to_light_document()를 믹스인 메서드로 사용할 수 있습니다 — 동일한 인터페이스가 라이브러리의 다른 리더 클래스와 공유됩니다.
일반적인 문제
| 문제 | 원인 | 수정 |
|---|---|---|
로드 후 doc.sections가 비어 있습니다 | 파일이 유효한 DOCX가 아닙니다 | 파일이 Word에서 열리는지 확인하고, 암호로 보호되지 않았는지 확인하십시오 |
doc.all_paragraphs가 비어 있습니다 | 문서에 본문 텍스트가 없습니다 (예: 머리글/바닥글만 존재) | 대신 doc.header_paragraphs와 doc.footer_paragraphs를 확인하십시오 |
AttributeError에 to_light_document() | 변환 전에 로드 메서드가 호출되지 않았습니다 | 먼저 load_file(), load_stream(), 또는 load_bytes() 중 하나를 호출하십시오 |
FAQ
LDM이란 무엇입니까?
라이트 문서 모델(light_document_model.Document)은 추상화된 Python 네이티브 객체 트리로, Word 문서를 나타냅니다. 이는 문서 논리를 DOCX 바이너리 형식과 분리하며, 프로그래밍을 통한 손쉬운 조작을 위해 설계되었습니다.
LDM을 수정하고 다시 DOCX에 기록할 수 있나요?
예. 수정한 후 doc 객체를, 다음에 전달하십시오 LdmDocxWriter.write(doc, path) 업데이트된 DOCX 파일을 생성합니다. 자세한 내용은 LdmDocxWriter 가이드 자세한 내용은.
LdmBuilderMixin은(는) 무엇을 추가하나요?
LdmBuilderMixin은(는) to_light_document() 메서드를 제공합니다. DocumentReader는 이를 상속받기 때문에 언제나 리더 인스턴스를 통해 이 메서드에 접근합니다.
API Reference 요약
| 클래스 / 메서드 | 설명 |
|---|---|
DocumentReader | DOCX 입력을 읽고 LDM 표현을 구축합니다 |
DocumentReader.load_file(filepath) | 경로를 통해 DOCX 파일을 로드합니다 |
DocumentReader.load_stream(stream) | 파일과 유사한 스트림에서 로드합니다 |
DocumentReader.load_bytes(data) | bytes 객체에서 로드합니다 |
DocumentReader.to_light_document() | 로드된 문서를 LDM Document 형태로 반환합니다 |
LdmBuilderMixin | to_light_document() 를 제공하는 Mixin |
Document.sections | 문서의 섹션 목록 |
Document.all_paragraphs | 모든 단락의 평탄한 목록 |
Document.text | 문서의 전체 순수 텍스트 내용 |
Document.page_count | 페이지 수 |
Document.headings(max_level) | max_level까지의 제목 단락 목록 |