Skip to content
Coverton Blog
In this article
Documents and text

Extract and clean text from DOCX, HWPX and copied documents

Saving a document as TXT makes its words reusable, but does not retain the original page layout or formatting. Select the tool for DOCX, PPTX, ODT, EPUB, HWPX or RTF, then compare important text with the source.

Use the free tool →

Check the Korean document format first

The HWPX reader handles supported paragraphs, tables and pictures. Legacy HWP is unsupported, and the reader does not reproduce original fonts, shapes, equations or pagination. Preserve the submission original and use the TXT result for text reuse.

Clean line breaks and invisible characters carefully

Copied-PDF cleanup joins internal line breaks while retaining blank-line-separated paragraphs. Review tables, poems and lists against the source. Optional invisible-character removal affects U+200B and BOM only. NFKC and NFKD normalization also change compatibility characters.

Conversion is different from privacy inspection

Markdown-to-HTML converts document markup; HTML-to-TXT discards tags and layout. The Office inspector reports author information, comments, hidden sheets and external-link traces. A report alone does not mean those items have been removed from the file.

Choose the right tool

Example inputs illustrate the workflow; they do not promise a particular file size, recognition accuracy or compatibility.

Extract text from DOCX

Extract text for searching or organizing. The TXT copy does not reproduce document layout or pictures.

Example

document.docx → extracted text.txt
  1. Choose files in DOCX format. Maximum files per batch: 1.
  2. Review Document separator, then start processing.
  3. Download the TXT result and check its content and quality.
What to check

Layout, pictures and tracked editing structure are not reproduced; scanned images are not OCR-processed.

Open the tool →

Extract text from PPTX

Collect slide text into a working transcript. Choose a separator and review extraction order.

Example

slides.pptx → slide-separated text.txt
  1. Choose files in PPTX format. Maximum files per batch: 1.
  2. Review Document separator, then start processing.
  3. Download the TXT result and check its content and quality.
What to check

Reading order may differ from the visual placement of text boxes. Images and speaker delivery are not transcribed.

Open the tool →

Extract text from ODT

Extract plain text from an OpenDocument file. Tables, fonts and pagination are not preserved in TXT.

Example

document.odt → extracted text.txt
  1. Choose files in ODT format. Maximum files per batch: 1.
  2. Review Document separator, then start processing.
  3. Download the TXT result and check its content and quality.
What to check

Page styling, embedded objects and full table layout are not preserved in plain text.

Open the tool →

Extract text from EPUB

Collect text from a supported EPUB for personal workflows. Check support for protected books and complex reading layouts.

Example

book.epub → plain reading text.txt
  1. Choose files in EPUB format. Maximum files per batch: 1.
  2. Review Document separator, then start processing.
  3. Download the TXT result and check its content and quality.
What to check

DRM-protected books are not supported. Navigation, typography and illustrations are not retained in TXT.

Open the tool →

Extract text from HWPX

Extract text from HWPX. Legacy HWP and precise reproduction of the original editing layout require other workflows.

Example

document.hwpx → section text.txt
  1. Choose files in HWPX format. Maximum files per batch: 1.
  2. Review Document separator, then start processing.
  3. Download the TXT result and check its content and quality.
What to check

HWPX and legacy binary HWP are different formats. This tool does not open .hwp files or preserve page layout.

Open the tool →

Convert RTF to plain text

Remove supported RTF control formatting for a text copy. Compare special characters and tables with the source.

Example

document.rtf → plain text.txt
  1. Choose files in RTF format. Maximum files per batch: 1.
  2. Check the input format and file order, then start processing.
  3. Download the TXT result and check its content and quality.
What to check

Complex RTF fields, embedded objects and unusual encodings may not be represented exactly.

Open the tool →

Convert Markdown to HTML

Export Markdown writing as HTML. Check how supported headings, lists and paragraphs render.

Example

# Heading + paragraphs → document.html
  1. Choose files in MD, MARKDOWN format. Maximum files per batch: 1.
  2. Review HTML document style, Open links in a new tab, then start processing.
  3. Download the HTML result and check its content and quality.
What to check

This is a bounded Markdown renderer; not every extension, plugin or embedded script is supported.

Open the tool →

Extract text from HTML

Extract text from an HTML file. This does not visit a URL to scrape a live website.

Example

saved-page.html → readable text.txt
  1. Choose files in HTML, HTM format. Maximum files per batch: 1.
  2. Check the input format and file order, then start processing.
  3. Download the TXT result and check its content and quality.
What to check

Content loaded only by JavaScript is not fetched or executed, and CSS layout is not preserved.

Open the tool →

Read HWPX

Read supported HWPX paragraphs, tables and pictures and extract text. Legacy HWP and precise layout editing are unsupported.

Example

document.hwpx → supported content preview + TXTOpen the tool →

Clean text list

Clean blanks, duplicates and whitespace in one-item-per-line lists. Sort only when changing order is intended.

Example

repeated line items → cleaned line listOpen the tool →

Unicode normalize

Normalize differently composed Unicode, including decomposed Korean text. Compatibility normalization can alter characters.

Example

decomposed text + NFC → composed textOpen the tool →

Invisible characters

Inspect copied text for invisible characters and review positions and code points before removal.

Example

text with invisible characters → location/code-point reportOpen the tool →

Remove line breaks

Join unwanted line breaks in copied prose. Tables, poems and lists can depend on intentional line breaks.

Example

wrapped prose + blank paragraphs → joined paragraphsOpen the tool →

Inspect hidden information in Office files

Inspect author metadata, comments and hidden-content traces before sharing Office files, then edit them in the source application.

Example

Office file → hidden-information inspection reportOpen the tool →

This guide describes the public web tools and their support limits. Keep your original and check the downloaded result. Coverton · Corrections and contact