Login
Free Sign Up
Docs
/

Text - Parse Document to Content

Type identifier: text:parseDocumentToContent Category: Text Operations

Description

Parses PDF files into structured content. Extracts text with page-level granularity, generates page images, and splits content into chunks.

Input Handles

Handle

Data Type

Required

Description

document

file

Yes

The document file to parse

pageScale

number

No

Dynamic page rendering scale

Input Handle Details

document

  • Type: file
  • Supported Formats: PDF

pageScale

  • Type: number (0-10)
  • Notes:

    • Controls image quality for page rendering

Output Handles

Handle

Data Type

Description

content

array

Text chunks from document

pages

array

Page-level data with images

metadata

object

Document metadata

error

trigger

Triggered on parsing error

Output Handle Details

content

  • Type: string[]
  • Notes:

    • Text split into chunks
    • Respects chunkSize and chunkOverlap
    • Ready for embedding or search indexing

pages

  • Type: Array of page objects
  • Includes:

    • Page number
    • Page text
    • Rendered page image (file)

metadata

  • Type: Object
  • Includes:

    • title: Document title
    • totalPages: Number of pages
    • documentId: PDF document ID
    • instanceId: PDF instance ID

Configuration Options

Property

Type

Required

Default

Description

label

string

Yes

"Parse Document to Content"

Display label

chunkSize

number

Yes

750

Characters per chunk

chunkOverlap

number

Yes

20

Overlap between chunks

pageScale

number (literal or dynamic)

No

{ type: "literal", value: 2 }

Page render scale

Configuration Details

chunkSize

  • Type: number
  • Default: 750 characters
  • Notes: Target size for each text chunk

chunkOverlap

  • Type: number
  • Default: 20 characters
  • Notes: Overlap between consecutive chunks to maintain context

pageScale

  • Type: number, literal or dynamic
  • Range: 0-10
  • Default: 2
  • Notes:

    • Higher values = better quality, larger images
    • Dynamic: the scale comes from the pageScale input handle

Behaviour

Execution Flow

  1. Resolves document input handle
  2. Validates the file input
  3. Loads the file
  4. Parses PDF
  5. Extracts text from each page
  6. Renders page images at configured scale
  7. Splits text into chunks
  8. Returns content, pages, and metadata

PDF Processing

  • Supports standard PDF features
  • Handles encrypted PDFs (if password-free)

Text Splitting

Chunks are split as follows:

  • Attempts to split at natural boundaries (paragraphs, sentences)
  • Falls back to character boundaries
  • Maintains overlap for context continuity

Error Handling

Scenario

Behaviour

No file entry

Throws error

Invalid file entry

Throws error

PDF parse failure

Triggers error

Corrupted PDF

Triggers error

Examples

Basic Document Parsing

{
  "type": "text:parseDocumentToContent",
  "data": {
    "label": "Parse PDF",
    "chunkSize": 750,
    "chunkOverlap": 20
  }
}

Large Chunk Processing

{
  "type": "text:parseDocumentToContent",
  "data": {
    "label": "Parse for Summarization",
    "chunkSize": 2000,
    "chunkOverlap": 100,
    "pageScale": { "type": "literal", "value": 1 }
  }
}

High Quality Page Images

{
  "type": "text:parseDocumentToContent",
  "data": {
    "label": "Parse for Display",
    "chunkSize": 750,
    "chunkOverlap": 20,
    "pageScale": { "type": "literal", "value": 4 }
  }
}

Related pages