Type identifier:
text:parseDocumentToContentCategory: Text Operations
Parses PDF files into structured content. Extracts text with page-level granularity, generates page images, and splits content into chunks.
Handle | Data Type | Required | Description |
|---|---|---|---|
|
| Yes | The document file to parse |
|
| No | Dynamic page rendering scale |
documentfilepageScalenumber (0-10)Notes:
Handle | Data Type | Description |
|---|---|---|
|
| Text chunks from document |
|
| Page-level data with images |
|
| Document metadata |
|
| Triggered on parsing error |
contentstring[]Notes:
chunkSize and chunkOverlappagesIncludes:
metadataIncludes:
title: Document titletotalPages: Number of pagesdocumentId: PDF document IDinstanceId: PDF instance IDProperty | Type | Required | Default | Description |
|---|---|---|---|---|
|
| Yes |
| Display label |
|
| Yes |
| Characters per chunk |
|
| Yes |
| Overlap between chunks |
|
| No |
| Page render scale |
chunkSizenumber750 characterschunkOverlapnumber20 characterspageScalenumber, literal or dynamicNotes:
pageScale input handledocument input handleChunks are split as follows:
Scenario | Behaviour |
|---|---|
No file entry | Throws error |
Invalid file entry | Throws error |
PDF parse failure | Triggers |
Corrupted PDF | Triggers |
{
"type": "text:parseDocumentToContent",
"data": {
"label": "Parse PDF",
"chunkSize": 750,
"chunkOverlap": 20
}
}{
"type": "text:parseDocumentToContent",
"data": {
"label": "Parse for Summarization",
"chunkSize": 2000,
"chunkOverlap": 100,
"pageScale": { "type": "literal", "value": 1 }
}
}{
"type": "text:parseDocumentToContent",
"data": {
"label": "Parse for Display",
"chunkSize": 750,
"chunkOverlap": 20,
"pageScale": { "type": "literal", "value": 4 }
}
}