Login
Free Sign Up
Docs
/

Text HTML Query and Transform

Type identifier: text:html:queryAndTransform Category: Text Operations

Description

The Text HTML Query and Transform node manipulates HTML through a sequence of configurable operations. It can query elements using CSS selectors to extract content (HTML, text, or attributes) and remove unwanted elements from the document. Each query operation creates a dynamic output handle for its results.

Input Handles

Handle

Type

Description

html

string

The HTML document to query and transform.

Output Handles

Handle

Type

Description

result

string

The transformed HTML document after all operations.

{operationId}

varies

Dynamic handles for each query operation (see Operations).

Configuration Options

Operations

The node supports a list of operations executed in sequence:

Query Operation

Extracts content from elements matching a CSS selector.

Field

Type

Default

Description

Selector

string

-

CSS selector to match elements.

Match All

boolean

true

If true, matches all elements; if false, matches only the first.

Output Type

"html" | "text" | "attributes"

"html"

Format of extracted content.

Attributes

string[]

-

When Output Type is "attributes", specifies which attributes to extract.

Output handle type varies based on configuration:

  • html output: string or string[] (if Match All)
  • text output: string or string[] (if Match All)
  • attributes output: object or object[] with requested attribute keys

Remove Operation

Removes elements matching a CSS selector from the document.

Field

Type

Default

Description

Selector

string

-

CSS selector for elements to remove.

Match All

boolean

true

If true, removes all matching elements.

Behaviour

  1. Parses the input HTML document
  2. Executes operations in sequence:

    • Query operations: Extract content and set to operation's output handle
    • Remove operations: Delete matching elements from the document
  3. Returns the transformed HTML document via the result handle

CSS Selector Support

Uses standard CSS selector syntax:

  • Element selectors: p, div, span
  • Class selectors: .class-name
  • ID selectors: #element-id
  • Attribute selectors: [attr], [attr=value]
  • Combinators: >, +, ~, (descendant)
  • Pseudo-classes: :first-child, :last-child

Output Type Details

Output Type

Description

Example Output

html

Full HTML including tags

<p class="intro">Hello</p>

text

Text content only

Hello

attributes

Specified attribute values

{ "class": "intro", "id": null }

Examples

Extract Article Content

Operations:

  1. Query: Selector article h1, Output Type text — the title
  2. Query: Selector article p, Output Type html — the paragraphs
  3. Remove: Selector script, style, nav

Input:

<html>
  <nav>Menu</nav>
  <article>
    <h1>Article Title</h1>
    <p>First paragraph.</p>
    <p>Second paragraph.</p>
  </article>
  <script>
    console.log("removed");
  </script>
</html>

Outputs:

  • Query 1: "Article Title"
  • Query 2: ["<p>First paragraph.</p>", "<p>Second paragraph.</p>"]
  • Result: HTML with nav, script, and style removed

Extract Links with Attributes

Operations:

  1. Query: Selector a, Attributes ["href", "title"], Output Type attributes

Input:

<p>Visit <a href="https://example.com" title="Example">our site</a>.</p>

Output:

[{ "href": "https://example.com", "title": "Example" }]

Clean HTML for Processing

Operations:

  1. Remove: Selector noscript, script, style, [hidden], [aria-hidden], link, svg

Use case: Preparing HTML for text extraction or AI processing.

Related pages