Skip to content

html5ever: Add TreeBuilder::with_open_elements and is_in_text_insertion_mode - #798

Open
staylor wants to merge 1 commit into
servo:mainfrom
staylor:staylor/tree-builder-open-elements
Open

staylor wants to merge 1 commit into
servo:mainfrom
staylor:staylor/tree-builder-open-elements

Conversation

@staylor

@staylor staylor commented Oct 9, 2026

Copy link
Copy Markdown

Disclosure: an AI coding assistant wrote the code, tests and this description, at my direction.

This adds two read-only methods to TreeBuilder. They expose state that embedders can't see today: the stack of
open elements, and whether the insertion mode is "text". Both are additive and parsing doesn't change.

Motivation

We drive html5ever::tree_builder::TreeBuilder from a different tokenizer, html5gum. We also wrap it in a layer
that enforces resource limits and decides which text to keep before tokens reach the tree builder. For each token,
that layer needs to know:

  • The current node and its open ancestors. It uses them to cap nesting depth and to tell whether a token is
    inside a script, template or noscript element.
  • Whether the tree builder is in the "text" insertion mode. If it is, the character tokens about to be
    forwarded will become the contents of the current node, which is a script, style, textarea or similar
    element.

The tree builder already tracks both, but open_elems and mode are private, so today we patch a vendored copy.
Rebuilding this outside the tree builder means replicating tree construction:

  • Some elements are implied: <table><td> also opens tbody and tr.
  • Others are closed implicitly.
  • The adoption agency algorithm reorders the stack.
  • In the fragment case, the context element isn't on the stack.

A TreeSink can't mirror the stack either. TreeSink::pop is called for only some of the ways an element leaves
the stack (#543). For example, pop_until and process_end_tag_in_body don't call it. A driver can approximate
text mode by watching for TokenSinkResult::RawData and the end tag that follows, but that duplicates state the
tree builder already has.

A depth check on the open elements also gives embedders a way to bound the cost described in #788 today. They can
stop or reject input once the stack gets too deep, without pre-scanning the document.

API

impl<Handle, Sink> TreeBuilder<Handle, Sink> {
    /// Call `f` with the stack of open elements.
    pub fn with_open_elements<F, R>(&self, f: F) -> R
    where
        F: FnOnce(&[Handle]) -> R;

    /// Is the insertion mode "text"?
    pub fn is_in_text_insertion_mode(&self) -> bool;
}

with_open_elements passes the stack as a slice, from the topmost node (the root html element) to the bottommost
node (the current node).

  • It's empty before the root element is inserted and after TokenSink::end.
  • In the fragment case, the context element isn't on it.
  • The stack is borrowed while f runs, so processing a token from inside f can panic. The docs say so under
    # Panics.

is_in_text_insertion_mode is true from the start tag of an element parsed as raw text or RCDATA until that
element's end tag or the end of the input. Those elements are script, style, title, textarea, xmp,
iframe, noembed, noframes, and noscript when scripting is enabled. In this insertion mode, character tokens
are inserted into the current node.

The method reports the tree builder's insertion mode, not the tokenizer's state. None of these cases is text mode:

  • <plaintext>
  • SVG script and style
  • "in table text"
  • a fragment whose context element is script

A wrapper can use them like this:

let depth = tree_builder.with_open_elements(|elements| elements.len());
let in_template = tree_builder.with_open_elements(|elements| {
    elements.iter().any(|element| {
        tree_builder.sink.elem_name(element).expanded() == expanded_name!(html "template")
    })
});
let raw_text_element = if tree_builder.is_in_text_insertion_mode() {
    tree_builder.with_open_elements(|elements| elements.last().cloned())
} else {
    None
};

Alternatives considered

  • fn open_elements(&self) -> Ref<'_, Vec<Handle>>. This is what we patch in today and the smallest diff, but
    it exposes both the RefCell and the Vec:

  • Ref<'_, [Handle]>. This hides the Vec, but the panic is the same and Ref is still in the public API.
    It's the most ergonomic option, and switching to it is a one-line change if you prefer it.

  • Returning an iterator. If the iterator keeps the borrow, the panic is the same and less visible. If it
    borrows again on every step, it has to clone each handle and can see the stack change mid-iteration.

  • Returning a Vec<Handle> snapshot. No borrow is exposed, but every call allocates and clones the whole stack.
    That adds up for a wrapper that checks every token.

  • Several narrower queries, such as current_node() or open_elements_len(). That's more API, and it still
    doesn't cover walking the ancestors.

  • Making InsertionMode public and adding insertion_mode(). This is more general, but it's a bigger
    commitment than one boolean, even with #[non_exhaustive]:

    Only "text" is needed here, but I can go that way if you'd rather.

Testing

The new rcdom/tests/html-tree-builder-state.rs uses RcDom through the public driver. It covers:

  • the stack across implied elements, implicit closing, foreign content and end
  • the fragment case
  • text mode for every element that enters it, with character tokens landing in the current node
  • leaving text mode at the end tag and at EOF
  • the not-text cases listed above

These all pass:

  • cargo test --workspace
  • cargo clippy --all-targets and cargo clippy --all-features --all-targets
  • cargo fmt --all -- --check
  • cargo doc, with no warnings
  • cargo check --lib --all-features on Rust 1.85.1, as in the MSRV job

Code that feeds a `TreeBuilder` from its own tokenizer, or wraps it to
enforce limits, needs to know where tokens will end up: how deeply
elements are nested, whether a token is inside a given element such as
`template` or `noscript`, and whether character tokens are about to
become the contents of an element like `script` or `style`. The tree
builder tracks all of this, but the stack of open elements and the
insertion mode are private, so such code has to replicate parts of tree
construction to recover them.

Add two read-only methods to `TreeBuilder`:

* `with_open_elements` calls a closure with the stack of open elements
  as a slice, from the root `html` element to the current node. Passing
  a slice to a closure, rather than returning a `Ref`, keeps the
  `RefCell` that holds the stack out of the public API, and means the
  borrow cannot be held by mistake while the tree builder processes the
  next token, which would panic.

* `is_in_text_insertion_mode` returns whether the insertion mode is
  "text", in which character tokens are inserted into the current node.
  A dedicated query avoids making the `InsertionMode` enum public.

Neither method changes how documents are parsed.
@github-actions github-actions Bot added the V-non-breaking A non-breaking change label Oct 9, 2026
@staylor
staylor marked this pull request as ready for review October 9, 2026 16:49
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

V-non-breaking A non-breaking change

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant