Repository navigation
Conversation
Code that feeds a `TreeBuilder` from its own tokenizer, or wraps it to enforce limits, needs to know where tokens will end up: how deeply elements are nested, whether a token is inside a given element such as `template` or `noscript`, and whether character tokens are about to become the contents of an element like `script` or `style`. The tree builder tracks all of this, but the stack of open elements and the insertion mode are private, so such code has to replicate parts of tree construction to recover them. Add two read-only methods to `TreeBuilder`: * `with_open_elements` calls a closure with the stack of open elements as a slice, from the root `html` element to the current node. Passing a slice to a closure, rather than returning a `Ref`, keeps the `RefCell` that holds the stack out of the public API, and means the borrow cannot be held by mistake while the tree builder processes the next token, which would panic. * `is_in_text_insertion_mode` returns whether the insertion mode is "text", in which character tokens are inserted into the current node. A dedicated query avoids making the `InsertionMode` enum public. Neither method changes how documents are parsed.
staylor
marked this pull request as ready for review
October 9, 2026 16:49
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Disclosure: an AI coding assistant wrote the code, tests and this description, at my direction.
This adds two read-only methods to
TreeBuilder. They expose state that embedders can't see today: the stack ofopen elements, and whether the insertion mode is "text". Both are additive and parsing doesn't change.
Motivation
We drive
html5ever::tree_builder::TreeBuilderfrom a different tokenizer, html5gum. We also wrap it in a layerthat enforces resource limits and decides which text to keep before tokens reach the tree builder. For each token,
that layer needs to know:
inside a
script,templateornoscriptelement.forwarded will become the contents of the current node, which is a
script,style,textareaor similarelement.
The tree builder already tracks both, but
open_elemsandmodeare private, so today we patch a vendored copy.Rebuilding this outside the tree builder means replicating tree construction:
<table><td>also openstbodyandtr.A
TreeSinkcan't mirror the stack either.TreeSink::popis called for only some of the ways an element leavesthe stack (#543). For example,
pop_untilandprocess_end_tag_in_bodydon't call it. A driver can approximatetext mode by watching for
TokenSinkResult::RawDataand the end tag that follows, but that duplicates state thetree builder already has.
A depth check on the open elements also gives embedders a way to bound the cost described in #788 today. They can
stop or reject input once the stack gets too deep, without pre-scanning the document.
API
with_open_elementspasses the stack as a slice, from the topmost node (the roothtmlelement) to the bottommostnode (the current node).
TokenSink::end.fruns, so processing a token from insidefcan panic. The docs say so under# Panics.is_in_text_insertion_modeis true from the start tag of an element parsed as raw text or RCDATA until thatelement's end tag or the end of the input. Those elements are
script,style,title,textarea,xmp,iframe,noembed,noframes, andnoscriptwhen scripting is enabled. In this insertion mode, character tokensare inserted into the current node.
The method reports the tree builder's insertion mode, not the tokenizer's state. None of these cases is text mode:
<plaintext>scriptandstylescriptA wrapper can use them like this:
Alternatives considered
fn open_elements(&self) -> Ref<'_, Vec<Handle>>. This is what we patch in today and the smallest diff, butit exposes both the
RefCelland theVec:Refalive while the next token is processed, the tree builder panics with "RefCellalready borrowed".
Ref<'_, [Handle]>. This hides theVec, but the panic is the same andRefis still in the public API.It's the most ergonomic option, and switching to it is a one-line change if you prefer it.
Returning an iterator. If the iterator keeps the borrow, the panic is the same and less visible. If it
borrows again on every step, it has to clone each handle and can see the stack change mid-iteration.
Returning a
Vec<Handle>snapshot. No borrow is exposed, but every call allocates and clones the whole stack.That adds up for a wrapper that checks every token.
Several narrower queries, such as
current_node()oropen_elements_len(). That's more API, and it stilldoesn't cover walking the ancestors.
Making
InsertionModepublic and addinginsertion_mode(). This is more general, but it's a biggercommitment than one boolean, even with
#[non_exhaustive]:types.rssays its types aren't exported.<select>parser (WHATWG proposal) #560 removedInSelectandInSelectInTable.Only "text" is needed here, but I can go that way if you'd rather.
Testing
The new
rcdom/tests/html-tree-builder-state.rsusesRcDomthrough the public driver. It covers:endThese all pass:
cargo test --workspacecargo clippy --all-targetsandcargo clippy --all-features --all-targetscargo fmt --all -- --checkcargo doc, with no warningscargo check --lib --all-featureson Rust 1.85.1, as in the MSRV job