Skip to content

Share lexical regions with syntax highlighting - #253

Merged
jserv merged 1 commit into
sysprog21:mainfrom
moon-jam:fix/editor-tokenizer-highlighting
Oct 7, 2026
Merged

jserv merged 1 commit into
sysprog21:mainfrom
moon-jam:fix/editor-tokenizer-highlighting

Conversation

@moon-jam

@moon-jam moon-jam commented Oct 7, 2026 •

Copy link
Copy Markdown
Collaborator

This follows up on #91 by using its shared tokenizer for syntax highlighting. Previously, highlighting maintained separate comment and string rules. For example, an unfinished /* ... affected indentation but remained uncolored until */ appeared.

Highlighting now uses the tokenizer’s comment, string, and regular expression ranges while retaining the existing keyword tables and colors. It reuses the editor’s cached ranges introduced in #238, so bracket matching and highlighting do not tokenize the same buffer twice. When a grammar finishes loading, highlighting refreshes without requiring another edit.

Regression tests cover unfinished comments, multiline strings, template interpolation, regular expressions, bracket markers, escaping, and source preservation. Browser tests also cover delayed and failed grammar downloads, language switching, caret position, and undo/redo. The full offline browser check also passes locally.

Retaining syntax trees for incremental parsing or sharing the same tree between highlighting and indentation could improve performance, but would add complexity. This could be considered in a follow-up if needed.

Refs #91


Summary by cubic

Follows up on #91 by sharing the tokenizer's comment, string, and regexp ranges with syntax highlighting, so highlighting and indentation stay consistent even while a region is unfinished. An unclosed /* ... now takes color immediately instead of waiting for the closing */, and when a grammar finishes loading, the active editor repaints without requiring another edit.

Highlighting keeps the existing keyword, literal, and number tables and reuses the cached token ranges, so the buffer is tokenized once for both indentation and highlighting.

Refactors

  • Highlighting falls back to scanner ranges while a grammar is still loading or after a download fails.
  • Regression tests cover unfinished comments, multiline strings, template interpolation, regexps after declarations, bracket markers, escaping, and preserved source text; browser tests cover delayed and failed grammar downloads, language switching, caret position, and undo/redo.

Written for commit cf5cde8. Summary will update on new commits.

Review in cubic

cubic-dev-ai[bot]

This comment was marked as resolved.

Comment thread web/highlight.js Outdated
Separate comment and string rules let highlighting disagree with
indentation, particularly while code was unfinished. Use the shared
tokenizer for these regions and repaint the active editor when its
grammar becomes ready, retaining existing keyword tables and colors.

Cover scanner and parser paths, delayed and failed grammar loads,
language switches, and editor text, caret and undo behavior.
@moon-jam
moon-jam force-pushed the fix/editor-tokenizer-highlighting branch from 05cebdb to cf5cde8 Compare October 7, 2026 06:06
@jserv
jserv requested a review from ColtenOuO October 7, 2026 06:13

@ColtenOuO ColtenOuO left a comment

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

LGTM, thanks

@jserv
jserv merged commit cf156e7 into sysprog21:main Oct 7, 2026
6 checks passed
@jserv

jserv commented Oct 7, 2026

Copy link
Copy Markdown
Contributor

Thank @moon-jam for contributing!

@moon-jam
moon-jam deleted the fix/editor-tokenizer-highlighting branch October 7, 2026 11:53
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants