Skip to content

fix(documents): render tables in TOC-path section text instead of fusing cells - #1361

Merged
dgunning merged 4 commits into
mainfrom
fix/wzgu-toc-section-table-cells
Sep 26, 2026
Merged

dgunning merged 4 commits into
mainfrom
fix/wzgu-toc-section-table-cells

Conversation

@dgunning

Copy link
Copy Markdown
Owner

What was wrong

Text for sections found through the table of contents glued each table row's cells together. For example, Regions' 10-K Item 5 read November 1-30, 20217,294,800: the year 2021 fused with a 7,294,800 share count. Across the 70 tracked 10-K/10-Q/20-F fixtures there were 144 such tokens, all on the TOC path; doc.text() had none. This turned up while fixing GH #1347, whose fix moved ExxonMobil's 10-Q onto the TOC path.

Root cause

SECSectionExtractor._extract_section_content builds TOC section text by walking the raw lxml tree and adds a break only after block elements. td and th aren't block elements, so neighbouring cells ran together. doc.text() renders tables through TableProcessor → TextExtractor/FastTableRenderer instead, so the same job had two implementations.

Fix

  • The TOC walk hands each whole <table> in range to TableProcessor and to a new TextExtractor.render_table, the rule doc.text() uses. A table that contains the section's end boundary keeps the element-wise walk, so the section still stops inside it.
  • Two layout differences from doc.text(), both deliberate:
    • Column width is unbounded. The default truncates cells over 500 chars.
    • Alignment padding collapses to two spaces and rule lines are dropped. The padding alone made table-heavy sections 1.5–2.6× longer.
  • Routing section text through the shared renderer exposed two data losses in that renderer. Both are fixed at the source, so doc.text() gains the content too:
    • FastTableRenderer._identify_meaningful_columns dropped columns of one- or two-character values, such as Regions' "$85" and Morgan Stanley's "WM".
    • TableProcessor._extract_cell_content kept only a cell's <div> text, which lost labels like Netflix's "Derivatives not designated as hedging instruments:".

The TOC-path SIGNATURES cutoff mentioned in the bead already landed in #1355.

Evidence for the re-pinned values

Checked across all 70 tracked 10-K/10-Q/20-F fixtures by counting every letter and digit in doc.text() and in each section, on main and on this branch:

Check Result
TOC-path fused tokens 144 → 0
Documents losing any letter/digit 0
Sections losing any letter/digit 0
Sections added or removed 0
doc.text() +2,474 letters/digits across 58 documents
Largest per-section gain 0.65%; by share, HD Item 4 +17 chars (a page-footer cell, "25 / Table of Contents")

So every changed pin is either collapsed table padding or recovered content. The 20-F's Item 18 going from 181,349 to 151,598 chars is padding. Content assertions were updated to the separated form ("Item 5. Other Information") and none were loosened. The FilingSummary baseline for AAPL reports R8, R13 and R14 changes only because cells now keep their own text ahead of their <div>s.

Verification

Bead: edgartools-wzgu

🤖 Generated with Claude Code

dgunning and others added 4 commits September 25, 2026 18:32
…ing cells

Section text for sections resolved from the table of contents glued each
table row's cells together: Regions' 10-K Item 5 read "November 1-30,
20217,294,800" (the year 2021 welded to a 7,294,800 share count). 106 such
tokens across the 69 tracked 10-K/10-Q/20-F fixtures, all on the TOC path;
doc.text() had none.

Root cause: SECSectionExtractor._extract_section_content builds TOC section
text by walking the raw lxml tree and emitting each element's text, adding a
break only after block elements. td/th are not in that set, so adjacent cells
ran together. doc.text() renders tables through a different implementation:
TableProcessor builds a TableNode and TextExtractor/FastTableRenderer render
it. One rule, two implementations.

Fix: the TOC walk now hands each whole <table> in range to TableProcessor and
TextExtractor.render_table (new public wrapper over the rule doc.text() uses),
and skips the table's subtree. A table that holds the section's end boundary
keeps the element-wise walk so the section still stops inside it. Two layout
differences from doc.text(), both deliberate: column width is unbounded (the
default truncates cells over 500 chars to "...", 161 new losses in section
text), and alignment padding is collapsed to two spaces with rule lines
dropped (the padding alone made table-heavy sections 1.5-2.6x longer and
pushed nine correctly-bounded sections past the size guardrail's bands).

Routing section text through the shared renderer exposed two data losses in
it, fixed at the source so doc.text() gains them too:
- FastTableRenderer._identify_meaningful_columns scored cells by length only,
  so a column of one- and two-character values was dropped as spacing (RF's
  "$85" consumer-loan figure, Morgan Stanley's "WM" segment header). Any cell
  with a letter or digit now makes its column content.
- TableProcessor._extract_cell_content, for a cell with more than one <div>,
  kept only the divs' text: Netflix's "<td><span>Derivatives not designated
  as hedging instruments:</span><div></div><div></div></td>" became an empty
  label, and nested divs were doubled. It now extracts from the whole cell.

Corpus (69 fixtures): glued tokens 106 -> 0; letters/digits lost from any
section 0; doc.text() gains 2,095 letters/digits in 56 documents and loses
none; section chars +0.4%. Pinned lengths and two Filing.text() hashes
re-captured; every change adds or separates tokens, none removes a figure.

The TOC path's missing SIGNATURES cutoff, reported in the same bead, lands
with PR #1355.

Bead: edgartools-wzgu

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
…table-cells

# Conflicts:
#	edgar/documents/extractors/toc_section_extractor.py
Pinned lengths moved because tables in TOC-resolved sections now render
cell by cell (the fix in 593858b) and because #1344/#1345/#1347/#1348,
merged since, added pins measured on the fused text.

Every re-pin was checked against a corpus-wide letter/digit count on
main vs this branch (70 tracked 10-K/10-Q/20-F fixtures): no document
and no section loses a single letter or digit; no section appears or
disappears; TOC-path fused tokens go 144 -> 0; doc.text() gains 2,474
letters/digits in 58 documents. Shorter pins (e.g. the 20-F's Item 18,
181,349 -> 151,598) are collapsed table padding; per-section gains are at
most 0.65% (the largest by share, HD Item 4 +17, is a page-footer cell
"25 / Table of Contents" the shared renderer used to drop). Content
assertions were updated to the separated form ("Item 5.  Other
Information"), never loosened.

The FilingSummary baseline (AAPL R8/R13/R14) changes because
TableProcessor now keeps a cell's own text ahead of its <div>s.

Bead: edgartools-wzgu

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
@dgunning
dgunning enabled auto-merge (squash) September 26, 2026 09:32
@dgunning
dgunning merged commit bf0441b into main Sep 26, 2026
8 of 10 checks passed
@dgunning
dgunning deleted the fix/wzgu-toc-section-table-cells branch September 26, 2026 09:48
dgunning added a commit that referenced this pull request Oct 1, 2026
…ssage counts (#1401)

test_2010_20f_resolves_all_eight_wrapped_items_without_the_legacy_parser
has failed locally since #1361 (bf0441b, bisected from 66e3d48 where
the pins held): Item 5 107,457 -> 106,096, Item 6 58,425 -> 57,677, Item
11 7,504 -> 7,081. It skips in CI because its fixture lives in the
gitignored text_boundary_corpus, so nobody saw it; #1386 then took 3
whitespace characters off Item 5 (word stream identical).

The shrinkage is #1361 removing duplicates, not losing text. Against the
source's own item spans: Item 5's technology transfer (July 2009) and W2E
license (January 2010) paragraphs occur once and were printed twice;
Item 6's "247,900", "Company Headcount", "100,000" (4), "50,000" (11) and
"(2)" (15) were each printed once too often; Item 11's forward-contract
paragraph ("notional amount of $1,695") occurs once in Item 11 (its other
copy is in Item 19's notes) and was printed twice. After #1361 every count
matches the source.

Re-pinned, with count assertions for those passages so the next change is
judged on content rather than length; they fail on bf0441b~1. The other
40 tests on the ignored corpus pass.

Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
dgunning added a commit that referenced this pull request Oct 2, 2026
… so CI runs them (#1403)

Three regression tests read fixtures from the gitignored
tests/fixtures/text_boundary_corpus and carried skipif(not exists), so
they skipped in CI and only ever ran on a developer machine. One of them
(dt1f1's 2010 20-F) went stale for five days after #1361 changed three item
lengths, because only the CI-visible pins were updated (re-pinned in #1401).

Copied into tests/fixtures/parity_gate, the convention the 99-001234,
10-073212 and 16-000635 fixtures already follow:
- 10-K/0000927356-01-000369.html  (0.3 MB; test_3dp_bare_textnode_headers)
- 10-K/0001193125-21-101193.html  (43 KB; test_dt1f1_item_9at)
- 20-F/0001144204-10-017467.html  (3.5 MB; test_dt1f1_wrapped_item_headers)
The tests point at the tracked copies, lose their skipif, and their
CORPUS NOTEs say so. The section parity ratchet lists the three in
TRACKED_GAP_FIXTURES so they cannot quietly disappear; parity_benchmark's
build_corpus labels both trees by accession and prefers the tracked copy,
so no local ratchet result moves.

Verified in a detached worktree (no ignored corpus, CI's view): the three
files plus the section ratchet, 23 passed, 0 skipped. Locally: the three
parity ratchets 16 passed, fast regression 3,472 passed,
check_regression_skips OK.

Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant