fix(documents): render tables in TOC-path section text instead of fusing cells - #1361
Merged
Merged
Conversation
…ing cells Section text for sections resolved from the table of contents glued each table row's cells together: Regions' 10-K Item 5 read "November 1-30, 20217,294,800" (the year 2021 welded to a 7,294,800 share count). 106 such tokens across the 69 tracked 10-K/10-Q/20-F fixtures, all on the TOC path; doc.text() had none. Root cause: SECSectionExtractor._extract_section_content builds TOC section text by walking the raw lxml tree and emitting each element's text, adding a break only after block elements. td/th are not in that set, so adjacent cells ran together. doc.text() renders tables through a different implementation: TableProcessor builds a TableNode and TextExtractor/FastTableRenderer render it. One rule, two implementations. Fix: the TOC walk now hands each whole <table> in range to TableProcessor and TextExtractor.render_table (new public wrapper over the rule doc.text() uses), and skips the table's subtree. A table that holds the section's end boundary keeps the element-wise walk so the section still stops inside it. Two layout differences from doc.text(), both deliberate: column width is unbounded (the default truncates cells over 500 chars to "...", 161 new losses in section text), and alignment padding is collapsed to two spaces with rule lines dropped (the padding alone made table-heavy sections 1.5-2.6x longer and pushed nine correctly-bounded sections past the size guardrail's bands). Routing section text through the shared renderer exposed two data losses in it, fixed at the source so doc.text() gains them too: - FastTableRenderer._identify_meaningful_columns scored cells by length only, so a column of one- and two-character values was dropped as spacing (RF's "$85" consumer-loan figure, Morgan Stanley's "WM" segment header). Any cell with a letter or digit now makes its column content. - TableProcessor._extract_cell_content, for a cell with more than one <div>, kept only the divs' text: Netflix's "<td><span>Derivatives not designated as hedging instruments:</span><div></div><div></div></td>" became an empty label, and nested divs were doubled. It now extracts from the whole cell. Corpus (69 fixtures): glued tokens 106 -> 0; letters/digits lost from any section 0; doc.text() gains 2,095 letters/digits in 56 documents and loses none; section chars +0.4%. Pinned lengths and two Filing.text() hashes re-captured; every change adds or separates tokens, none removes a figure. The TOC path's missing SIGNATURES cutoff, reported in the same bead, lands with PR #1355. Bead: edgartools-wzgu Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
…table-cells # Conflicts: # edgar/documents/extractors/toc_section_extractor.py
Pinned lengths moved because tables in TOC-resolved sections now render cell by cell (the fix in 593858b) and because #1344/#1345/#1347/#1348, merged since, added pins measured on the fused text. Every re-pin was checked against a corpus-wide letter/digit count on main vs this branch (70 tracked 10-K/10-Q/20-F fixtures): no document and no section loses a single letter or digit; no section appears or disappears; TOC-path fused tokens go 144 -> 0; doc.text() gains 2,474 letters/digits in 58 documents. Shorter pins (e.g. the 20-F's Item 18, 181,349 -> 151,598) are collapsed table padding; per-section gains are at most 0.65% (the largest by share, HD Item 4 +17, is a page-footer cell "25 / Table of Contents" the shared renderer used to drop). Content assertions were updated to the separated form ("Item 5. Other Information"), never loosened. The FilingSummary baseline (AAPL R8/R13/R14) changes because TableProcessor now keeps a cell's own text ahead of its <div>s. Bead: edgartools-wzgu Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
dgunning
enabled auto-merge (squash)
September 26, 2026 09:32
This was referenced Sep 26, 2026
dgunning
added a commit
that referenced
this pull request
Oct 1, 2026
…ssage counts (#1401) test_2010_20f_resolves_all_eight_wrapped_items_without_the_legacy_parser has failed locally since #1361 (bf0441b, bisected from 66e3d48 where the pins held): Item 5 107,457 -> 106,096, Item 6 58,425 -> 57,677, Item 11 7,504 -> 7,081. It skips in CI because its fixture lives in the gitignored text_boundary_corpus, so nobody saw it; #1386 then took 3 whitespace characters off Item 5 (word stream identical). The shrinkage is #1361 removing duplicates, not losing text. Against the source's own item spans: Item 5's technology transfer (July 2009) and W2E license (January 2010) paragraphs occur once and were printed twice; Item 6's "247,900", "Company Headcount", "100,000" (4), "50,000" (11) and "(2)" (15) were each printed once too often; Item 11's forward-contract paragraph ("notional amount of $1,695") occurs once in Item 11 (its other copy is in Item 19's notes) and was printed twice. After #1361 every count matches the source. Re-pinned, with count assertions for those passages so the next change is judged on content rather than length; they fail on bf0441b~1. The other 40 tests on the ignored corpus pass. Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
dgunning
added a commit
that referenced
this pull request
Oct 2, 2026
… so CI runs them (#1403) Three regression tests read fixtures from the gitignored tests/fixtures/text_boundary_corpus and carried skipif(not exists), so they skipped in CI and only ever ran on a developer machine. One of them (dt1f1's 2010 20-F) went stale for five days after #1361 changed three item lengths, because only the CI-visible pins were updated (re-pinned in #1401). Copied into tests/fixtures/parity_gate, the convention the 99-001234, 10-073212 and 16-000635 fixtures already follow: - 10-K/0000927356-01-000369.html (0.3 MB; test_3dp_bare_textnode_headers) - 10-K/0001193125-21-101193.html (43 KB; test_dt1f1_item_9at) - 20-F/0001144204-10-017467.html (3.5 MB; test_dt1f1_wrapped_item_headers) The tests point at the tracked copies, lose their skipif, and their CORPUS NOTEs say so. The section parity ratchet lists the three in TRACKED_GAP_FIXTURES so they cannot quietly disappear; parity_benchmark's build_corpus labels both trees by accession and prefers the tracked copy, so no local ratchet result moves. Verified in a detached worktree (no ignored corpus, CI's view): the three files plus the section ratchet, 23 passed, 0 skipped. Locally: the three parity ratchets 16 passed, fast regression 3,472 passed, check_regression_skips OK. Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
What was wrong
Text for sections found through the table of contents glued each table row's cells together. For example, Regions' 10-K Item 5 read
November 1-30, 20217,294,800: the year 2021 fused with a 7,294,800 share count. Across the 70 tracked 10-K/10-Q/20-F fixtures there were 144 such tokens, all on the TOC path;doc.text()had none. This turned up while fixing GH #1347, whose fix moved ExxonMobil's 10-Q onto the TOC path.Root cause
SECSectionExtractor._extract_section_contentbuilds TOC section text by walking the raw lxml tree and adds a break only after block elements.tdandtharen't block elements, so neighbouring cells ran together.doc.text()renders tables throughTableProcessor→TextExtractor/FastTableRendererinstead, so the same job had two implementations.Fix
<table>in range toTableProcessorand to a newTextExtractor.render_table, the ruledoc.text()uses. A table that contains the section's end boundary keeps the element-wise walk, so the section still stops inside it.doc.text(), both deliberate:doc.text()gains the content too:FastTableRenderer._identify_meaningful_columnsdropped columns of one- or two-character values, such as Regions' "$85" and Morgan Stanley's "WM".TableProcessor._extract_cell_contentkept only a cell's<div>text, which lost labels like Netflix's "Derivatives not designated as hedging instruments:".The TOC-path SIGNATURES cutoff mentioned in the bead already landed in #1355.
Evidence for the re-pinned values
Checked across all 70 tracked 10-K/10-Q/20-F fixtures by counting every letter and digit in
doc.text()and in each section, on main and on this branch:doc.text()So every changed pin is either collapsed table padding or recovered content. The 20-F's Item 18 going from 181,349 to 151,598 chars is padding. Content assertions were updated to the separated form (
"Item 5. Other Information") and none were loosened. The FilingSummary baseline for AAPL reports R8, R13 and R14 changes only because cells now keep their own text ahead of their<div>s.Verification
-m fast -n 4: 7,691 passed, 0 failed.test_wzgu_toc_section_table_cells.py: 6 passed.obj['Item 7A']returns Item 7 #1345 13 passed, TOC: a label cell that carries its own title ("Item 1. Business") yields no item label — FirstEnergy'sobj['Item 7']returns Item 8 #1347 12 passed.check_regression_provenance.pyandcheck_regression_skips.py: OK on 395 files.Bead: edgartools-wzgu
🤖 Generated with Claude Code