Skip to content

Commit d522a6a

Browse files
committed
feat(ingest): parse CPU family and model columns
Handle rowspan family cells, stacked CPU headers, and th-only SKU rows; preserve unknown threads and deduplicate CPU names against the target dataset. Refs #99
1 parent 3038908 commit d522a6a

8 files changed

Lines changed: 11392 additions & 37 deletions

File tree

‎app/ingest/artifacts/amd-cpu-dry-run.json‎

Lines changed: 10675 additions & 0 deletions
Large diffs are not rendered by default.
Lines changed: 89 additions & 0 deletions
Original file line numberDiff line numberDiff line change
@@ -0,0 +1,89 @@
1+
# AMD CPU ingest audit
2+
3+
The reusable CPU table parser now handles stacked CPU headers, separate family
4+
and model columns, rowspan carry, and SKU rows made entirely of `th` cells.
5+
It reads units from headers, excludes GPU/NPU columns, retains variant suffixes,
6+
and does not infer threads from a core count. Architecture comes from the source
7+
column or section, never the product family. Opteron is registered in the existing
8+
CPU collector and therefore uses the normal ingest pipeline and weekly workflow.
9+
10+
CPU additions are deduplicated symmetrically against both names and slugs in the
11+
requested TechAPI data root, ignoring punctuation and manufacturer prefixes.
12+
PRO placement is normalized while PRO/non-PRO and H/HS/U/X/HE variants stay
13+
distinct. Records cite their exact Wikipedia page and stay `verified: false`.
14+
15+
## Dry run against TechAPI develop
16+
17+
Dataset revision: `bf4a0597381cdf208c1fdeceedc1559f2b770bb7` (1,479 AMD CPU
18+
records). The artifact records the dataset content hash and source HTML hashes.
19+
20+
| Live tables | Unique models | Ready additions | Already curated | Incomplete |
21+
| --- | ---: | ---: | ---: | ---: |
22+
| Ryzen | 445 | 52 | 383 | 10 |
23+
| Opteron | 287 | 0 | 102 | 185 |
24+
| Total | 732 | 52 | 485 | 195 |
25+
26+
All 185 unresolved Opteron models lack a stated thread count; 76 also lack a
27+
per-row core count. Across both pages, 22 unresolved models lack a usable release
28+
date. These reasons overlap. Missing table specifications are not evidence that
29+
no public specification exists elsewhere; other authoritative sources can fill
30+
them later without guessing.
31+
32+
The dry-run JSON lists all 52 complete records with their intended output paths,
33+
all incomplete models with reasons, and every already represented model.
34+
No TechAPI files were written and no TechAPI data PR was opened. The existing
35+
weekly workflow creates data PRs against `develop`, but it cannot select just
36+
these two pages and this branch's collector is not deployed on `main`; dispatching
37+
it would also ingest unrelated CPU pages. This is the requested dry-run fallback.
38+
39+
## Reconciliation of issue #19's 759 gaps
40+
41+
The live coverage scraper reproduces **759** exact-slug misses and exactly the
42+
30 rows shown in issue #19: Ryzen contributes 277 entries, Opteron 283, and EPYC
43+
199. Thus the original 759 includes a third page beyond the two requested pages.
44+
45+
| Reproduced gap status | Entries |
46+
| --- | ---: |
47+
| Ready to fill through normal ingest | 21 |
48+
| Already curated under a complete branded name | 324 |
49+
| Missing required table specifications | 167 |
50+
| Family tiers, stepping captions, dates, or part-number cells | 48 |
51+
| EPYC entries outside this two-page audit | 199 |
52+
| Total | 759 |
53+
54+
Of the 759, **21 can be filled by the proposed additions; 738 are not new
55+
additions from this audit**, partitioned above. The 167 incomplete gap entries
56+
include 161 with missing threads, 52 with missing cores, and 17 with missing dates
57+
(overlapping reasons). The other **31 of the 52 ready additions** were omitted
58+
by the coverage scraper's first-cell traversal. Every reproduced entry and its
59+
canonical candidate association appears in the JSON's `coverage_reconciliation`.
60+
61+
These are proposed additions, not applied data changes. The coverage issue's
62+
exact comparison of unqualified table cells with branded curated slugs will
63+
continue to produce false positives until the coverage collector is separately
64+
updated; that collector is outside this task's ownership.
65+
66+
## Reproduce
67+
68+
```powershell
69+
python -m app.ingest.cpu_audit `
70+
--page List_of_AMD_Ryzen_processors `
71+
--page List_of_AMD_Opteron_processors `
72+
--coverage-page List_of_AMD_Ryzen_processors `
73+
--coverage-page List_of_AMD_Opteron_processors `
74+
--coverage-page List_of_AMD_Epyc_processors `
75+
--data-root ../TechAPI/data `
76+
--output app/ingest/artifacts/amd-cpu-dry-run.json
77+
```
78+
79+
For an offline replay, add `--html-dir PATH` containing the three downloaded
80+
`List_of_AMD_*_processors.html` files. Without that flag, the engine's regular
81+
Wikipedia fetcher downloads the pages using its declared user-agent.
82+
83+
## Validation
84+
85+
Repository-wide `ruff check app tests`, `mypy app`, and `python -m app.validate`
86+
pass. The full suite passes: **517 tests**, with **77.16% coverage** against the
87+
60% threshold. The focused parser/pipeline selection contains 36 passing tests.
88+
89+
Refs #99

‎app/ingest/cpu_audit.py‎

Lines changed: 176 additions & 0 deletions
Original file line numberDiff line numberDiff line change
@@ -0,0 +1,176 @@
1+
"""Read-only, reproducible CPU ingest audit; never writes to the dataset.
2+
3+
Example::
4+
5+
python -m app.ingest.cpu_audit --page List_of_AMD_Ryzen_processors \
6+
--page List_of_AMD_Opteron_processors --data-root ../TechAPI/data \
7+
--output amd-cpu-dry-run.json
8+
9+
``--html-dir`` replays previously downloaded HTML instead of fetching pages.
10+
The JSON includes every proposed record and every unresolved unique model.
11+
"""
12+
13+
from __future__ import annotations
14+
15+
import argparse
16+
import hashlib
17+
import json
18+
from collections import Counter
19+
from datetime import UTC, datetime
20+
from pathlib import Path
21+
22+
from app.coverage.sources.wikipedia import fetch_wikipedia_html
23+
from app.coverage.sources.wikipedia_cpu import WikipediaCpu
24+
25+
from .pipeline import run
26+
from .sources.base import IngestCandidate
27+
from .sources.wikipedia_cpu import PAGES, WikipediaCpuIngest
28+
29+
30+
def _entry(candidate: IngestCandidate) -> dict[str, object]:
31+
return {
32+
"output_path": candidate.output_path.as_posix(),
33+
"record": candidate.record,
34+
"missing_fields": list(candidate.missing_fields),
35+
}
36+
37+
38+
def _coverage_audit(
39+
html_by_page: dict[str, str],
40+
candidates: list[IngestCandidate],
41+
ready: set[str],
42+
existing: set[str],
43+
) -> dict[str, object]:
44+
points = {
45+
point.slug: point
46+
for page, html in html_by_page.items()
47+
for point in WikipediaCpu._extract(html, "amd", page)
48+
}
49+
entries = []
50+
for slug, point in sorted(points.items()):
51+
matches = [
52+
c
53+
for c in candidates
54+
if c.source_url == point.url and (c.slug == slug or c.slug.endswith("-" + slug))
55+
]
56+
if not any(c.source_url == point.url for c in candidates):
57+
status = "outside_requested_pages"
58+
elif any(c.slug in ready for c in matches):
59+
status = "ready_to_add"
60+
elif any(c.slug in existing for c in matches):
61+
status = "already_curated"
62+
elif matches:
63+
status = "missing_required_specs"
64+
else:
65+
status = "non_model_or_unparsed_cell"
66+
entries.append(
67+
{
68+
"coverage_slug": slug,
69+
"source_url": point.url,
70+
"status": status,
71+
"candidate_slugs": sorted({c.slug for c in matches}),
72+
"missing_fields": sorted({field for c in matches for field in c.missing_fields}),
73+
}
74+
)
75+
return {
76+
"total": len(points),
77+
"counts": dict(Counter(entry["status"] for entry in entries)),
78+
"entries": entries,
79+
}
80+
81+
82+
def main(argv: list[str] | None = None) -> int:
83+
parser = argparse.ArgumentParser(description=__doc__)
84+
parser.add_argument("--page", action="append", required=True, choices=[p[1] for p in PAGES])
85+
parser.add_argument("--data-root", required=True, type=Path)
86+
parser.add_argument("--output", required=True, type=Path)
87+
parser.add_argument("--html-dir", type=Path)
88+
parser.add_argument(
89+
"--coverage-page",
90+
action="append",
91+
default=[],
92+
help="Optional AMD pages whose raw coverage entries should be reconciled.",
93+
)
94+
args = parser.parse_args(argv)
95+
if not args.data_root.is_dir():
96+
parser.error("--data-root must be an existing TechAPI data directory")
97+
candidates: list[IngestCandidate] = []
98+
sources = []
99+
html_by_page: dict[str, str] = {}
100+
for manufacturer, page, family in PAGES:
101+
if page not in args.page:
102+
continue
103+
html = (
104+
(args.html_dir / f"{page}.html").read_text(encoding="utf-8")
105+
if args.html_dir
106+
else fetch_wikipedia_html(page)
107+
)
108+
sources.append(
109+
{
110+
"url": f"https://en.wikipedia.org/wiki/{page}",
111+
"html_sha256": hashlib.sha256(html.encode()).hexdigest(),
112+
}
113+
)
114+
html_by_page[page] = html
115+
candidates.extend(WikipediaCpuIngest._extract(html, manufacturer, page, family))
116+
result = run(candidates, data_root=args.data_root, dry_run=True)
117+
existing = {c.slug: c for c in result.skipped_existing}
118+
incomplete = {c.slug: c for c in result.skipped_incomplete if c.slug not in existing}
119+
missing = Counter(field for c in incomplete.values() for field in c.missing_fields)
120+
snapshot = hashlib.sha256()
121+
manufacturers = {c.manufacturer for c in candidates}
122+
curated_paths = sorted(
123+
path for maker in manufacturers for path in (args.data_root / "cpu" / maker).rglob("*.json")
124+
)
125+
for path in curated_paths:
126+
snapshot.update(path.relative_to(args.data_root).as_posix().encode())
127+
snapshot.update(b"\0")
128+
snapshot.update(path.read_bytes())
129+
payload = {
130+
"generated_at": datetime.now(UTC).isoformat(),
131+
"dry_run": True,
132+
"include_drafts": False,
133+
"sources": sources,
134+
"curated_cpu_snapshot": {
135+
"records": len(curated_paths),
136+
"sha256": snapshot.hexdigest(),
137+
},
138+
"counts": {
139+
"candidate_rows": len(candidates),
140+
"unique_models": len({c.slug for c in candidates}),
141+
"would_add": len(result.written),
142+
"already_existing": len(existing),
143+
"incomplete": len(incomplete),
144+
"missing_fields": dict(missing),
145+
},
146+
"would_add": [_entry(c) for c in result.written],
147+
"already_existing": sorted(existing),
148+
"incomplete": [
149+
{"slug": c.slug, "missing_fields": list(c.missing_fields), "source_url": c.source_url}
150+
for c in incomplete.values()
151+
],
152+
}
153+
if args.coverage_page:
154+
for page in args.coverage_page:
155+
if page not in html_by_page:
156+
html_by_page[page] = (
157+
(args.html_dir / f"{page}.html").read_text(encoding="utf-8")
158+
if args.html_dir
159+
else fetch_wikipedia_html(page)
160+
)
161+
payload["coverage_reconciliation"] = _coverage_audit(
162+
{page: html_by_page[page] for page in args.coverage_page},
163+
candidates,
164+
{c.slug for c in result.written},
165+
set(existing),
166+
)
167+
args.output.parent.mkdir(parents=True, exist_ok=True)
168+
args.output.write_text(
169+
json.dumps(payload, indent=2, ensure_ascii=False) + "\n", encoding="utf-8"
170+
)
171+
print(json.dumps(payload["counts"]))
172+
return 0
173+
174+
175+
if __name__ == "__main__":
176+
raise SystemExit(main())

‎app/ingest/cpu_identity.py‎

Lines changed: 51 additions & 0 deletions
Original file line numberDiff line numberDiff line change
@@ -0,0 +1,51 @@
1+
"""Conservative, symmetric CPU identity checks for additions-only ingest."""
2+
3+
from __future__ import annotations
4+
5+
import json
6+
import re
7+
from pathlib import Path
8+
9+
from app.coverage.normalize import slugify
10+
11+
12+
def cpu_key(value: str, manufacturer: str) -> str:
13+
tokens = slugify(value, manufacturer=manufacturer).split("-")
14+
if tokens and tokens[0] == manufacturer:
15+
tokens.pop(0)
16+
# Wikipedia uses both "1700X PRO" and "PRO 1700X" for the same SKU.
17+
if "pro" in tokens:
18+
tokens = [token for token in tokens if token != "pro"] + ["pro"]
19+
return "".join(tokens)
20+
21+
22+
def same_cpu(left: str, right: str) -> bool:
23+
if not left or not right:
24+
return False
25+
if left == right:
26+
return True
27+
short, long = sorted((left, right), key=len)
28+
if len(short) < 4 or not long.endswith(short):
29+
return False
30+
# Bare 1200 must not match 41200. Suffixes (X, U, HE, PRO) remain identity.
31+
prefix = long[: -len(short)]
32+
return (
33+
not (short[0].isdigit() and prefix[-1].isdigit())
34+
or re.search(r"[a-z]\d{1,2}$", prefix) is not None
35+
)
36+
37+
38+
def curated_cpu_keys(data_root: Path, manufacturer: str) -> set[str]:
39+
keys: set[str] = set()
40+
for path in (data_root / "cpu" / manufacturer).rglob("*.json"):
41+
try:
42+
record = json.loads(path.read_text(encoding="utf-8"))
43+
except (OSError, json.JSONDecodeError):
44+
continue
45+
if not isinstance(record, dict):
46+
continue
47+
for field in ("slug", "name"):
48+
value = record.get(field)
49+
if isinstance(value, str):
50+
keys.add(cpu_key(value, manufacturer))
51+
return keys

‎app/ingest/pipeline.py‎

Lines changed: 17 additions & 2 deletions
Original file line numberDiff line numberDiff line change
@@ -9,6 +9,7 @@
99

1010
from app.coverage.curated import curated_slugs
1111

12+
from .cpu_identity import cpu_key, curated_cpu_keys, same_cpu
1213
from .sources.base import IngestCandidate
1314

1415

@@ -59,12 +60,24 @@ def run(
5960
result = IngestResult()
6061
curated_by_category: dict[tuple[str, str], set[str]] = {}
6162
written_slugs: set[tuple[str, str, str]] = set()
63+
cpu_keys: dict[str, set[str]] = {}
6264

6365
for candidate in candidates:
6466
key = (candidate.category, candidate.manufacturer)
67+
if candidate.category == "cpu":
68+
if candidate.manufacturer not in cpu_keys:
69+
cpu_keys[candidate.manufacturer] = curated_cpu_keys(
70+
data_root, candidate.manufacturer
71+
)
72+
identity = cpu_key(candidate.slug, candidate.manufacturer)
73+
if any(same_cpu(identity, other) for other in cpu_keys[candidate.manufacturer]):
74+
result.skipped_existing.append(candidate)
75+
continue
6576
if key not in curated_by_category:
66-
curated_by_category[key] = curated_slugs(
67-
candidate.category, candidate.manufacturer
77+
curated_by_category[key] = (
78+
set()
79+
if candidate.category == "cpu"
80+
else curated_slugs(candidate.category, candidate.manufacturer)
6881
)
6982
if candidate.slug in curated_by_category[key]:
7083
result.skipped_existing.append(candidate)
@@ -78,6 +91,8 @@ def run(
7891
continue
7992

8093
written_slugs.add(run_key)
94+
if candidate.category == "cpu":
95+
cpu_keys[candidate.manufacturer].add(identity)
8196
target = data_root / candidate.output_path
8297
if not dry_run:
8398
target.parent.mkdir(parents=True, exist_ok=True)

0 commit comments

Comments
 (0)