feat(ingest): fill absent software and website fields from Wikidata - #106
Merged
Merged
Conversation
Seungpyo1007
force-pushed
the
feat/wikidata-fill
branch
from
September 27, 2026 08:10
0f0a367 to
b392118
Compare
P571 is the owner's inception (e.g. a newspaper founded in 1785), so dates before 1991 are skipped for launch_date. Refs #99
This was referenced Sep 27, 2026
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Software and website records often lack fields already available on their cited Wikidata entities. Add a dry-run-first CLI that fills only absent values using the existing schema fields and compares proposed records with the current offline scorer.
Batch 50 entities per request with an identified User-Agent, maxlag=5, sequential requests spaced by at least one second, incremental JSON caching, and retries that leave failures uncached. Explicit response-size truncation retries only omitted entities in smaller batches. Resolve item properties to English labels, skip missing/redirect entities and deprecated statements, and prefer preferred rank. Dates require day precision and the Gregorian calendar because the current schemas require YYYY-MM-DD; month/year claims are skipped without inventing days. Apply preserves key order, two-space indentation, trailing newline, BOM and existing line endings.
Validation: 497 tests passed with an external data checkout and explicit basetemp; mypy passed across 108 app files; ruff passed across app/tests. Fifteen mocked tests cover mappings, rank, precision, no-overwrite, redirects/missing entities, rate spacing, cache reuse, error handling, response truncation and dry-run/apply serialization. Full-category dry runs did not apply any dataset records.
The 30 website green losses come from pre-Web newspaper inception dates that trigger launch_date_plausible; the requested direct P571 mapping is preserved. For example, Aftenposten receives 1860-05-14 and moves from 78.8 to 73.3. Review this semantic ambiguity before applying. Offline green projections do not promote records or prove source liveness.
Refs #99
Refs GetTechAPI/TechAPI#295