Skip to content

fix(verify): audit integrity_check.py for threshold rules blind to confounding variables - #112

Merged
Seungpyo1007 merged 1 commit into
mainfrom
fix/integrity-threshold-audit
Sep 29, 2026
Merged

Seungpyo1007 merged 1 commit into
mainfrom
fix/integrity-threshold-audit

Conversation

@Seungpyo1007

Copy link
Copy Markdown
Member

What & why

Follow-up audit to #111 (cores-aware era rule). That fix showed one integrity
rule had a specific shape of bug: it compared a raw value against a fixed
reference
while ignoring a variable that legitimately shifts what "normal"
looks like
(core count). This PR sweeps every rule in integrity_check.py
for the same shape and fixes the one other place it occurs.

Rules audited

Rule Kind Fixed reference? Verdict
structural (dup slug/name, slug≠file, verified-without-source) hard no — exact identity checks fine
CPU name/tier consistency advisory no — regex on model number fine
CPU single>multi hard no — self-relative invariant, guarded by threads>1 fine
CPU era-vs-score advisory was flat, fixed in #111 already fixed
CPU cross-source ratio outliers advisory yes — single global median±MAD fixed here
GPU cross-source ratio outliers advisory global median±MAD checked, left as-is (see below)

The bug (CPU cross-source ratio outliers)

mad_outliers() computed one global median±MAD over the whole CPU catalog
for ratios like cinebench_r23_multi / geekbench_multi. That ratio is not
scale-free — it's confounded by core count, because the two benchmarks scale
differently with parallelism: Cinebench R23 multi scales near-linearly with
cores while Geekbench multicore compresses at high core counts.

Verified against live TechAPI data (develop, 4,512 CPUs):

  • Per-thread-band median R23/GB ratio climbs monotonically:
    1.05 (1-4T) → 1.22 → 1.34 → 1.43 → 1.37 → 1.48 (65T+).
  • Pearson corr(threads, log-ratio) = +0.52 for R23/GB, −0.59 for
    PassMark/R23.
  • The global median flagged 90 of 739 R23/GB pairs — flagged-thread median
    56 vs 16 overall. The high-ratio side (69 parts, thread-median 96)
    was the entire EPYC / Threadripper / Xeon many-core cluster; the low-ratio
    side (21 parts, thread-median 4) was low-core parts.
  • Spot-checked raw values are genuine, e.g. EPYC 9654 R23 140,000 / GB 58,000
    (192T) = 2.41, EPYC 9754 155,000 / 62,000 (256T) = 2.50 — real scores that
    only look like outliers against a desktop-dominated global median.

Same shape as the flat era ceiling: a fixed reference blind to a variable
(core count) that legitimately shifts normal — so an entire legitimate
population gets flagged.

The fix

Mirrors #111 (which divided the score by thread count): regress the
confounder out
. mad_outliers() now takes an optional per-part covariate,
fits a robust Theil–Sen line of log-ratio vs log(covariate), and runs
the median±MAD test on the residuals, so each part is judged against the
ratio expected for its own core count. CPU pairs pass thread count (falling
back to cores, then 1); callers without a covariate (GPUs) get the original
single-population behaviour unchanged. The check is not deleted — it stays
advisory and still catches a part anomalous for its own class.

Result on live data: R23/GB flags drop 90 → 10, and the survivors are
genuine per-class outliers (Xeon Platinum 8452Y at 5.48, the Skylake-X i9 HEDT
chips, Snapdragon X). The whole legitimate many-core cluster is no longer
flagged. The era / structural / tier / single>multi sections stay clean.

Coarse thread-banding was tried and rejected: it removes the between-band
trend but shrinks the within-band MAD envelope, netting more false positives
(90 → 101). Detrending is the direct analog of #111's per-thread normalisation.

Checked and deliberately left alone

The GPU cross-source ratios (e.g. passmark_g3d_mark / fp32_tflops) also
produce many flags, but the confounder is not a single clean variable: they
pair a theoretical spec (fp32_tflops) with empirical benchmarks across
gaming vs. compute cards and many hardware eras. There is no single stratifier
of the era-rule quality, so — per the "don't weaken a rule without showing the
flagged set is a legitimate pattern it fails to account for" bar — this rule is
left as a single population, advisory-only, and documented as such in the code.

Tests & verification

  • New tests/unit/test_integrity_cross_source.py (mirrors the fix(verify): make era advisory rule cores-aware #111 era-rule
    tests): trend-followers don't flag; the same data without a covariate
    reproduces the old false positives; a part anomalous for its own core count
    still flags; missing covariate == original global behaviour; <8 points never
    flags; zero/None values skipped; reported ratio is the raw a/b; Theil–Sen
    recovers a known slope.
  • pytest tests/unit/test_integrity_cross_source.py tests/unit/test_integrity_era.py → 16 passed.
  • ruff check app tests → clean. mypy app → clean. Full pytest → green.

Refs #98

Audit follow-up to #111 (cores-aware era rule): swept every rule in
integrity_check.py for the same shape of bug — a value compared against a
fixed reference that ignores a variable which legitimately shifts what
"normal" looks like — and found one more instance.

The CPU cross-source ratio detector ran a single global median±MAD over the
whole catalog. But ratios like cinebench_r23_multi/geekbench_multi are
confounded by core count: R23 multi scales near-linearly with cores while
Geekbench multicore compresses, so the ratio climbs monotonically with thread
count (measured on live data: ~1.05 at 1-4T up to ~1.48 at 65T+, Pearson
corr(threads, log-ratio) +0.52; PassMark/R23 falls, corr -0.59). A global
median therefore flagged entire legitimate core-count strata — 90 of 739
R23/GB pairs, median 56 threads vs 16 overall, the whole EPYC/Threadripper/
Xeon many-core cluster with genuine scores — as contamination. Same shape as
the flat era ceiling.

Fix mirrors #111 (which divided score by threads): regress the confounder out.
mad_outliers now accepts an optional per-part covariate, fits a robust
Theil-Sen line of log-ratio vs log(covariate), and runs the median±MAD test on
the residuals, so each part is judged against the ratio expected for its own
core count. CPU pairs pass thread count; the systematic core-count gradient no
longer flags while a part anomalous for its own class still does (live R23/GB
flags 90 -> 10, survivors are genuine per-class outliers). Callers without a
covariate (GPUs) are unchanged. Coarse thread-banding was rejected: it removes
the between-band trend but shrinks the within-band envelope, netting more false
positives.

The GPU cross-source ratios were checked and left as a single population on
purpose: they mix a theoretical spec (fp32_tflops) with empirical benchmarks
across gaming vs. compute cards and many hardware eras, with no single clean
stratifying variable, and stay advisory-only.

Adds tests/unit/test_integrity_cross_source.py.

Refs #98
@Seungpyo1007
Seungpyo1007 merged commit 147d27d into main Sep 29, 2026
1 check passed
@Seungpyo1007
Seungpyo1007 deleted the fix/integrity-threshold-audit branch September 29, 2026 13:54
@Seungpyo1007 Seungpyo1007 added this to the Data & API correctness milestone Sep 29, 2026
@Seungpyo1007 Seungpyo1007 added the bug Something isn't working label Sep 29, 2026
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

bug Something isn't working

Projects

Status: Done

Development

Successfully merging this pull request may close these issues.

2 participants