Skip to content

fix(ui): stop error toast spam and stuck spinner for stale pod rows - #854

Open
pujitha24 wants to merge 1 commit into
KusionStack:mainfrom
pujitha24:auto/issue-767
Open

pujitha24 wants to merge 1 commit into
KusionStack:mainfrom
pujitha24:auto/issue-767

Conversation

@pujitha24

Copy link
Copy Markdown

What type of PR is this?

/kind bug

What this PR does / why we need it:

On the topology view, expanding a ReplicaSet's Pod table renders one PodStatusCell per row, each polling /rest-api/v1/insight/summary for that pod's live status. When a pod is still present in the search index but has already been deleted from the live cluster (a maintainer-acknowledged, expected transient inconsistency between the search index and the cluster, per the discussion on the linked issue), that call returns {success:false, message: "pods \"X\" not found"} over HTTP 200.

Today every such row triggers a disruptive global error toast, and the row's status cell is left spinning forever, because the shared useAxios hook's handleResponse never updates its response state on a failed /rest-api call -- it only fires the notification and returns, so callers can't distinguish "still loading" from "request finished but failed".

This PR adds an opt-in silentError option to useAxios (default false, so every other existing call site is unaffected). usePodStatus passes silentError: true and now distinguishes "no response yet" (stays "Loading") from "response arrived but failed" (resolves to "Unknown", a status value already recognized by PodStatusCell's existing color mapping).

This does not address the underlying search-index/cluster staleness itself (a separate, harder problem the maintainer is still investigating on the issue thread); it only makes the UI degrade gracefully -- no error-toast spam, no stuck spinner -- when that staleness is encountered.

Validation: This sandbox's disk was exhausted at the OS/container level (confirmed via df -h, unrelated to this repo), so npm install/npm run lint/npm run build could not be run here. Instead, a standalone dependency-free node -e script mirroring the exact branch logic added to both handleResponse and usePodStatus's status resolver was executed directly, asserting: existing (non-opted-in) callers keep notifying and skip setResponse on failure exactly as before; the new silentError: true path skips the notification but still resolves response; and the three-way status resolution (no response -> Loading, success -> resolved status, failed response -> Unknown) behaves as intended. All assertions passed. Flagging this plainly since it's a logic-equivalence check, not an in-app or CI run -- happy to re-verify with npm run lint/npm run build if CI or a reviewer surfaces anything.

Which issue(s) this PR fixes:

Fixes #767

Motivation:
On the topology view, expanding a ReplicaSet's Pod table renders one
PodStatusCell per row, each polling `/rest-api/v1/insight/summary` for
that pod's live status. When a pod is still present in the search
index but has already been deleted from the live cluster (a
maintainer-acknowledged, expected transient inconsistency between the
search index and the cluster), that call returns
`{success:false, message: "pods \"X\" not found"}` over HTTP 200. Every
such row currently triggers a disruptive global error toast, and the
row's status cell is left spinning forever, because the shared
`useAxios` hook's `handleResponse` never updates its `response` state
on a failed `/rest-api` call -- it only fires the notification and
returns, so callers have no way to distinguish "still loading" from
"request finished but failed".

Approach:
Add an opt-in `silentError` option to `useAxios` (default `false`, so
every other existing call site is unaffected -- confirmed no other
caller references it and the failure branch is otherwise unchanged).
`usePodStatus` passes `silentError: true` and now distinguishes "no
response yet" (stays "Loading") from "response arrived but failed"
(resolves to "Unknown", a status value already recognized by
PodStatusCell's existing color mapping in sourceTable/index.tsx).

This does not address the underlying search-index/cluster staleness
itself (a separate, harder problem the maintainer is still
investigating on the issue thread); it only makes the UI degrade
gracefully instead of spamming error toasts and spinning forever when
that staleness is encountered.

Validation:
This sandbox's disk was exhausted at the OS/container level (~116-143Mi
free of 228Gi, confirmed via `df -h`, entirely unrelated to this repo
-- no node_modules exists and the whole checkout is ~10MB), so
`npm install`/`npm run lint`/`npm run build` could not be run, and no
scratch file could even be written to disk. What was run instead: a
standalone `node -e '...'` script requiring no dependencies, mirroring
the exact branch logic added to both `handleResponse` and
`usePodStatus`'s status resolver, asserting (a) callers that don't set
`silentError` still notify-and-skip-setResponse on failure exactly as
before, (b) the new `silentError: true` path skips the notification
but still resolves `response`, and (c) the three-way status
resolution (no response -> Loading, success -> resolved status, failed
response -> Unknown) behaves as intended. All assertions passed. This
is disclosed as a logic-equivalence check, not an in-app or CI run.

Report: KusionStack#767
Signed-off-by: Pujitha Paladugu <10557236+pujitha24@users.noreply.github.com>
Assisted-by: claude-sonnet-5 (via Claude Code)
@github-actions

github-actions Bot commented Sep 12, 2026 •

Copy link
Copy Markdown

CLA Assistant Lite bot All contributors have signed the CLA ✍️ ✅

@pujitha24

Copy link
Copy Markdown
Author

I have read the CLA Document and I hereby sign the CLA

@pujitha24

Copy link
Copy Markdown
Author

recheck

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

Topology get Error Pod message

1 participant