Skip to content

Seeker ends initial gossip sync on a timer the stream cannot advance, then discards the channel list #9547

Description

@bigg-bb

Summary

On a fresh v26.06.8 node, gossipd's seeker declares the startup peer finished while that peer is still streaming, then issues one full-range query_channel_range and abandons the result. After that it loops forever: ask for 8000 unknown scids, start a random 1024-block probe when the reply completes, return to NORMAL. The graph does not fill. This is not caused by bitcoind latency or by any documented config. The same sequence happens against a local-latency public explorer (about 7s per verified height) and against a remote bitcoind (about 50s per height).

NORMAL here means the idle maintenance loop, not a completed sync. A two-minute trace had 26,705 Adding ... to pending lines and 7 got reply lines.

What the log does

Fresh gossip_store, v26.06.8, autoconnect-seeker-peers left at the default. Seeker created at 16:35:47. Startup peer chosen at 16:36:01 (starting gossip (EVERYTHING)). No seeker log at the first 60s tick. At 16:37:47.451:

  • seeker: startup peer finished
  • query_channel_range for 504500+463795 (from when_lightning_became_cool to the tip)
  • peer answers in 140ms with 24,407 scids; last chunk is 967794+501 (388 scids)
  • next query is 503498+1002, reply is 0 scids
  • 16:37:47.601 seeker: state = NORMAL No unannounced nodes

Then every 60s:

  • ASKING_FOR_UNKNOWN_SCIDS Asking for 8000 scids
  • about 15s later, not on the timer, PROBING_SCIDS for a random 1024-block window
  • that probe doubles backward until a non-empty reply adds no new unknown scid (530339+6506 contained 242 scids) and the state returns to NORMAL

Cause

Three separate problems in gossipd/seeker.c.

1. Startup progress reads a counter the stream never increments.

peer_made_progress requires query_reply_counter to rise by GOSSIP_SEEKER_INTERVAL (60). That counter is only increased from handle_reply_channel_range (gossipd/queries.c), by tal_count(scids) / 20. The startup path sends gossip_timestamp_filter, not a range query, so the counter stays 0 while channel updates and announcements are arriving.

selected_peer then does:

seeker->prev_gossip_count = peer->query_reply_counter - GOSSIP_SEEKER_INTERVAL(seeker);

Both are size_t. With a counter of 0 this wraps. The next check evaluates 0 >= prev + 60, which wraps back to 0 >= 0 and counts as progress, and stores prev_gossip_count = 0. The check after that sees 0 >= 60, fails, and check_firstpeer logs startup peer finished. That is the 106s gap in the log (peer chosen at :01, silent tick, finished on the second tick). peer_supplied_good_gossip updates gossip_counter, and only after an update or node announcement is actually stored. The watchdog does not read it.

2. The follow-up range uses the last reply slice, then an empty reply ends the walk.

handle_reply_channel_range invokes the callback with the first_blocknum and number_of_blocks from that reply, not from the query. The last slice was +501. next_block_range doubles that (prev_num_blocks * 2) and, because scid_probe_start > 0, steps backward:

504500 - 1002 = 503498, query 503498+1002.

That window is before when_lightning_became_cool and returns 0 scids. process_scid_probe continues only when new_unknown_scids is true, so the probe stops and the state becomes NORMAL. add_unknown_scid has already returned false past 10,000 (seeker->num_unknown_scids > 10000), so most of the 24,407 ids from the real reply were dropped before this empty step.

3. Completing an unknown-scid query always starts a random probe.

scid_query_done does not look at seeker state. It always calls probe_some_random_scids (1024 blocks). The in-flight 8000-scid fetch is therefore replaced as soon as its reply_short_channel_ids_end arrives. The random probe doubles backward until add_unknown_scid refuses every id in a reply (the 10,000 map is full), new_unknown_scids is false, and the state returns to NORMAL. One minute later seeker_check asks for 8000 again.

PENDING_LIMIT in gossipd/gossmap_manage.c is also 10,000, and channel_announcement: Adding %s to pending is logged before map_add, so the Adding count includes announcements that are then dropped. One height is verified at a time through gettxout. That limits how fast the ids that survived are checked. It is not why the seeker stops: the same stop happened when a height completed in 6.6s.

Expected

A cold start should keep the initial stream until the peer stops sending, or should walk the channel-range replies until the requested span is stored, and should not treat a slice-sized backward step or a random 1024-block sample as the end of that sync.

Config

No option changes this. autoconnect-seeker-peers only sets how many peers to hold. --dev-fast-gossip shortens the same timer from 60s to 5s. --dev-fast-gossip-prune and --dev-throttle-gossip do not touch this path. Reproduced with bcli disabled and trustedcoin on public explorers, and separately with a remote bitcoind over a 113ms link.

Version

Core Lightning v26.06.8. when_lightning_became_cool is 504500 (bitcoin/chainparams.c).

Note

this is completely genered by ai alongside human assistence

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions