Zebrad stalling

My node has same issues described in this thread, my first full sync took ~15h about 8 months ago, now I resynced from scratch because of v5 release and after 56h I’m at ~97% on mainnet (testnet needed 48h).

At this point I’ll finish sync with v5 release, waiting for next release.

2 Likes

Ok, mainnet is synced after 64h and many zebra restarts…

1 Like

Big congratulations! You’ve really been through a lot with that. I’m still stuck around 70% myself. I think I’ll just sit back and wait for the next release after version 5.0.0.

2 Likes

I would like to express my sincere gratitude to Autotunafish for introducing the fix/5709-sync-stall branch (v5.0.0+15). It has been incredibly helpful. My zebrad is now syncing perfectly without any stalls.[ing 77%]

5 Likes

Gracias, me alegraron el dia. ya queria estrellar esto contra la pared.

1 Like

I completed node synchronization to 100% two days ago using the fix/5709-sync-stall branch, but the issue does not appear to be fully resolved.

To be fair, synchronization does continue and eventually progresses, albeit very slowly. However, I am still observing behavior that suggests the underlying problem may persist.

Please refer to the data below for additional context.

2026-06-08T23:48:15.031373Z INFO zebrad::commands::start: initializing syncer

2026-06-08T23:48:15.031389Z INFO zebrad::commands::start: initializing mempool
2026-06-08T23:48:15.031399Z INFO zebrad::commands::start: fully initializing inbound peer request handler
2026-06-08T23:48:15.031465Z WARN zebrad::commands::start: configure an indexer_listen_addr to start the indexer RPC server
2026-06-08T23:48:15.031475Z INFO zebrad::commands::start: spawning block gossip task
2026-06-08T23:48:15.031534Z INFO zebrad::commands::start: spawning mempool queue checker task
2026-06-08T23:48:15.031556Z INFO zebrad::commands::start: spawning mempool transaction gossip task
2026-06-08T23:48:15.031560Z INFO zebrad::commands::start: spawning delete old databases task
2026-06-08T23:48:15.031570Z INFO zebrad::commands::start: spawning progress logging task
2026-06-08T23:48:15.031574Z INFO zebrad::commands::start: initializing health endpoints
2026-06-08T23:48:15.031578Z INFO zebrad::commands::start: spawning end of support checking task
2026-06-08T23:48:15.031583Z INFO zebrad::commands::start: spawning mempool crawler task
2026-06-08T23:48:15.031586Z INFO zebrad::commands::start: spawning syncer task
2026-06-08T23:48:15.031583Z INFO zebra_state::config: checking for old database versions db_kind=“state”
2026-06-08T23:48:15.031592Z INFO zebrad::commands::start: spawned initial Zebra tasks
2026-06-08T23:48:15.031623Z INFO zebrad::components::sync::gossip: initializing block gossip task
2026-06-08T23:48:15.031636Z INFO zebrad::components::mempool::queue_checker: initializing mempool queue checker task
2026-06-08T23:48:15.031641Z INFO zebrad::components::mempool::gossip: initializing transaction gossip task
2026-06-08T23:48:15.031671Z WARN zebrad::components::sync::progress: chain tip metrics channel closed err=SendError { .. }
2026-06-08T23:48:15.031685Z INFO zebra_state::config: finished old database version cleanup task
2026-06-08T23:48:15.031738Z INFO zebrad::components::sync::end_of_support: Starting end of support task
2026-06-08T23:48:15.031775Z INFO zebrad::components::mempool::crawler: initializing mempool crawler task
2026-06-08T23:48:15.031891Z INFO sync:try_to_sync: zebrad::components::sync: starting sync, obtaining new tips state_tip=Some(Height(3370103))
2026-06-08T23:48:18.032476Z INFO crawl_and_dial{new_peer_interval=61s}:dial{candidate=MetaAddr { addr: v4redacted:8233, services: None, untrusted_last_seen: None, last_response: None, rtt: None, ping_sent_at: None, last_attempt: Some(Instant { tv_sec: 1553, tv_nsec: 817103481 }), last_failure: None, misbehavior_score: 0, last_connection_state: AttemptPending, is_inbound: false }}: zebra_network::peer_set::initialize: failed to make outbound connection to peer error=Elapsed(()) candidate.addr=v4redacted:8233
2026-06-08T23:48:25.032926Z INFO zebrad::components::sync::end_of_support: Checking if Zebra release is inside support range …
2026-06-08T23:48:25.032973Z INFO zebrad::components::sync::end_of_support: Zebra release is supported until block 3479160, please report bugs at Issues · ZcashFoundation/zebra · GitHub
2026-06-08T23:48:25.412256Z INFO init{config=Config { checkpoint_sync: true } network=Mainnet}: zebra_consensus::router: finished state checkpoint validation
2026-06-08T23:48:28.194944Z INFO zebra_state::service::finalized_state::disk_format::upgrade: database format is valid running_version=27.0.0 initial_disk_version=27.0.0
2026-06-08T23:48:35.033314Z INFO peer_cache_updater: zebra_network::config: updated cached peer IP addresses cached_ip_count=14 peer_cache_file=“/root/.cache/zebra/network/mainnet.peers”
2026-06-08T23:49:15.052996Z INFO zebrad::components::sync::progress: estimated progress to chain tip sync_percent=99.970% current_height=Height(3370279) network_upgrade=Nu6_2 remaining_sync_blocks=1002 time_since_last_state_block=25s
2026-06-08T23:50:15.072440Z INFO zebrad::components::sync::progress: estimated progress to chain tip sync_percent=99.970% current_height=Height(3370279) network_upgrade=Nu6_2 remaining_sync_blocks=1003 time_since_last_state_block=1m 25s
2026-06-08T23:51:15.091616Z INFO zebrad::components::sync::progress: estimated progress to chain tip sync_percent=99.970% current_height=Height(3370279) network_upgrade=Nu6_2 remaining_sync_blocks=1004 time_since_last_state_block=2m 25s
2026-06-08T23:52:15.113498Z INFO zebrad::components::sync::progress: estimated progress to chain tip sync_percent=99.970% current_height=Height(3370279) network_upgrade=Nu6_2 remaining_sync_blocks=1005 time_since_last_state_block=3m 25s
2026-06-08T23:53:15.131654Z INFO zebrad::components::sync::progress: estimated progress to chain tip sync_percent=99.970% current_height=Height(3370279) network_upgrade=Nu6_2 remaining_sync_blocks=1006 time_since_last_state_block=4m 25s
2026-06-08T23:53:35.034771Z INFO peer_cache_updater: zebra_network::config: updated cached peer IP addresses cached_ip_count=57 peer_cache_file=“/root/.cache/zebra/network/mainnet.peers”
2026-06-08T23:54:15.150503Z INFO zebrad::components::sync::progress: estimated progress to chain tip sync_percent=99.970% current_height=Height(3370279) network_upgrade=Nu6_2 remaining_sync_blocks=1006 time_since_last_state_block=5m 25s
2026-06-08T23:55:15.167626Z INFO zebrad::components::sync::progress: estimated progress to chain tip sync_percent=99.970% current_height=Height(3370279) network_upgrade=Nu6_2 remaining_sync_blocks=1007 time_since_last_state_block=6m 25s
2026-06-08T23:56:15.186432Z INFO zebrad::components::sync::progress: estimated progress to chain tip sync_percent=99.970% current_height=Height(3370279) network_upgrade=Nu6_2 remaining_sync_blocks=1008 time_since_last_state_block=7m 25s
2026-06-08T23:56:47.812589Z WARN sync:try_to_sync:try_to_sync_once: zebrad::components::sync: error downloading and verifying block e=ValidationRequestError { error: Elapsed(()), height: Height(3370287), hash: block::Hash(“00000000005b78976eeeb795f5f78fc34d0efe7823661b0fb891e790b2a274cb”) }
2026-06-08T23:56:47.812710Z INFO sync: zebrad::components::sync: waiting to restart sync timeout=67s state_tip=Some(Height(3370279))
2026-06-08T23:57:15.203308Z INFO zebrad::components::sync::progress: estimated progress to chain tip sync_percent=99.970% current_height=Height(3370279) network_upgrade=Nu6_2 remaining_sync_blocks=1009 time_since_last_state_block=8m 25s

1 Like

Another thing I did was also delete the existing mainnet.peers in .cache/zebra/network. i don’t know if it helped or not but you could try

1 Like

I’ve cleared my peer address book about 30 times over the last few days. It does help temporarily. If I happen to connect with a good peer, the synchronization speeds up for a few minutes or maybe tens of minutes. But eventually, things go right back to the same problematic environment. It’s not a real solution. Honestly, I can barely even call it a temporary fix at this point.

I really wish the developers and the Foundation would bring this issue out into the open. I think they need to discuss things with the whole community—clearly explaining what the actual problem is, how ZEC network participants can help solve it, and what practical roadblocks they are hitting along the way.

1 Like

2026-06-09T10:24:29+07:00 /Zebra:5.0.0/ height=3371449 peers=54 errors=“attempted to add a banned peer addr to address book”

Saw this error message today. Seems like there are a lot of bad peers causing this stalling issue. Someone sabotaging the nodes?

1 Like

Fixes the sync stall where Zebra nodes freeze for minutes during initial block download, requiring repeated restarts to make progress. Root cause: FindBlocks responses didn’t register block availability in the inventory router, so getdata requests were load-balanced to peers that didn’t have the block — causing a NotFound storm that poisoned the registry and left frontier blocks permanently un-routable.

Three coordinated changes that together produce a 0-error genesis-to-tip sync:

  • Stop poisoning the inventory registry on transient errors (client.rs). Only mark inventory as missing for explicit notfound responses — timeouts and connection drops no longer poison routing.

  • Re-request dropped blocks (sync.rs). Re-queue a block whose download failed with NotFound instead of silently dropping it. Bounded by MAX_BLOCK_REOBTAIN_RETRIES (3 attempts). Without this, a single missing block at the checkpoint frontier is dropped and never re-fetched, wedging the verify pipeline until the 8-minute BLOCK_VERIFY_TIMEOUT fires.

  • Duplicate-tolerant batch dispatch (sync.rs). Catches DuplicateBlockQueuedForDownload errors and continues processing the remaining batch instead of dropping unprocessed hashes — preventing frontier gaps from missed blocks.

Why this surfaced now

The underlying bug was always present — FindBlocks responses never registered block availability, and the inventory registry always poisoned on transient errors. But several converging factors turned this latent defect into a visible stall:

  • Sandblasting-era blocks (~1.4M–2M) take 5–8 minutes to verify. Before sandblasting, blocks were cheap and the misrouting was invisible — a dropped block was recovered on the next restart cycle before anyone noticed the 8-minute timeout.

  • Degraded peer set. As more nodes stalled or ran old versions, at-tip peer ratio dropped as low as 2–5%. Fewer good peers means more misrouted requests hit peers that genuinely don’t have the block, accelerating the poisoning cascade.

2 Likes

Well done!

1 Like
2 Likes

My node was able to sync to the tip with 5.0.0. After upgrading to 5.1.0, it stops syncing!

1 Like

You could technically revert the other version worked better though I would say give it a little time, but otherwise let us know what happens.

1 Like

Just updated from synced v5.0.0, many connections refused from peers (~50), but all works as expected.

1 Like

Was your node similar to how it is now, even before the STALL issue occurred at the end of May?

I changed nothing, only the docker image… Now it’s on v5.1.0 and all works fine (no resync from v5.0 to v5.1).

I’m still staying on v5.0.0+15. I thought about whether I need to move to v5.1, but I haven’t been able to reach a conclusion.

Why not? You can stop, update, restart…

I think I’ve got a bit of trauma. After spending ten whole days syncing zebrad, I keep having these useless imaginations that things are going to blow up again.