Skip to content

fix(pointer): harden pointer replication against thin views and floods - #240

Open
grumbach wants to merge 1 commit into
WithAutonomi:mainfrom
grumbach:fix/pointer-replication-hardening
Open

grumbach wants to merge 1 commit into
WithAutonomi:mainfrom
grumbach:fix/pointer-replication-hardening

Conversation

@grumbach

@grumbach grumbach commented Sep 29, 2026 •

Copy link
Copy Markdown
Member

Hardening for merged pointer replication (#231), from a review of it. Each change brings the code back in line with what ADR-0016 already says it does.

  • Repair on a thin view. Repair counted its quorum over however many members of the close group this node could see. A node still filling its routing table could see one, and adopt an owner-signed state nobody paid to store on that one peer's word. It now counts over the whole configured group, so members it cannot see are unanswered rather than a smaller quorum, and it does not repair at all while bootstrapping.
  • Serve floods. Every fetch and state request spawned a task before it waited for one of 32 serve permits, so nothing bounded how many waited. Requests are now admitted before a task exists, fairly across peers, as chunk fetches are: at most 128 outstanding and 16 per peer. A state request larger than an honest one is dropped before admission.
  • Honest peers under that limit. This node now keeps at most 8 pointer requests outstanding at any one peer, across repair, possession checks and pruning together. That is half the allowance a peer gives it, so its requests are never dropped as a flood and an honest peer is never judged on a request it never saw.
  • Capability set. Any sender, routed or not, was remembered as speaking pointers, with nothing bounding the set. Only routing-table peers are remembered now, and the set is capped.
  • Signature checks on the executor. A record fetched from a peer was verified on the async executor. It is now verified on the blocking pool, as ADR-0016 says every pointer signature check is.
  • Lost records. A commit for a record this node had lost accepted any state, so a write verified before the loss could roll the node back past it. It now takes the lost state or a newer one, never an older one.

Not changed: a peer asking which state this node holds is answered from the index, not the file. The module does that on purpose, so a vote costs no disk read, and a lost file is caught on its first read.

Linear issue

Closes V2-1277

Risk tier

  • T0 — docs / tooling / CI / pure UX-output. Repo CI only.
  • T1 — client-only, no network-facing behavior change. CI + prod compat smoke.
  • T2 — node/client logic with behavioral surface, no protocol/format/economics change. Dev testnet + ADR.
  • T3 — protocol / storage format / payments / routing. T2 evidence + adversarial testing.

Compatibility

  • Wire: none. The same messages, answered or dropped under load as before.
  • Storage: none.
  • API: none. Only crate-internal functions changed.

Semver impact

  • breaking
  • feature
  • fix

Test evidence

  • a_thin_view_of_the_group_cannot_adopt_on_one_vote: one visible peer can no longer decide a repair. It fails when the quorum width goes back to the peers seen.
  • a_commit_after_a_loss_takes_nothing_older_than_what_was_lost: put_bytes of an older state after a loss is Stale, the lost state itself is restored. It fails when the commit rule is reverted.
  • outbound_permits_in_use_are_never_dropped, plus a compile-time check that the per-peer serve allowance is at least twice what an honest node asks of one peer.
  • Every pointer e2e test, 12 of them over QUIC, passes: repair by neighbour sync, a late joiner catching up, a lone unpaid state never adopted, pruning, possession checks, audits. pointer_convergence 15/15.
  • Every test step CI runs, locally on macOS at this head: cargo test --lib --features test-utils 1,230 passed; e2e 110 passed, 3 ignored as on main; migration_reclaims_disk 2, migration_crash_safety 5, migration_shared_volume 5, storage_scale 2, webrtc_direct_devnet 3, poc_commitment_audit_attacks 19, poc_audit_handler_live 16, poc_bootstrap_stall 3, poc_shutdown_lmdb_drain 1, and pointer_convergence 15, all passing.
  • cargo clippy --all-targets --all-features -- -D warnings, cargo fmt --check, cargo doc with --deny=warnings pass.
  • Dev testnet: not run. That gate is the release manager's call.

New dependency

none

ADR

https://github.com/WithAutonomi/ant-node/blob/main/docs/adr/ADR-0016-pointers-immutable-owner.md

Mitigation / rollback

Revert. Nothing is persisted or changed on the wire, so a node running the previous code behaves as before.

- Repair counts its quorum over the whole configured close group, so a
  node that sees few members of it cannot adopt an owner-signed state
  nobody paid for on one peer's word, and it does not repair while it is
  still bootstrapping.
- Fetch and state requests are admitted fairly before a task exists, at
  most 128 outstanding and 16 per peer, as chunk fetches are. A state
  request larger than an honest one is dropped before admission.
- This node keeps at most 8 pointer requests outstanding at any one peer,
  across repair, possession checks and pruning, so it never exceeds the
  allowance a peer gives it and is never dropped as a flood.
- A peer is remembered as speaking pointers only while it is in the
  routing table, and that set is capped.
- A fetched record's signature is verified off the async executor, as
  ADR-0016 says every pointer signature check is.
- A commit for a record this node lost takes the lost state or a newer
  one, never an older one, so a write verified before the loss cannot
  roll the node back.
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant