Skip to content

feat(pointer): paid mutable references with an immutable owner - #34

Merged
jacderida merged 15 commits into
mainfrom
feat/pointers-immutable-owner
Sep 27, 2026
Merged

jacderida merged 15 commits into
mainfrom
feat/pointers-immutable-owner

Conversation

@grumbach

@grumbach grumbach commented Sep 18, 2026 •

Copy link
Copy Markdown
Member

Linear issue

Closes V2-1277

Risk tier

  • T0 — docs / tooling / CI / pure UX-output. Repo CI only.
  • T1 — client-only, no network-facing behavior change. CI + prod compat smoke.
  • T2 — node/client logic with behavioral surface, no protocol/format/economics change. Dev testnet + ADR.
  • T3 — protocol / storage format / payments / routing. T2 evidence + adversarial testing.

New wire messages, a new stored record format, and a payment identity that is not the storage address. T3 on every axis it touches.

Compatibility

  • Wire: additive. Four variants appended to the end of ChunkMessageBody, so every existing discriminant keeps its value and a peer built before pointers rejects them cleanly as an unknown discriminant rather than misreading one. No existing message changed, and no protocol family is bumped.
  • Storage: additive. A new record type; no change to chunks or their addressing.
  • API: additive. New pointer module and its re-exports. Nothing removed or changed — main has never carried a pointer API, so this is a feature bump, not a breaking one.

Semver impact

  • breaking
  • feature
  • fix

2.4.0 → 2.5.0.

Test evidence

cargo test — 103 passed, 0 failed (16 pointer tests).
cargo clippy --all-targets --all-features -- -D warnings — clean.
cargo fmt --all -- --check — clean.
cargo build --no-default-features — clean.

Adversarial cases covered by test, not just by construction:

  • a tampered record fails its signature
  • an unknown version is refused, and all 256 values give distinct paid identifiers, so a version-only change cannot spend another version's receipt
  • records of different owners are not comparable, so one owner's update cannot displace another's pointer
  • a counter jump is not a paid successor — every value tried, including u64::MAX and a wrap back to 0
  • equal state never replaces, however many valid signatures are made over it
  • the merge rule is a strict total order — irreflexive, antisymmetric, transitive — so no fold over arriving replies depends on their order
  • a terminal counter still resolves equal-counter conflicts deterministically, so replicas cannot split there
  • re-encoding a parsed record is byte-identical to what was signed

New dependency

None. Display is hand-written rather than pulling in thiserror, matching error.rs.

ADR

https://github.com/WithAutonomi/ant-node/blob/feat/pointers-immutable-owner/docs/adr/ADR-0016-pointers-immutable-owner.md — pointers, immutable owner. Lands in the paired ant-node PR, WithAutonomi/ant-node#231.

Mitigation / rollback

Self-contained: one new module plus four appended enum variants. Reverting the commit removes the module and the variants; nothing existing is modified, so no data or peer is left depending on it. A peer that never sends a pointer message is unaffected either way.

grumbach and others added 11 commits September 24, 2026 11:19
A pointer is a mutable, owner-signed reference at an address derived from
its owner's public key. Public-key addressed and self-verifying: the
ML-DSA-65 key is inlined, so validating one needs nothing but the record
itself — no fetch, no cache, no second close group.

Ownership is fixed at creation. A former owner keeps its key forever, so
transferable ownership cannot be made fork-proof by a local rule: hash
tie-breaks are grindable in ~2 keygens and payment-order ties fall to a
pre-buy. Declining transfer is what lets the design be this small.

Paid per increment. A quote is paid against state_id = BLAKE3(domain ||
body), never the address: the address is stable for the pointer's life, so
paying against it would make every update after the first free — 1.0's
defect. Creation is counter 0; each update is counter + 1, so one payment
buys exactly one increment and no owner can jump the counter.

Merge is (larger counter, smaller target bytes): a total order on the
states of one address, so equivocating forks resolve identically on every
node. Equal state never replaces, whatever the signature bytes — ML-DSA
signing is randomized, so ordering record bytes would let an owner sign
one paid state repeatedly and have every submission win.

Wire messages are appended to ChunkMessageBody, so existing discriminants
keep their values and older peers reject them cleanly as unknown.
The struct carried a byte cache and a derived-state cache alongside the
owner key, so reading it told you about the implementation rather than
about a pointer. It now holds exactly what a pointer is: version, owner,
counter, target, signature.

Dropping the byte cache is safe because the encoding is fixed-width with
no optionality — there is exactly one byte sequence for a given record, so
re-encoding is always identical to what was signed. A test asserts that
across counters and target tags rather than leaving it to argument.

The derived identifiers go the same way: an address is a hash of the owner
key and a state_id is a hash of the body, so both are computed rather than
stored.
bytes_hash described a storage commitment pointers do not take part in,
and three error variants were never constructed. Deleting them keeps the
record's surface to what the feature actually uses.
known_state_id and the Unchanged response let a replica ask 'anything
newer than this?' without pulling 5 KB. Nothing asks: nodes do not
replicate pointers yet, so the only reader is a client fetching a value it
does not have. It belongs with the replication that would use it.
`merge` and `cmp_merge` restated what `replaces` already decides and had no
caller outside their own tests; `MergeRank` and `rank()` were public only to
serve them. The rule now has exactly one public spelling, and the order it
rests on is private.

The order-independence test is stronger for it: instead of merging one pair
both ways round it checks the three properties a fold actually needs —
irreflexive, never both ways, transitive — over a set of contending states.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
A pointer address was BLAKE3 of a prefix and the owner key, and a chunk address
is BLAKE3 of its content. So a chunk whose content is that prefix followed by
the key landed on exactly that owner's pointer address -- no collision needed,
just a preimage anyone can write down. Two things followed from it: an attacker
who knew a key could squat the address before its owner created the pointer, and
since state_id was built the same way, one settled quote could pay for both a
pointer and a chunk sitting at the identifier it was quoted against.

Both identities now use BLAKE3's derive-key mode, which is a different function:
no plain hash of any content can produce a value it returns. The cross-kind
refusal in the node now guards a real hash collision rather than a preimage.

Also: version is written back by the encoder rather than assumed, so a record is
re-encoded as what was signed rather than as what this build would sign; and the
version()/signature() accessors, which nothing called, are gone.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
The doc comments still said BLAKE3 over a domain prefix, which is the
construction the preimage attack used.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
…impossible

The two modes share a 32-byte range. What changed is the cost of crossing it:
a preimage under one mode for an output of the other, rather than a string
anyone can write down. The test is renamed to say what it checks -- that the
preimages the old construction handed out now miss -- since no test can prove
no content exists.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
The paid-update rule required exactly counter + 1, so a node holding one state
refused another at the same counter -- even though both were separately paid and
the merge rule says which of them wins. Deliver two such states to two nodes in
opposite orders and each keeps a different record for ever: the fork the merge
rule exists to prevent, reintroduced by the gate standing in front of it. The
convergence tests never saw it because they drive the store directly, below the
gate.

The rule is now "wins the merge, and does not skip": is_paid_update_of replaces
is_successor_of, admitting the next counter or a better target at the counter
already held, and still refusing any jump. One payment still buys one increment
-- a tie-break advances nothing and is paid for on its own state_id.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
The changelog still defined state_id as a hash of a prefix and still said every
update is exactly counter + 1. Neither is what the code does: the identities are
derive-key outputs, and a separately paid state at the counter already held is
admitted when it wins the merge rule's tie-break, which is how two concurrent
updates converge instead of splitting the group.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
@grumbach
grumbach force-pushed the feat/pointers-immutable-owner branch from b216ffb to cf51144 Compare September 24, 2026 02:26
ADR-0015 was taken by direct browser clients over WebRTC-direct while this
branch was open.
The release train cut ant-protocol 3.0.0 and moved saorsa-core, saorsa-pqc
and evmlib from git pins to published versions. ant-node's main has already
taken it and now requires ant-protocol 3.0.0, so a 2.x pointer branch cannot
resolve in that graph at all.

Same three dependency moves here, and 3.1.0 rather than 3.0.0 because the
pointer record and its wire messages are additive on top of that release. The
release itself changed no source, so this is the same code against the
published crates -- ML-DSA included, which is what the record is signed with.
The counter orders states; it does not meter them. Every stored state is
paid for against its own state_id, so a number that is skipped is never
stored and owes nothing. The reasons given for the +1 rule did not hold:
jumping ahead avoids no payment, and "stranding" a pointer at u64::MAX is
something only its owner can do to themself -- and it does not even freeze
the pointer, since equal counters still resolve by target.

What the rule did do was stop catching up. A write ends once 5 of the 7
close-group peers answer, so a peer can miss an update, and one that joins
later holds nothing. Under +1 either refused every later update for good.

PointerState::is_paid_update_of and is_genesis are removed; replaces() is
the whole rule. create() still starts at 0 and update() still adds one,
because one past the held record is the smallest counter that beats it.
@jacderida

Copy link
Copy Markdown
Member

Testnet evidence 1/2 — smoke run on 100 nodes (DEV-01, registry id 614, 2026-09-26)

Deployment-side evidence for this PR set. A dedicated pointer client tier (saorsa-deploy, V2-1339) drove ant pointer create / update / get / resolve continuously against a testnet built from these branches, with the results recorded to Postgres per operation. Full report on Linear V2-1344.

Build: ant-node and ant from feat/pointers-immutable-owner — ant-node a61abe870ffc15c2a263aa8086ff1aac65520dff, ant-client fb22bd6089759f902dee16b7a521ea49e29c37bc, ant-protocol pin 4412b6af4fab2ceb0a4e1efea1efca08e68c6e4e. Branch heads verified identical before and after the build. Fleet: 7 bootstraps + genesis, 100 nodes (DigitalOcean + Vultr, 5 services/VM), no NAT, 2 × 10 MB uploaders + 2 downloaders supplying real chunk targets, 2 pointer clients (DO lon1, DO syd1).

Workload: per client, create a new pointer (fresh ML-DSA-65 key) every 60 s, update a random owned pointer every 30 s, read back after every write, sweep every owned pointer plus the other client's published pointers at least every 600 s; 10% of updates re-point at another owned pointer and are followed with resolve. 1h warm-up, 2h measurement window.

Result — every criterion passed, all failure classes zero:

measure measurement window (2h)
Writes (create + update) 170/170 (100%) — 72 create, 98 update
Reads (get + sweep) 1,236/1,236 (100%)
Read-your-writes (get after write returned the written counter + target) 170/170
Stale reads (own or foreign counter below expected / going backwards) 0
Lost writes (update refused as "moved while in flight") 0
Anomalies (counter above owner's last write) 0
not_found on own pointers / read shortfalls 0 / 0
Chain resolve 6/6 (12/12 over the full run)
Double payments / paid-but-not-stored 0 / 0
RPC-error writes 0
Write latency p50 / p95 21.2 s / 35.0 s
Read latency p50 / p95 5.1 s / 11.8 s
Mean cost per write 0.01183 ANT + 2.58e-6 ETH (Arbitrum Sepolia)

Full run (3.1h): 303/303 writes, 1,658/1,658 reads, still all zeros.

Decay with age (the ADR-0016 question): sweep reads bucketed by pointer age at read, own and foreign, 30-minute buckets out to 150–180 min — 0.00% stale and 0 not_found in every bucket. No decay.

Node side (Elasticsearch, full run): pointer_put_rpc 2,118 — 2,003 success (p50 192 ms / p95 568 ms), 113 unchanged, 2 payment_required, 0 stale, 0 error. pointer_get_rpc 13,819 — 12,898 success, 921 not_found (individual peers not holding the address; client-side not_found was 0), 0 error. Pointer commit failed 0; chunk/pointer address collisions 0 in both directions. The only non-zero pointer WARN was does not hold pointer … it was offered (ant_node::replication::pointer): 5 events on 5 hosts over the run, no client-visible effect. 113 ant-* units across 33 hosts, 0 restarts, 0 failed units, 0 auto-upgrade events.

One observation worth a look: write latency p50 ~21 s / p95 ~35 s means a paid 5-ack write frequently outlasts the 30 s update interval, so the serial client produced ~100 writes/hour rather than the nominal 180. Nothing failed; throughput is write-latency-bound. Whether p95 ~35–38 s is inherent to a single-node-quote paid write is the open question.

@jacderida

Copy link
Copy Markdown
Member

Testnet evidence 2/2 — staging scale, 990 nodes (DEV-02, registry id 615, 2026-09-26/27)

Same branches, same SHAs and same pointer workload as the smoke run, on the full staging shape: 7 bootstraps + genesis, 990 nodes across DigitalOcean, Vultr, OVH and OVH 3-AZ (66 VMs × 15 services), NAT 10%, 10 native uploaders (20–1000 MB) + 2 downloaders, 3 WASM (browser-path) uploaders + 1 WASM downloader also built from this ant-client branch, 2 pointer clients (DO lon1, DO syd1). 1h warm-up, 4h measurement window. Full report on Linear V2-1345.

Result — every criterion passed, all failure classes zero, same as at 100 nodes:

measure measurement window (4h)
Writes (create + update) 250/250 (100%) — 112 create, 138 update
Reads (get + sweep) 2,219/2,219 (100%)
Read-your-writes 250/250
Stale reads / lost writes / anomalies 0 / 0 / 0
not_found on own pointers / read shortfalls 0 / 0
Chain resolve 13/13 (21/21 full run)
Double payments / paid-but-not-stored 0 / 0
RPC-error writes 0
Write latency p50 / p95 26.5 s / 38.2 s
Read latency p50 / p95 7.6 s / 13.3 s
Mean cost per write 0.01288 ANT + 2.64e-6 ETH

Full run (~5h): 371/371 writes, 2,545/2,545 reads, all zeros.

Decay with age: own and foreign sweep reads in 30-minute buckets out to 270–300 min — 0.00% stale, 0 not_found in every bucket. A 10x fleet and a 2x longer window changed nothing about correctness.

Scale cost, 100 → 990 nodes: write p50 +25% (21.2 → 26.5 s), p95 +9%; read p50 +48% (5.1 → 7.6 s); cost per write +8.8% ANT / +2.0% ETH. Latency only.

Node side (Elasticsearch, full run, 124M forwarded node log lines): pointer_put_rpc 2,597 — 2,480 success (p50 281 ms / p95 1.08 s), 115 unchanged, 2 payment_required, 0 stale, 0 error. pointer_get_rpc 21,024 — 19,831 success, 1,193 peer-level not_found (client-side not_found 0), 0 error. Pointer commit failed 0; address collisions 0. The only non-zero pointer WARN was again does not hold pointer … it was offered: 3 events on 3 hosts, all during warm-up, 0 in the measurement window — fewer than on the 100-node run despite a 3.3x larger fleet, so it reads as a start-up transient in replication offers rather than a steady-state defect. 1,015 ant-* units on 91 hosts, 0 restarts, 0 failed units, 0 (deleted) exes, 0 auto-upgrade events across ten checks.

Alongside (not pointer-related): native transfers 4,517/4,517 uploads and 796/796 downloads; node CPU median 31.8%, per-service RSS p95 459 MB across 997 services. WASM downloads 568/569; WASM uploads failed almost only at 900 MB (8 ok / 145 fail, V2-1305 quote-collection livelock) while the 300 MB and 20 MB WASM uploaders ran at 93% and 99%.

Carried forward: the write-latency observation from the smoke run is confirmed and slightly worse at scale — at p50 26.5 s against a 30 s update interval the serial client produced ~74 writes/hour, not 180. If a target write count matters for a staging train it has to be derived from measured write latency at fleet size.

@jacderida
jacderida merged commit 5ddf576 into main Sep 27, 2026
12 checks passed

@mickvandijke mickvandijke left a comment

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Really clean piece of work. The hand-rolled fixed-width encoding keeps the signed path free of any serializer, deriving both identities with BLAKE3 derive-key separates the pointer and chunk namespaces properly, and "equal state never replaces" neatly closes the re-signing replay. The tests that check the merge rule is a strict total order are a nice touch. Thanks Anselme!

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants