Conversation
|
Before you submit for review:
If you did not complete any of these, then please explain below. |
PQ benchmark results — performance not clearedCompleted 12 fresh PQ builds (two per arm per dataset), 12 retained-index query retests, and four separate CAP JFR construction profiles. Baseline Configuration: JDK 23, Panama 512-bit SIMD, verified Means below; QPS is from the separate retained-index retest with sequential warming, verified initial residency, three real-query warmup passes, and three samples of at least 10 seconds. Initial residency does not guarantee residency throughout execution.
CorrectnessEvery fixed graph has zero duplicate neighbor IDs at every level. Baseline extra duplicate entries per build: Ada002 180,869 / 180,227; CAP 199,123 / 199,720; Cohere 770,411 / 769,181. The 27 focused regression tests passed. Recall differences are small, but two builds do not establish statistical significance. Timing limitations and profilingDo not interpret construction means as an isolated deduplication cost. Raw build times (seconds): Ada main 249.24/916.22, fix 201.76/886.65; CAP main 316.40/241.58, fix 645.70/886.56; Cohere main 187.54/172.90, fix 784.38/673.22. All outliers are retained. Query timing also varies across JVMs despite warming. Separate CAP JFR runs locate most slow-run time in insertion. Insertion times main/fix/fix/main were 1337.00/1277.67/227.96/869.51 seconds (diagnostic timings, not pooled above). Slow runs on both branches heavily sample Recommendation: keep this PR draft. Correctness is demonstrated; performance is not cleared. Investigate the shared PQ compilation instability separately before requesting readiness. No ASH, reader implementation, query scorer, or format changes are included in this PR. |
CAP-100k construction diagnosis: shared PQ/JIT instabilityFollow-up on the construction-time variability above. Tested exactly the first 100,000 CAP vectors, with the same shared PQ codebook, graph parameters, JDK23, 48 construction workers, Panama512, and verified MemorySegmentReader. No recall measurement for this prefix. Production artifacts remain unchanged. Uninstrumented insertion times in main/fix/fix/main order: 15.84 / 55.53 / 21.09 / 83.05 seconds. Both branches exhibit the problem. Four separate JFR/compiler-log runs gave insertion 57.82 / 21.48 / 19.71 / 21.86 seconds. All reached C2 early. The slow kernel lacks Separate timing instrumentation measured 1.251B vs 1.258B insertion diversity comparisons (+0.59%), but sampled cost rose from 670 ns to 2,869 ns per comparison. Search scoring did not slow down. Wrappers can change JIT decisions; these are diagnostic measurements, not isolated branch-overhead estimates. Allocation is a major part of this: GC logs show approximately 5.21 TB cumulatively reclaimed in the slow-main profiled JVM versus 12.9 GB in fast main (rounded before/after heap sizes, whole runs, not resident memory). JFR allocation stacks identify temporary vector objects and backing arrays in PQ conversions/gather/add operations. Short GC pauses alone substantially understate the allocation cost. Controls, again main/fix/fix/main insertion seconds:
During the observed 164.79-second insertion, system I/O pressure was only 0.007 seconds over the 162.36-second monitored window, system iowait 0.006%, and no Java D-state threads were sampled. This argues against disk stalls as the main cause of that run. Separate writeInline contention and one 24.59-second flush outlier remain documented; they are not being discarded. Keep this PR draft. The shared PQ compilation/allocation problem is established, but the exact trigger for the default-mode store-intrinsic failure and a production remedy remain unresolved. Stabilize that scorer before interpreting the large 1M build differences as deduplication overhead. Neither diagnostic JVM flag is a recommended production setting. |
Summary
Follow-up to #522. This takes a narrower approach on current main, preserving score-sorted neighbors and the existing diverse-prefix optimization rather than changing storage to ID order.
No scorer, query algorithm, or disk-format changes. Existing graphs are not repaired on load; normal candidate merges may deduplicate individual neighborhoods. Different IDs with identical vectors remain distinct.
Recap
Duplicate node IDs show up in graphs built using asymmetric scoring because A dot quant(B) /= B dot quant(A) and the current code only rejects nodes with BOTH the same node ID and score. This damages search due to redundant search paths at all levels in the graph. This does not affect graph indexes built with full precision.
Validation
27 focused tests passed on JDK23 with the jdk20 Maven profile: TestNodeArray, TestNeighbors, GraphIndexBuilderTest. Coverage includes different-score duplicates, full arrays, merge tails, tie ordering, concurrent offers, degrees 2–2048, duplicate rejection without pruning, and full-precision graph construction. Replacement is tested to evict the lowest score rather than the highest ID.
Performance — pending
Draft pending a BenchYAML comparison of frozen main e31aa59 versus this fix: FP and PQ on Ada002-1M, CAP-1M, Cohere-1M; repeated fresh graphs in balanced order with shared PQ compressors; construction time, single-thread QPS, visited count, recall@10 at 1x/2x, and duplicate audits at every level. Both branches will use identical benchmark harness code. Material regressions will be investigated and reported before requesting readiness.