Skip to content

GH-33295: [C++][Python] Add list_contains compute function - #51611

Open
jonasdedden wants to merge 12 commits into
apache:mainfrom
jonasdedden:GH-33295-list-contains
Open

jonasdedden wants to merge 12 commits into
apache:mainfrom
jonasdedden:GH-33295-list-contains

Conversation

@jonasdedden

@jonasdedden jonasdedden commented Sep 28, 2026 •

Copy link
Copy Markdown
Contributor

Rationale for this change

There is no direct way to check whether each list in a list array contains a value. This PR closes #33295.

What changes are included in this PR?

A binary scalar function list_contains(lists, value) returning a boolean per list, for list, large list, list view, large list view and fixed-size list inputs.

value can be a scalar, or an array of the same length, in which case each list is searched for the value at the same index. A scalar lists paired with an array of values is broadcast.

Each list type provides its child values in list order (a slice for lists and fixed-size lists, Flatten() for list views), so each list's range follows the previous one. These values are compared once with equal, so implicit casts apply (e.g. an int8 list and an int64 value). For an array of values, each value is first repeated for its list's child values with take, and compared element-wise with the same semantics. Each list's range in the resulting bitmap is then scanned, stopping at the first match.

Semantics:

  • Null lists give null.
  • Null list values don't match a non-null value.
  • A null value matches lists that hold a null, as in Polars. Its type must still be comparable with the list value type, like a non-null value.
  • A NaN value (including float16) matches NaN list values, as in is_in and Polars (equal says NaN != NaN).

Are these changes tested?

Yes, in scalar_nested_test.cc and test_compute.py. Cases include sliced, chunked and out-of-order list-view inputs, arrays of values, scalar lists with arrays of values, lists spanning several bitmap words, null and NaN values (including float16), implicit casts, several value types, and type errors for null values.

Are there any user-facing changes?

Yes, a new compute function, pyarrow.compute.list_contains in Python.

Was AI used for this PR?

PR code and description written by:

  • Human
  • AI

Reviewed before submission by:

  • Human
  • AI
  • Not reviewed

@jonasdedden

Copy link
Copy Markdown
Contributor Author

@pitrou @zanmato1984 and/or @AlenkaF

@jonasdedden

jonasdedden commented Sep 30, 2026 •

Copy link
Copy Markdown
Contributor Author

972bc4c speeds up the per-list scan: short lists are now 2.4-4.1x faster, long lists are unchanged. The scan used to build an OptionalBinaryBitBlockCounter per list. Now it folds validity into the match bitmap once, reads a single unaligned word for lists spanning <= 64 bits, and looks up the offsets/sizes/width once per batch.

List length Match rate equal alone list_contains before (95d6b87) list_contains after (972bc4c) Speedup
1 1e-5 18 ms 346 ms 108 ms 3.2×
3 1e-5 19 ms 140 ms 48 ms 2.9×
3 30% 19 ms 196 ms 48 ms 4.1×
10 30% 19 ms 68 ms 28 ms 2.4×
100 30% 18 ms 26 ms 26 ms 1.0×
1,000 1e-5 19 ms 21 ms 20 ms 1.0×
1,000 30% 18 ms 19 ms 19 ms 1.0×
Setup and script

50M int64 values in one list<int64> array, int64 item, best of 7, best of two runs per build (runs differed by ≤ 6 ms). Release build, 8 cores, system allocator; same PyArrow, only libarrow_compute swapped between the two commits. equal alone is the comparison over the same values, a lower bound for the kernel.

"""Time pc.list_contains against the `equal` pass it is built on. Run once per build."""

import json
import sys
import time

import numpy as np
import pyarrow as pa
import pyarrow.compute as pc

N_VALUES = 50_000_000
CASES = [(1, 1e-5), (3, 1e-5), (3, 0.3), (10, 0.3), (100, 0.3), (1000, 1e-5), (1000, 0.3)]


def best_ms(func, repeats=7):
    times = []
    for _ in range(repeats):
        start = time.perf_counter()
        func()
        times.append(time.perf_counter() - start)
    return min(times) * 1e3


results = {}
for list_length, match_rate in CASES:
    rng = np.random.default_rng(42)
    values = rng.integers(2, 1 << 40, N_VALUES)
    values[rng.random(N_VALUES) < match_rate] = 1
    offsets = np.arange(0, N_VALUES + 1, list_length, dtype=np.int32)
    lists = pa.ListArray.from_arrays(offsets, pa.array(values))
    flat = lists.values
    item = pa.scalar(1, pa.int64())
    results[f"{list_length},{match_rate}"] = {
        "list_contains": best_ms(lambda: pc.list_contains(lists, item)),
        "equal": best_ms(lambda: pc.equal(flat, item)),
    }
json.dump(results, sys.stdout)

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Projects

None yet

Development

Successfully merging this pull request may close these issues.

[C++] Add a "list_contains" kernel

1 participant