Skip to content

pgwire extended protocol collapses duplicate output column names to one cell #337

Description

@farhan-syah

Version / build tested against

origin/main @ d3d73be

Deployment mode

Origin — single node (local)

Engine(s) involved

Not engine-specific / unsure

Summary

Over the extended-query protocol (Parse/Bind/Execute), a SELECT with two output columns of the same name renders one cell in both columns. The Describe phase rebuilds the result projection from the announced PG fields with lookup_key = display_name = field name (nodedb/src/control/server/pgwire/handler/prepared/execute.rs, the OutputSchema built from stmt.result_fields). That projection overrides the planner's OutputSchema at nodedb/src/control/server/pgwire/handler/routing/execute.rs (shaping.projection.or(Some(&output_schema))). The shaper then reads both cells through the same key. Every producer stores cells under the unique keys from cell_keys (response_shape/project.rs), so the second column's cell (<name>_1) is never read. The simple-query protocol uses the planner schema and is correct after #327.

Found by static review of #327. Derived from the code, not reproduced on a running server.

Steps to reproduce

-- Run through a prepared statement (tokio_postgres `client.query`, psycopg `execute`,
-- any driver on Parse/Bind/Execute). `psql` simple-query mode does not reproduce.
CREATE SEQUENCE s START 1 INCREMENT 1;
SELECT nextval('s'), nextval('s');   -- expected (1, 2); returns (1, 1) after #327, (2, 2) before

-- Same class, join with a repeated bare name:
CREATE COLLECTION w (id TEXT PRIMARY KEY, b_id TEXT) WITH (engine='document_strict');
CREATE COLLECTION b (id TEXT PRIMARY KEY) WITH (engine='document_strict');
INSERT INTO b (id) VALUES ('b1');
INSERT INTO w (id, b_id) VALUES ('w1', 'b1');
SELECT w.id, b.id FROM w JOIN b ON w.b_id = b.id;   -- expected (w1, b1); both columns read one cell

Expected behavior

Each output column renders its own cell on every protocol. The lookup keys come from one place: the planner's build_output_schema, which already derives per-column unique keys via cell_keys. The Describe-built projection must not replace those keys. It can supply display names, types, and result formats only.

Actual behavior

On Bind/Execute the duplicate columns both read the first stored cell (nextval → 1, 1). Stored data is intact. Simple-query protocol returns the correct row.

What actually happened? (check all that are true)

  • Acknowledged/committed data was lost, corrupted, or silently wrong
  • The server crashed, hung, or failed to start
  • A security or isolation boundary was crossed
  • Core functionality is broken with no acceptable workaround
  • A workaround exists (rewrite the query, avoid one path, etc.)

Workarounds: alias each column to a distinct name, or use the simple-query protocol.

Proposed severity

SEV-2 — High: major functionality broken or silently-wrong results; stored data intact

Reproducibility

Always — every attempt

Last known-good version / commit (if a regression)

Never worked. The override predates cell_keys.

Environment & logs

Linux x86_64. No server log output: the wrong row is returned without error.

Before submitting

  • I searched existing issues and this is not a duplicate.
  • I reproduced this on a released tag or a current main build (not a stale local branch). — Derived from code reading at d3d73bed, not run.
  • This is not a security vulnerability (those go to a private advisory).

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

Labels

area:pgwirePostgreSQL wire protocol / client compatsev:2-highMajor functionality broken; no acceptable workaroundstatus:in-progressActively being worked ontype:bugA defect — broken, incorrect, or lost data

Type

No type

Projects

No projects

    Milestone

    No milestone

    Relationships

    None yet

    Development

    No branches or pull requests

    Issue actions