Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
2 changes: 2 additions & 0 deletions NEWS.md
Original file line number Diff line number Diff line change
@@ -1,3 +1,5 @@
**10/01/2026:** **Behavior change:** the numeric columns of the `waterdata` OGC getters and `ngwmn` — `value`, `altitude`, `drainage_area`, and the other measurement columns — are always `float64`. They were `int64` whenever every value in a response happened to be whole, so the same column changed dtype between calls, breaking concatenation, schema checks, and typed stores (#428). **Behavior change:** a value in a numeric or datetime column that cannot be parsed still becomes `NaN`/`NaT`, but now emits a `UserWarning` naming the column and the count, so a parse failure is distinguishable from a missing value. Repeat the call with `convert_type=False` to see the raw values. Expect it from `waterdata.get_monitoring_locations()`: its `construction_date` holds partial dates (`1925`, `194805`, `19100000`) that a datetime column cannot represent, and they were already being dropped without notice — 7,306 of 24,089 values in a query for Maryland groundwater sites.

**10/01/2026:** **Bug fix:** `waterdata.get_nearest_continuous` turned a missing target (`NaT`, `None`, `NaN`, or `""`) into a `(time >= 'nan' AND time <= 'nan')` clause and sent it to the service, including when only one entry of a list was missing. It now raises `ValueError` before any request, naming how many targets are missing and the position of the first. Targets pandas cannot parse now raise a `ValueError` that names `targets`, rather than pandas' message, which named no argument and suggested a `format=` the getter does not accept. **Bug fix:** the same getter's `window` was not validated: a missing window (`None`, `'NaT'`) failed with an `AttributeError` while the filter was built, and a negative one (`'-PT5M'`) inverted every bound, so the query matched nothing and returned an empty frame indistinguishable from a gap in the data. Both now raise `ValueError` naming `window`; a zero window remains an exact-match query. **Behavior change:** a bare-number `window` now raises `TypeError`; pandas read `window=450` as 450 nanoseconds, so it matched almost nothing. Pass a duration such as `window='PT450S'`. Numeric `targets` raise `TypeError` for the same reason: `targets=[1.5e9]` meant as epoch seconds was read as 1970-01-01T00:00:01.5Z. **Behavior change:** calls that previously sent a malformed filter now fail locally, so a caller passing targets with gaps must drop or fill them first.

**09/09/2026:** **Bug fix:** code and identifier columns keep their leading zeros. A bare `pandas.read_csv` infers a zero-padded code as a number, so `waterdata.get_samples()` returned parameter code `00060` as `60` and HUC12 `070700050502` as `70700050502`, and `nwis.get_info()` returned `huc_cd` `02060005` as `2060005`. One rule now decides what a code column is — a name ending in `code`, the RDB abbreviation `_cd`, or a name containing `identifier`, `huc`, or `fips` — and every delimited response is parsed through it: the Samples and WQP CSV readers, `rdb.read_rdb` (which reads the names from the RDB header rather than the caller listing them), and the Water Use CSV pages. **Behavior change:** these columns now hold strings. `waterdata.get_samples()`: `USGSpcode`, `Location_HUCEightDigitCode`, `Location_HUCTwelveDigitCode`, `SampleCollectionMethod_Identifier` (`get_samples_summary()` shares the parse; no column in its current profile was affected). `nwis.get_info()`, `nwis.what_sites()`, and `nwis.get_record(service="site")`: `huc_cd`, `state_cd`, `county_cd`, `district_cd`. A comparison against a number — `df["USGSpcode"] == 60` — or a merge onto a numeric key now matches nothing instead of raising, so compare against the padded string (`== "00060"`) or call `.astype(int)` where the number is what you want. **Behavior change:** a count whose name reads as an identifier is numeric again. WQP's `AlternateLocation_IdentifierCount` has been read as text since 05/31/2026 because "Identifier" appears in its name; a name ending in `count` is now excluded from the rule, so the same column has one dtype in every service that reports it. Measurement columns are unchanged, and the `waterdata` OGC getters and `ngwmn` were never affected: their JSON responses deliver codes as strings and numeric coercion there is limited to a fixed list of measurement columns. **Correction to the 1.2.0 notes:** the same fix was applied to the nine `wqp` getters on 05/31/2026 and never recorded here — `wqp.get_results()` and the `what_*` getters have returned HUCs, parameter codes, and FIPS codes as strings since that release.
Expand Down
44 changes: 40 additions & 4 deletions dataretrieval/ogc/shaping.py
Original file line number Diff line number Diff line change
Expand Up @@ -12,6 +12,7 @@

import logging
import re
import warnings
from typing import Any

import httpx
Expand Down Expand Up @@ -286,17 +287,52 @@ def _type_cols(df: pd.DataFrame, dialect: OgcDialect) -> pd.DataFrame:
-------
pd.DataFrame
The DataFrame with columns cast to appropriate types.

Notes
-----
Numeric columns are always ``float64``. ``pd.to_numeric`` alone infers
``int64`` when every value happens to be whole, so the same column would
change dtype between calls (issue #428). The columns are measurements, and
``float64`` also holds the ``NaN`` that marks a missing value.

A value that cannot be parsed is still set to ``NaN``/``NaT``, so one bad
value does not fail a whole download, but a warning names the column and
the count, so a parse failure is distinguishable from a missing value.
"""
cols = set(df.columns)
for col in cols.intersection(dialect.time_cols):
df[col] = pd.to_datetime(df[col], errors="coerce")
for col in sorted(cols.intersection(dialect.time_cols)):
parsed = pd.to_datetime(df[col], errors="coerce")
_warn_unparsed(df[col], parsed, col, "datetimes")
df[col] = parsed

for col in cols.intersection(dialect.numerical_cols):
df[col] = pd.to_numeric(df[col], errors="coerce")
for col in sorted(cols.intersection(dialect.numerical_cols)):
parsed = pd.to_numeric(df[col], errors="coerce").astype("float64")
_warn_unparsed(df[col], parsed, col, "numbers")
df[col] = parsed

return df


def _warn_unparsed(raw: pd.Series, parsed: pd.Series, col: str, kind: str) -> None:
"""Warn when coercion turned values that were present into ``NaN``/``NaT``.

A value counts as present when it is neither null nor an empty string, so
missing values reported either way do not trigger the warning.
"""
if not parsed.isna().any():
return
lost = int((parsed.isna() & raw.notna() & (raw != "")).sum())
if lost:
values = "1 value" if lost == 1 else f"{lost} values"
warnings.warn(
f"{values} in column {col!r} could not be parsed as {kind} and "
"were set to missing. To inspect the raw values, repeat the call "
"with convert_type=False.",
UserWarning,
stacklevel=2,
)


def _sort_rows(df: pd.DataFrame, dialect: OgcDialect) -> pd.DataFrame:
"""
Sorts rows by the API ``dialect``'s ``sort_cols`` (in priority order).
Expand Down
13 changes: 13 additions & 0 deletions tests/waterdata_test.py
Original file line number Diff line number Diff line change
Expand Up @@ -777,6 +777,19 @@ def test_get_daily(httpx_mock):
assert hasattr(md, "url") and hasattr(md, "query_time")


def test_get_daily_value_is_float_when_every_value_is_whole(httpx_mock):
"""Issue #428: whole-number values used to infer ``int64``, so ``value``
changed dtype between calls. It is ``float64`` regardless of the data."""
body = _fixture("daily")
for feature in body["features"]:
feature["properties"]["value"] = "42"
_mock_items(httpx_mock, "daily", body=body)

df, _ = get_daily(monitoring_location_id="USGS-05427718", parameter_code="00060")

assert df["value"].dtype == "float64"


def test_get_daily_sends_date_only_time_interval(httpx_mock):
"""The Water Data dialect marks ``daily`` date-only, so an open-ended
interval goes out as ``2025-01-01/..`` with no time component."""
Expand Down
58 changes: 58 additions & 0 deletions tests/waterdata_utils_test.py
Original file line number Diff line number Diff line change
Expand Up @@ -3,6 +3,7 @@
import functools
import json
import logging
import warnings
from unittest import mock

import httpx
Expand Down Expand Up @@ -30,12 +31,14 @@
_parse_retry_after,
_raise_for_non_200,
)
from dataretrieval.ogc.policy import OgcDialect
from dataretrieval.ogc.schema import _check_ogc_requests
from dataretrieval.ogc.shaping import (
_arrange_cols,
_deal_with_empty,
_get_resp_data,
_to_snake_case,
_type_cols,
)
from dataretrieval.ogc.shaping import _finalize_ogc as _ogc_finalize
from dataretrieval.waterdata import get_stats_date_range, get_stats_por
Expand Down Expand Up @@ -855,6 +858,61 @@ def test_arrange_cols_keeps_geometry_when_present():
assert "geometry" in result.columns


# --- _type_cols --------------------------------------------------------------

_TYPING_DIALECT = OgcDialect(
time_cols=frozenset({"time"}), numerical_cols=frozenset({"value"})
)


@pytest.mark.parametrize(
"values",
[["1", "2"], [1, 2], [], [None, None]],
ids=["whole-strings", "json-integers", "empty", "all-null"],
)
def test_type_cols_numeric_columns_are_always_float(values):
"""Whole, integer, empty, and all-null columns all come back ``float64``
(#428): ``pd.to_numeric`` alone would keep the integers as ``int64``."""
df = pd.DataFrame({"value": pd.Series(values, dtype=object)})
assert _type_cols(df, _TYPING_DIALECT)["value"].dtype == "float64"


def test_type_cols_missing_values_do_not_warn():
"""Null and empty-string values are missing, not parse failures."""
df = pd.DataFrame(
{
"time": ["2024-01-01T00:00:00Z", None, ""],
"value": ["1.5", None, ""],
}
)
with warnings.catch_warnings():
warnings.simplefilter("error")
_type_cols(df, _TYPING_DIALECT)


@pytest.mark.parametrize(
("col", "raw", "kind"),
[
("value", ["1.5", None, "", "abc", "n/a"], "numbers"),
("time", ["2024-01-01T00:00:00Z", None, "", "not a date", "n/a"], "datetimes"),
],
ids=["numbers", "datetimes"],
)
def test_type_cols_counts_only_unparseable_values(col, raw, kind):
"""The warning counts values the parser dropped, not values that were
already missing, so a parse failure cannot be mistaken for a gap (#428)."""
df = pd.DataFrame({col: raw})
with pytest.warns(UserWarning, match=rf"^2 values in column '{col}'.*{kind}"):
out = _type_cols(df, _TYPING_DIALECT)
assert out[col].isna().sum() == 4


def test_type_cols_warning_is_singular_for_one_value():
df = pd.DataFrame({"value": ["1.5", "abc"]})
with pytest.warns(UserWarning, match=r"^1 value in column 'value'"):
_type_cols(df, _TYPING_DIALECT)


# --- _format_api_dates -------------------------------------------------------


Expand Down
Loading