Skip to content
Open
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
19 changes: 19 additions & 0 deletions README.md
Original file line number Diff line number Diff line change
Expand Up @@ -151,6 +151,25 @@ By adhering to the DB API 2.0 specification, the mssql-python module ensures com

The driver offers a suite of Pythonic enhancements that streamline database interactions, making it easier for developers to execute queries, manage connections, and handle data more efficiently.

### Text encoding

SQL statements and Python `str` parameters always use UTF-16LE. Text parameters are
bound as ODBC `SQL_C_WCHAR` on every supported platform, for both `execute()` and
`executemany()`, including calls that use `setinputsizes()`. Declaring a `VARCHAR`
SQL type does not switch to narrow C buffers; SQL Server performs the conversion
to the destination column's character set.

`Connection.setencoding()` retains requested settings for compatibility, but does
not change statement encoding or parameter binding. Requests that pass validation
but differ from `encoding="utf-16le", ctype=SQL_WCHAR` emit `UserWarning`, including
an explicitly requested or automatically selected `SQL_CHAR`. Invalid codec names,
invalid ctypes, and incompatible combinations (such as UTF-8 with `SQL_WCHAR`)
raise `ProgrammingError` before any warning is emitted or settings are stored.
`getencoding()` returns the requested settings, not the effective binding. For
example, requesting ASCII with `SQL_CHAR` does not cause non-ASCII parameters to
raise encoding errors. Use `setencoding()` with no arguments to restore the
supported defaults. `setdecoding()` independently controls how result data is read.

## Getting Started Examples
Connect to SQL Server and execute a simple query:

Expand Down
53 changes: 41 additions & 12 deletions mssql_python/connection.py
Original file line number Diff line number Diff line change
Expand Up @@ -1183,15 +1183,24 @@ def setautocommit(self, value: bool = False) -> None:

def setencoding(self, encoding: Optional[str] = None, ctype: Optional[int] = None) -> None:
"""
Sets the text encoding for SQL statements and text parameters.
Records the requested text encoding settings for compatibility.

Since Python 3 only has str (which is Unicode), this method configures
how text is encoded when sending to the database.
SQL statements and str parameters are always sent as UTF-16LE; text
parameters are bound as SQL_C_WCHAR on every platform. This applies to
execute() and executemany(), with or without setinputsizes(). This method
does not change that behavior or enforce the requested codec.

Requests that pass validation but differ from UTF-16LE with SQL_WCHAR emit
UserWarning. Invalid codec names, invalid ctypes, and incompatible
combinations (such as UTF-8 with SQL_WCHAR) raise ProgrammingError before
any warning is emitted or settings are stored. Accepted settings are still
returned by getencoding(), not the effective binding. Use setdecoding()
separately to configure how results are read.

Args:
encoding (str, optional): The encoding to use. This must be a valid Python
encoding (str, optional): The requested encoding. This must be a valid Python
encoding that converts text to bytes. If None, defaults to 'utf-16le'.
ctype (int, optional): The C data type to use when passing data:
ctype (int, optional): The requested C data type:
SQL_CHAR or SQL_WCHAR. If not provided, SQL_WCHAR is used for
UTF-16 variants (see UTF16_ENCODINGS constant). SQL_CHAR is used
for all other encodings.
Expand All @@ -1200,15 +1209,20 @@ def setencoding(self, encoding: Optional[str] = None, ctype: Optional[int] = Non
None

Raises:
ProgrammingError: If the encoding is not valid or not supported.
ProgrammingError: If the encoding or ctype is invalid, or their
combination is incompatible.
InterfaceError: If the connection is closed.

Warns:
UserWarning: If the request passes validation but its encoding or
ctype cannot be honored.

Example:
# For databases that only communicate with UTF-8
cnxn.setencoding(encoding='utf-8')
# Restore the supported default.
cnxn.setencoding()

# For explicitly using SQL_CHAR
cnxn.setencoding(encoding='utf-8', ctype=mssql_python.SQL_CHAR)
# Warns: parameters still use UTF-16LE / SQL_C_WCHAR.
cnxn.setencoding(encoding='cp1252', ctype=mssql_python.SQL_CHAR)
"""
logger.debug(
"setencoding: Configuring encoding=%s, ctype=%s",
Expand Down Expand Up @@ -1276,20 +1290,35 @@ def setencoding(self, encoding: Optional[str] = None, ctype: Optional[int] = Non
if ctype == ConstantsDDBC.SQL_WCHAR.value:
_validate_utf16_wchar_compatibility(encoding, ctype, "SQL_WCHAR")

if encoding != "utf-16le" or ctype != ConstantsDDBC.SQL_WCHAR.value:
warnings.warn(
"setencoding() does not change SQL statement encoding or text parameter binding: "
"statements and str parameters always use UTF-16LE, and text parameters are "
"bound as SQL_C_WCHAR. The requested settings are retained by getencoding() "
"for compatibility but are not applied. Use setencoding() with no arguments "
"to restore the supported defaults.",
UserWarning,
stacklevel=2,
)

# Store the encoding settings (thread-safe with lock)
with self._encoding_lock:
self._encoding_settings = {"encoding": encoding, "ctype": ctype}

# Log with sanitized values for security
logger.info(
"Text encoding set to %s with ctype %s",
"Requested text encoding stored as %s with ctype %s",
sanitize_user_input(encoding),
sanitize_user_input(str(ctype)),
)

def getencoding(self) -> Dict[str, Union[str, int]]:
"""
Gets the current text encoding settings (thread-safe).
Gets the requested text encoding settings (thread-safe).

These settings are retained for compatibility. They do not describe the
effective binding: SQL statements and str parameters always use UTF-16LE,
and text parameters are bound as SQL_C_WCHAR.

Returns:
dict: A dictionary containing 'encoding' and 'ctype' keys.
Expand Down
1 change: 1 addition & 0 deletions mssql_python/constants.py
Original file line number Diff line number Diff line change
Expand Up @@ -75,6 +75,7 @@ class ConstantsDDBC(Enum):
SQL_C_VARBINARY = -3
SQL_C_LONGVARBINARY = -4
SQL_C_LONGVARCHAR = -1
# Legacy alias: text parameters bind as ODBC SQL_C_WCHAR (-8), not SQL_C_CHAR (1).
SQL_C_CHAR = -8
SQL_C_NUMERIC = 2
SQL_C_DECIMAL = 3
Expand Down
21 changes: 10 additions & 11 deletions mssql_python/pybind/ddbc_bindings.cpp
Original file line number Diff line number Diff line change
Expand Up @@ -2136,14 +2136,10 @@ SQLRETURN SQLExecute_wrap(const SqlHandlePtr statementHandle,
(SQLPOINTER)SQL_CONCUR_READ_ONLY, 0);
}

// The encoding-settings dict has the form {"encoding": str, "ctype": int}.
// Note: the Python layer's SQL_C_CHAR constant is numerically -8, the same
// as ODBC's SQL_C_WCHAR. As a result, the only path that genuinely uses
// byte-level character encoding is when the user explicitly opts in via
// setencoding(..., ctype=mssql_python.SQL_CHAR) (which sends ctype=1, the
// real ODBC SQL_CHAR). We default to utf-8 and only honor the dict's
// encoding when ctype == 1 (real ODBC SQL_CHAR). Otherwise the user's
// "encoding" value is meant for the wide-char path and we leave it alone.
// This codec only applies to parameters already typed as real SQL_C_CHAR (1).
// Public text parameter detection uses SQL_C_WCHAR (-8), including the
// Python layer's legacy SQL_C_CHAR alias. setencoding() does not change
// paramCType and warns when the requested settings cannot be applied.
std::string charEncoding = "utf-8";
if (encoding_settings.contains("ctype") && encoding_settings.contains("encoding")) {
int ctype = encoding_settings["ctype"].cast<int>();
Expand Down Expand Up @@ -2959,10 +2955,13 @@ SQLRETURN SQLExecuteMany_wrap(const SqlHandlePtr statementHandle, const std::u16
}
LOG("SQLExecuteMany: Parameter analysis - hasDAE=%s", hasDAE ? "true" : "false");

// Extract char encoding from encodingSettings dictionary
// Match SQLExecute_wrap: a wide-char codec must never encode narrow buffers.
std::string charEncoding = "utf-8"; // default
if (encodingSettings.contains("encoding")) {
charEncoding = encodingSettings["encoding"].cast<std::string>();
if (encodingSettings.contains("ctype") && encodingSettings.contains("encoding")) {
int ctype = encodingSettings["ctype"].cast<int>();
if (ctype == SQL_C_CHAR) {
charEncoding = encodingSettings["encoding"].cast<std::string>();
Comment thread
jahnvi480 marked this conversation as resolved.
}
}

if (!hasDAE) {
Expand Down
Loading
Loading