Summary
ClientSession.call_tool() issues a tools/list request after every successful tools/call whenever the called tool is not in the session's output-schema cache. On a short-lived session — one ClientSession per tools/call, the pattern gateways and proxies use when the downstream caller is stateless — the cache is empty on every call, so every call_tool costs two round-trips instead of one (initialize + notifications/initialized + tools/call + tools/list = 4 POSTs on Streamable HTTP instead of 3).
When the server behind the session is itself an aggregator whose tools/list fans out to N backends, the extra request is N upstream calls, and the slowest backend's tools/list latency is added to every call_tool — including calls to tools that declare no outputSchema and have nothing to validate.
There is no way to opt out short of subclassing ClientSession or patching the method.
Where
mcp 1.29.1, src/mcp/client/session.py:
# :386-415
async def call_tool(self, name, arguments=None, read_timeout_seconds=None, progress_callback=None, *, meta=None):
...
result = await self.send_request(...)
if not result.isError:
await self._validate_tool_result(name, result)
return result
# :417-421
async def _validate_tool_result(self, name: str, result: types.CallToolResult) -> None:
"""Validate the structured content of a tool result against its output schema."""
if name not in self._tool_output_schemas:
# refresh output schema cache
await self.list_tools()
...
Still present on main @ 6affe5c0 (2026-09-16) as the public validate_tool_result, :1118-1127 — the cache is populated only by list_tools() (_absorb_tool_listing), and _tool_output_schemas is per-ClientSession, so a fresh session always pays the refresh.
Reproduction
import anyio
from mcp import ClientSession
from mcp.client.streamable_http import streamablehttp_client
async def main():
async with streamablehttp_client("http://127.0.0.1:8000/mcp") as (r, w, _):
async with ClientSession(r, w) as s:
await s.initialize()
await s.call_tool("echo", {"text": "hi"}) # server access log: POST initialize, POST initialized, POST tools/call, POST tools/list
anyio.run(main)
Any FastMCP server with a tool that returns unstructured content shows the fourth POST. Measured against a proxy whose tools/list fans out to six backends: call_tool wall time p50 ≈ 3 s / p99 ≈ 36 s through the proxy vs p99 0.24 s calling the backend directly — the gap is entirely the post-result tools/list waiting on the slowest backend.
Proposed change (any of these would do)
- A constructor opt-out, e.g.
ClientSession(..., validate_tool_results: bool = True); when False, call_tool returns the result without calling validate_tool_result. Callers that already validate structured output elsewhere (a gateway with its own schema plugin, a server that validates before responding) can turn the client-side re-validation off.
- Do not refresh on an empty cache. If the session has never listed tools,
validate_tool_result cannot know whether the tool has an outputSchema; today it spends a round-trip to find out. Skipping the refresh when not self._tool_output_schemas (and keeping it when the cache is populated but lacks the tool — a tool added since the last listing) removes the cost on per-call sessions while leaving long-lived sessions unchanged. A DEBUG log line on the skip keeps it observable.
- Validate only when the result carries
structuredContent. A result with no structuredContent from a tool with no cached schema has nothing to check; the refresh then only serves to raise RuntimeError("… has an output schema but did not return structured content") for a tool the client never listed — a stricter contract than the server side enforces.
Option 2 is what we are running as a build-time patch on a vendored 1.29.1 (one three-line hunk in _validate_tool_result); happy to open a PR for whichever shape the maintainers prefer.
Environment
mcp 1.29.1 (Python 3.12); also reproduces on main @ 6affe5c0
- Transport: Streamable HTTP, stateless server (no
Mcp-Session-Id), one ClientSession per call
Summary
ClientSession.call_tool()issues atools/listrequest after every successfultools/callwhenever the called tool is not in the session's output-schema cache. On a short-lived session — oneClientSessionpertools/call, the pattern gateways and proxies use when the downstream caller is stateless — the cache is empty on every call, so everycall_toolcosts two round-trips instead of one (initialize+notifications/initialized+tools/call+tools/list= 4 POSTs on Streamable HTTP instead of 3).When the server behind the session is itself an aggregator whose
tools/listfans out to N backends, the extra request is N upstream calls, and the slowest backend'stools/listlatency is added to everycall_tool— including calls to tools that declare nooutputSchemaand have nothing to validate.There is no way to opt out short of subclassing
ClientSessionor patching the method.Where
mcp1.29.1,src/mcp/client/session.py:Still present on
main@6affe5c0(2026-09-16) as the publicvalidate_tool_result,:1118-1127— the cache is populated only bylist_tools()(_absorb_tool_listing), and_tool_output_schemasis per-ClientSession, so a fresh session always pays the refresh.Reproduction
Any FastMCP server with a tool that returns unstructured content shows the fourth POST. Measured against a proxy whose
tools/listfans out to six backends:call_toolwall time p50 ≈ 3 s / p99 ≈ 36 s through the proxy vs p99 0.24 s calling the backend directly — the gap is entirely the post-resulttools/listwaiting on the slowest backend.Proposed change (any of these would do)
ClientSession(..., validate_tool_results: bool = True); whenFalse,call_toolreturns the result without callingvalidate_tool_result. Callers that already validate structured output elsewhere (a gateway with its own schema plugin, a server that validates before responding) can turn the client-side re-validation off.validate_tool_resultcannot know whether the tool has anoutputSchema; today it spends a round-trip to find out. Skipping the refresh whennot self._tool_output_schemas(and keeping it when the cache is populated but lacks the tool — a tool added since the last listing) removes the cost on per-call sessions while leaving long-lived sessions unchanged. A DEBUG log line on the skip keeps it observable.structuredContent. A result with nostructuredContentfrom a tool with no cached schema has nothing to check; the refresh then only serves to raiseRuntimeError("… has an output schema but did not return structured content")for a tool the client never listed — a stricter contract than the server side enforces.Option 2 is what we are running as a build-time patch on a vendored 1.29.1 (one three-line hunk in
_validate_tool_result); happy to open a PR for whichever shape the maintainers prefer.Environment
mcp1.29.1 (Python 3.12); also reproduces onmain@6affe5c0Mcp-Session-Id), oneClientSessionper call