WSGITransport PATH_INFO changes UTF-8 bytes and breaks URL reconstruction #3798
Replies: 2 comments
|
Yes, this is a confirmed bug in WSGITransport. The WSGI spec (PEP 3333) requires PATH_INFO to be a native string with the URL-decoded path. httpx currently assigns it from request.url.path, which is URL-decoded, so far so good, but only for ASCII-safe paths. The real problem is that the WSGI transport never uses the raw request target at all; it reconstructs the WSGI environ from URL.path/URL.query, and on Python 3 the decoded string is then encoded by the WSGI server/app boundary using environ["PATH_INFO"].encode("latin-1") for URL reconstruction. When the path contains UTF-8 bytes that were percent-encoded, the decoded string round-trips through latin-1 and produces the wrong bytes, breaking URL reconstruction. The fix is for WSGITransport to expose the original percent-encoded path, or at least to avoid the double-decode/encode round-trip. The adjacent ASGITransport already does the right thing: it passes request.url.raw_path.split(b"?")[0] as raw_path. WSGITransport has no equivalent handling and simply assigns request.url.path to PATH_INFO. |
|
Independently reproduced on Linux / Python 3.12.14 at master I checked eight paths, including literal and escaped UTF-8, One clarification to the previous reply: PATH_INFO should not retain the percent-encoded spelling. In this representation I also checked an adjacent boundary: with These are focused in-memory checks, not a full test-suite run or approval of your unpublished patch. |
Uh oh!
There was an error while loading. Please reload this page.
Reproduced on current
master(b5addb64) with Python 3.12.13 on Windows and httpcore 1.0.9. No server or third-party WSGI framework is needed:Actual output:
Expected output:
WSGITransportcurrently setsPATH_INFOtorequest.url.path, which is decoded Unicode. PEP 3333 instead requires metadata strings to preserve bytes through Latin-1. That distinction changes the bytes oféand makes the standard-library URL reconstruction fail for中.This is related to #1365/#1391, but unquoting alone doesn't preserve that WSGI byte representation. Taking the raw path without its query, percent-decoding to bytes, then decoding with Latin-1 fixes these cases without changing HTTPX's public
URL.pathbehavior. It also preserves non-UTF-8 escaped bytes such as%FFrather than replacing them.I have a small local patch with 10 regression cases covering Unicode, escaped Unicode, non-UTF-8 bytes, encoded slashes, double escapes and query separation. The original fails 5 of those cases; the patch passes all 22 WSGI tests and all 136 tests in the combined WSGI/ASGI/URL suites. Ruff also passes. Would you be open to a PR for this?
All reactions