Repository navigation
Buildup of uncollectible garbage when using Client and AsyncClient
#3276
Replies: 1 comment
|
Hiya, it's years later but if this is helpful for anyone, here's an answer: It's not uncollectible. The reason you'd use that diagnostic is to find out what the GC is doing rather than what it's not. The point is that most garbage collection gets done without the garbage collector. Python frees the memory for an object the moment the last reference to it disappears. It doesn't do this part in a GC sweep like other languages. However, there is still a cyclic garbage collector. When objects reference each other in a cycle their ref counts won't ever reach zero. Finding those is more expensive so it happens in sweeps over 3 generations and a full garbage collection won't happen at all unless certain triggers are hit. That's the kind of tricky area you would look into with Still, I'll assume you do have that kind of problem given that I did with httpx and here's what I found: This can happen when you have a long running request with with big blobs of bytes. By the time the request finishes the cyclic GC has already checked whether it can free the relevant objects containing a few times and by now it has put them in the 3rd Generation. That generation is the pile that means "Maybe this is permanent, stop frequently checking". Yes, the full garbage collection can still be triggered but only with a good amount of object creation/deletion activity. Writing/clearing a lot of bytes in an blob is a lot of memory activity and could give you an OOM but unfortunately it's not the kind of activity that triggers anything with CPython. Solutions:
I went for option 2, making shims for use with from collections.abc import AsyncIterator
from contextlib import asynccontextmanager
from typing import Any
import httpx
@asynccontextmanager
async def memory_safe_stream(method: str, url: str, *, timeout: float | httpx.Timeout, **kwargs: Any) -> AsyncIterator[httpx.Response]:
"""A shim over `httpx`'s `Response.stream` to ensure memory is freed.
Yields a response open for reading inside the block, uses a fresh client with ``timeout``.
After the block, note that `response.stream` is wiped. its status, headers and any body read with ``aread()`` stay readable.
"""
async with httpx.AsyncClient(timeout=timeout) as client, client.stream(method, url, **kwargs) as response:
try:
yield response
finally:
try:
await response.aclose()
finally:
# Response.stream links back to the Response; replacing it breaks a reference cycle so
# the object is immediately freed. Otherwise, only a full GC sweep will get it
# and it can take a lot of object churn before that's triggered.
response.stream = httpx.ByteStream(b"")
async def request(method: str, url: str, *, timeout: float | httpx.Timeout, **kwargs: Any) -> httpx.Response:
"""A drop-in shim over `httpx.request` that avoids a garbage collection problem with long-running large requests.
Sends one request on a fresh client and returns its response with the body already read.
"""
async with memory_safe_stream(method, url, timeout=timeout, **kwargs) as response:
await response.aread()
return responseThese should work as drop in replacements but don't cover every argument you might pass and clearly you'll need to adapt it if you don't want to create a fresh new client on every request. |
Uh oh!
There was an error while loading. Please reload this page.
Uh oh!
There was an error while loading. Please reload this page.
httpx version:
0.27.0Current Behavior:
When using the
ClientandAsyncClientobjects, there seems to be a buildup of uncollectible garbage from their usage. Strangely, the number of uncollectible objects identified by the GC changes depending on whether the return of the client request is assigned to a variable or not.Expected Behavior:
All objects created by the
httpxlibrary should be garbage collectible so as to avoid memory leakage.Steps To Reproduce:
Sync using response:
leads to:
Sync without using response:
leads to:
Async using response:
leads to:
Async without using response:
leads to:
Here are some forward-ref chain graphs from the second example above (sync without using response):
refs_44
refs_48
refs_76
refs_78
These graphs were made using this code:
Looping through the objects in
gc.garbageshows that they are all related tohttpxand its underlying dependencies onhttpcore&socket/_socket. I have also found it useful to use the objgraph library to render graphs of the reference dependencies of the uncollectible objects.I may be misunderstanding the action of Python's Garbage Collector as this is the first time that I'm digging into it to understand the problem I'm facing. So if there's some aspect that I may be missing then it would be great to understand that in the context of this issue!
Cheers 😁
All reactions