Skip to content

fixing memory leaks in python services: diagnostics and dump collection

Python services under load can silently consume memory until the cgroup limit triggers an OOM kill. Without systematic dump collection and introspection, root-cause analysis devolves into hypothesis spinning. Below is a practical set of commands and scripts for diagnostics, heap dump collection, and leak mitigation.

1. Memory Consumption Diagnosis Commands

Basic level — psutil. Installed in one line and works without process restart.

pip install psutil

Current process consumption:

import psutil, os
p = psutil.Process(os.getpid())
info = p.memory_info()
print(f"RSS: {info.rss}  VMS: {info.vms}")

Extended structure in one command:

full = p.memory_full_info()
print(full)
# Attributes: python, rss, vms, shared, text, lib, data, dt

Key attributes psutil.Process.memory_full_info() quick reference table:

AttributeDescription
pythonMemory allocated inside the Python interpreter
rssResident Set Size — physical memory in RAM
vmsVirtual Memory Size — virtual address space
sharedShared memory (shared libraries, mmap)
text, lib, dataELF code, library, data segments

For tracking growth dynamics over a short interval, psutil in a loop or watch -n 1 psutil ... can be used, but in production tracemalloc, built into CPython, is more common.

import tracemalloc
tracemalloc.start()
# ... service work ...
snapshot = tracemalloc.take_snapshot()
top_stats = snapshot.statistics('lineno')
for stat in top_stats[:10]:
    print(stat)
tracemalloc adds a small overhead (~1–2 %). Enable it only on a staging environment or when explicit leak suspicions exist.

2. Tools for Tracking Growth and Dump Collection

When RSS begins creeping upward unnoticed, deeper inspection is required. The toolset depends on debugger availability and ptrace permissions.

gdb + gcore — classic method to extract a full process dump without stopping it (provided coredump is enabled).

# 1. Find the PID
pgrep -f your_service

# 2. Attach gdb and dump the core
gdb -p <PID>
# inside gdb:
(gdb) gcore /tmp/heap_dump.core
(gdb) quit

The resulting gcore file is binary and can be analyzed locally:

# Show segment info
gdb -ex "info files" -ex "quit" /tmp/heap_dump.core
# Or via gdb python plugins (see further)

objgraph — quick answer to “who is holding this object”.

pip install objgraph
import objgraph
# Most frequent object types in memory
objgraph.most_common_types(limit=20)
# Find reference chains
objgraph.find_backref_chains(some_object, 'owner')
objgraph works with Python objects only. For native C extensions or ctypes, gdb or valgrind are required.

tracemalloc + heapdump — Python 3.4+ native mechanism can save a heap snapshot to a file.

import tracemalloc, sys
tracemalloc.start()
# ... ...
snapshot = tracemalloc.take_snapshot()
snapshot.dump('/tmp/tracemalloc.dump')

The dump can be opened in a visualizer, but for deep analysis gdb remains more convenient.

If the service runs in a container with cap_sys_ptrace restrictions, gcore collection may require host-level access or kubectl exec with appropriate capabilities.

3. Common Anti-patterns and What to Avoid

Anti-patternConsequenceRecommendation
Ignoring gc.collect() as a “magic button”False sense of security, leak persistsUse gc.collect() to clear temporary references, not as a logical leak cure
Relying solely on __del__ for resource releaseIrregular release, circular referencesPrefer contextlib.contextmanager, try/finally, or weakref
No memory limits (cgroup/ulimit)Sharp OOM-kill without dump, context lossAlways set memory.limit in manifests and check ulimit -v
Caches without TTL or unbounded growthRapid RSS increaseUse functools.lru_cache(maxsize=N) or external stores with expiration
Accumulation of objects in global lists/modules“Death” memory in long-running processesPeriodically check list lengths, move cleanup to periodic tasks
One of the most insidious patterns — del obj inside __del__ with circular references. Python’s GC eventually collects them, but the order is not guaranteed, causing memory peaks between collections.

4. OOM Incident Response Scenario

  1. Confirmation — check the cluster event: kubectl get events -n <ns> | grep OOM or dmesg | grep out of memory. Verify the process terminated with code 137.
  2. Quick RSS glance — if the service is still alive, run one command:
python3 -c "import psutil, os; p=psutil.Process(os.getpid()); print(p.memory_info().rss // 1024, 'KB')"
  1. Dump collection — choose the method based on tool availability:
    • If gdb is available with rights: gdb -p <PID> -batch -ex gcore -ex quit /tmp/oom_dump.core
    • If a quick snapshot is needed: python3 -c "import tracemalloc; tracemalloc.start(); ..." followed by tracemalloc.take_snapshot().statistics('traceback')
  2. Analysis — use objgraph.most_common_types() or gdb commands info files, bt for call stacks. Look for unexpected dict, list, or domain class object counts.
  3. Fix — eliminate the root cause: add maxsize to caches, remove circular imports, replace global lists with bounded loops, adopt weakref for observers.
  4. Prevention — add an alert at rss > 80% of limit in Prometheus/Grafana and document the dump-collection command in the runbook.

Completing diagnostics at this stage either restores the service or gathers sufficient data for a tracker ticket with specific object types and call stacks. Here’s a thinking process:

  1. Analyze the Request:
  • Topic: Fixing memory leaks in Python services: diagnostics and dump collection
  • Role: Practicing Lead DevOps
  • Format: IT notes for a Hugo blog (Mana theme)
  • Style: Short paragraphs, concrete commands, flag tables, code with fences, no fluff, no marketing, no emojis, no author signature at end
  • Allowed: GitHub admonitions (> [!NOTE], > [!TIP], > [!WARNING])
  • Prohibited: Made-up flags, URLs, versions; uncertain facts → cautious tone; no YAML/TOML front matter; no wrapping in ```; start with a lid (2-4 sentences); sections with ##; practical commands, tables if needed; 800-1600 words; end on last substantive section; don’t repeat rules, don’t write “User wants”, "
  • Key constraint: Write the English article as a parallel original, not a word-for-word translation. Same structure and facts as the Russian draft.