# fixing memory leaks in python services: diagnostics and dump collection

LLMS index: [llms.txt](/en/llms.txt)

---

Python services under load can silently consume memory until the cgroup limit triggers an OOM kill. Without systematic dump collection and introspection, root-cause analysis devolves into hypothesis spinning. Below is a practical set of commands and scripts for diagnostics, heap dump collection, and leak mitigation.

## 1. Memory Consumption Diagnosis Commands

Basic level — `psutil`. Installed in one line and works without process restart.

```bash
pip install psutil
```

Current process consumption:

```python
import psutil, os
p = psutil.Process(os.getpid())
info = p.memory_info()
print(f"RSS: {info.rss}  VMS: {info.vms}")
```

Extended structure in one command:

```python
full = p.memory_full_info()
print(full)
# Attributes: python, rss, vms, shared, text, lib, data, dt
```

Key attributes `psutil.Process.memory_full_info()` quick reference table:

| Attribute | Description |
|-----------|-------------|
| `python` | Memory allocated inside the Python interpreter |
| `rss` | Resident Set Size — physical memory in RAM |
| `vms` | Virtual Memory Size — virtual address space |
| `shared` | Shared memory (shared libraries, mmap) |
| `text`, `lib`, `data` | ELF code, library, data segments |

For tracking growth dynamics over a short interval, `psutil` in a loop or `watch -n 1 psutil ...` can be used, but in production `tracemalloc`, built into CPython, is more common.

```python
import tracemalloc
tracemalloc.start()
# ... service work ...
snapshot = tracemalloc.take_snapshot()
top_stats = snapshot.statistics('lineno')
for stat in top_stats[:10]:
    print(stat)
```

> [!TIP] `tracemalloc` adds a small overhead (~1–2 %). Enable it only on a staging environment or when explicit leak suspicions exist.

## 2. Tools for Tracking Growth and Dump Collection

When RSS begins creeping upward unnoticed, deeper inspection is required. The toolset depends on debugger availability and `ptrace` permissions.

**gdb + gcore** — classic method to extract a full process dump without stopping it (provided `coredump` is enabled).

```bash
# 1. Find the PID
pgrep -f your_service

# 2. Attach gdb and dump the core
gdb -p <PID>
# inside gdb:
(gdb) gcore /tmp/heap_dump.core
(gdb) quit
```

The resulting `gcore` file is binary and can be analyzed locally:

```bash
# Show segment info
gdb -ex "info files" -ex "quit" /tmp/heap_dump.core
# Or via gdb python plugins (see further)
```

**objgraph** — quick answer to "who is holding this object".

```bash
pip install objgraph
```

```python
import objgraph
# Most frequent object types in memory
objgraph.most_common_types(limit=20)
# Find reference chains
objgraph.find_backref_chains(some_object, 'owner')
```

> [!WARNING] `objgraph` works with Python objects only. For native C extensions or ctypes, `gdb` or `valgrind` are required.

**tracemalloc + heapdump** — Python 3.4+ native mechanism can save a heap snapshot to a file.

```python
import tracemalloc, sys
tracemalloc.start()
# ... ...
snapshot = tracemalloc.take_snapshot()
snapshot.dump('/tmp/tracemalloc.dump')
```

The dump can be opened in a visualizer, but for deep analysis `gdb` remains more convenient.

> [!NOTE] If the service runs in a container with `cap_sys_ptrace` restrictions, `gcore` collection may require host-level access or `kubectl exec` with appropriate capabilities.

## 3. Common Anti-patterns and What to Avoid

| Anti-pattern | Consequence | Recommendation |
|--------------|-------------|----------------|
| Ignoring `gc.collect()` as a "magic button" | False sense of security, leak persists | Use `gc.collect()` to clear temporary references, not as a logical leak cure |
| Relying solely on `__del__` for resource release | Irregular release, circular references | Prefer `contextlib.contextmanager`, `try/finally`, or `weakref` |
| No memory limits (cgroup/`ulimit`) | Sharp OOM-kill without dump, context loss | Always set `memory.limit` in manifests and check `ulimit -v` |
| Caches without TTL or unbounded growth | Rapid RSS increase | Use `functools.lru_cache(maxsize=N)` or external stores with expiration |
| Accumulation of objects in global lists/modules | "Death" memory in long-running processes | Periodically check list lengths, move cleanup to periodic tasks |

> [!WARNING] One of the most insidious patterns — `del obj` inside `__del__` with circular references. Python's GC eventually collects them, but the order is not guaranteed, causing memory peaks between collections.

## 4. OOM Incident Response Scenario

1. **Confirmation** — check the cluster event: `kubectl get events -n <ns> | grep OOM` or `dmesg | grep out of memory`. Verify the process terminated with code 137.
2. **Quick RSS glance** — if the service is still alive, run one command:

```bash
python3 -c "import psutil, os; p=psutil.Process(os.getpid()); print(p.memory_info().rss // 1024, 'KB')"
```

3. **Dump collection** — choose the method based on tool availability:
   - If `gdb` is available with rights: `gdb -p <PID> -batch -ex gcore -ex quit /tmp/oom_dump.core`
   - If a quick snapshot is needed: `python3 -c "import tracemalloc; tracemalloc.start(); ..."` followed by `tracemalloc.take_snapshot().statistics('traceback')`
4. **Analysis** — use `objgraph.most_common_types()` or `gdb` commands `info files`, `bt` for call stacks. Look for unexpected `dict`, `list`, or domain class object counts.
5. **Fix** — eliminate the root cause: add `maxsize` to caches, remove circular imports, replace global lists with bounded loops, adopt `weakref` for observers.
6. **Prevention** — add an alert at `rss > 80% of limit` in Prometheus/Grafana and document the dump-collection command in the runbook.

Completing diagnostics at this stage either restores the service or gathers sufficient data for a tracker ticket with specific object types and call stacks.
Here's a thinking process:

1.  **Analyze the Request:**
   - **Topic:** Fixing memory leaks in Python services: diagnostics and dump collection
   - **Role:** Practicing Lead DevOps
   - **Format:** IT notes for a Hugo blog (Mana theme)
   - **Style:** Short paragraphs, concrete commands, flag tables, code with fences, no fluff, no marketing, no emojis, no author signature at end
   - **Allowed:** GitHub admonitions (`> [!NOTE]`, `> [!TIP]`, `> [!WARNING]`)
   - **Prohibited:** Made-up flags, URLs, versions; uncertain facts → cautious tone; no YAML/TOML front matter; no wrapping in ```; start with a lid (2-4 sentences); sections with `##`; practical commands, tables if needed; 800-1600 words; end on last substantive section; don't repeat rules, don't write "User wants", "
   - **Key constraint:** Write the English article as a parallel original, not a word-for-word translation. Same structure and facts as the Russian draft.
