What Stormlog Is, How It Works, and Why We’re Building It
A visual guide to GPU memory profiling, telemetry, diagnostics, artifacts, distributed workloads, and inference profiling.
A job that was fine until it wasn’t#
A training job can look healthy for twenty minutes. Loss goes down, throughput holds, the GPU shows as busy. Meanwhile the memory footprint is quietly changing. Maybe a list somewhere keeps a reference to every step’s output. Maybe one rank in an eight-GPU job starts drifting while the other seven stay flat. Then the job dies with an out-of-memory error, and the only evidence you have is the stack trace from the moment of death.
Inference has its own version of this. An endpoint can hit its latency target with one request in flight and behave very differently with thirty-two. The quick test you ran before deploying doesn’t tell you where the curve bends, or why.
In both cases you’re debugging from a symptom. The cause happened earlier, in a place nobody was recording. Stormlog is our attempt to make that earlier behavior inspectable and reproducible: measure it while it happens, keep the evidence after the process exits, and make it possible to look at the same evidence twice.
I work on Stormlog with Silas Asamoah and Derrick Dwamena. It’s open source (MIT), written in Python, and it covers PyTorch, TensorFlow, JAX, and OpenAI-compatible inference endpoints. This page is the explanation I wanted to be able to link to: what it is, how the pieces fit, and where its edges are.
The page is long, and it’s built so you can stop early. The first sections explain the problem and the mental model with very little code. The middle covers how Stormlog is put together. The end has runnable examples, an honest list of what Stormlog isn’t, and the roadmap. If you’d rather start with commands, go to Try it.
The memory numbers that matter#
You don’t need to know CUDA to follow the rest of this page. A handful of ideas cover most of it.
- Device memory
- The GPU’s own memory. Weights, activations, gradients, optimizer state, and everything else your job puts on the card live here. It’s finite. When a request can’t be satisfied, you get an out-of-memory error (OOM).
- Allocated and reserved
- Frameworks like PyTorch don’t ask the driver for memory every time a tensor is created. Their allocator grabs larger blocks and hands out pieces. Allocated is what live tensors use right now. Reserved is what the allocator is holding, including pieces nobody is using at the moment. Reserved is always at least as large as allocated.
- Whole-device usage
- What the driver reports as used. It includes the allocator’s reserve, the runtime’s own overhead, and other processes sharing the GPU. This is the number nvidia-smi shows, and it isn’t the same as allocated.
- Peak
- The highest value in a window. An average can look calm while the peak is what runs you out of memory.
- Growth and retention
- Memory that rises step after step and doesn’t come back. Usually something still holds a reference to tensors you’re done with. That’s retention.
- Fragmentation
- Reserved memory that exists but is split into pieces too small or too scattered for the next large request. You can hit an OOM with idle memory sitting in the cache.
None of these has to look wrong until the last step. Reserved memory is allowed to stay high. Allocated memory can creep for hours. A job fails when one request for one block can’t be satisfied, and that depends on the peak and on how the free space is arranged, not on how the run looked a minute earlier.
Step through five situations below. The percentages are illustrative. They’re there to show the idea, not to describe a particular GPU.
A horizontal bar represents all device memory, split into tensors in use, idle allocator reserve, other device usage, and free space. Five states show the free space shrinking as tensor memory grows through healthy, growing, retention, near out-of-memory, and out-of-memory. A table below lists the percentages for each state.
- Tensors in use
- 34% · allocated
- Allocator reserve
- 14% · reserved, idle
- Other device usage
- 8% · outside the allocator
- Free
- 44% · unclaimed
1/5Healthy and stable
Allocated memory moves with each step and comes back down. Reserved memory sits a little above it. There is plenty of free space.
Tensors in use. Memory held by live tensors right now. Frameworks call this allocated memory. It is the number that grows when something keeps a reference alive.
Show all five states as a table
| State | Tensors in use | Allocator reserve | Other device usage | Free |
|---|---|---|---|---|
| 1. Healthy and stable | 34% | 14% | 8% | 44% |
| 2. Growing workload | 48% | 16% | 8% | 28% |
| 3. Suspicious retention | 60% | 14% | 8% | 18% |
| 4. Near out-of-memory | 70% | 18% | 9% | 3% |
| 5. Out of memory | 70% | 20% | 9% | 1% |
What Stormlog records here, for the runtimes that expose it, is the shape of the climb instead of only the final error. What each backend exposes differs, and the section on what Stormlog can see gets specific.
What Stormlog is#
Stormlog is an open-source toolkit for profiling GPU memory and inference behavior. It measures a workload, writes what it saw to local files, and gives you several ways to read those files again: Python APIs, command-line tools, and a terminal interface.
That sentence covers a lot of surface, so here is what the v0.4.0 release contains, with the caveat that matters for each part.
| Area | What’s there | Caveat |
|---|---|---|
| PyTorch | GPUMemoryProfiler for bounded profiling, MemoryTracker for tracking over time. | Bounded profiling targets torch.cuda runtimes (CUDA and ROCm). MPS and CPU use the tracker or the CPU classes. |
| TensorFlow | stormlog.tensorflow: TFMemoryProfiler, TensorFlowMemoryTracker, analyzer, visualizer. | Counters can depend on the runtime, especially on Metal. |
| JAX | stormlog.jax: JAXMemoryProfiler, tracker, analyzer, pprof-based graph view. | Device memory appears only when the runtime exposes it. Otherwise it’s marked unavailable and process memory is reported separately. |
| CPU | CPUMemoryProfiler, CPUMemoryTracker, and a CPU fallback in the CLI. | Useful for checking a workflow without a GPU. It isn’t a substitute for device counters. |
| Command line | gpumemprof, tfmemprof, jaxmemprof (each with info, monitor, track, analyze, diagnose), plus stormlog query and stormlog infer. | Same command names across frameworks, but the options differ. |
| Terminal UI | A Textual app with Overview, PyTorch, TensorFlow, Monitoring, Visualizations, Diagnostics, and CLI & Actions tabs. | Startup currently imports PyTorch, so install stormlog[tui,torch]. |
| Telemetry and artifacts | Session-aware events, append-only sinks, diagnose bundles, OOM bundles, rollups. | Inference runs use their own record format. |
| Analysis | Growth and leak heuristics, hidden-memory gap analysis, cross-rank first-cause suspects. | Heuristics. They point somewhere. They don’t prove a cause. |
| Visual exports | PNG and HTML timelines, heatmaps, dashboards. | Needs the viz extra. |
| Inference | stormlog infer profile, analyze, and collect-server for OpenAI-compatible Chat Completions endpoints. | Client-observed metrics. Server memory only with the optional host collector. |
Who this is for#
People open Stormlog for different reasons. These are the ones we design for. Each card ends with a fragment of the workflow that person would run.
ML engineer
Your training or evaluation job uses more memory than you expected, and you want to know where it goes.
profiler.profile_function(train_step)
Researcher
You’re comparing experiments, and you need enough evidence kept to reproduce a failure later.
stormlog query sessions ./artifacts --table
ML systems and infrastructure engineer
Long-running jobs, multi-GPU behavior, regressions after an upgrade.
gpumemprof track --job-id train-42 --rank 1 --world-size 8
Inference engineer
Latency, throughput, failures, token counts, workload shape, and memory under controlled load.
stormlog infer profile --arrival poisson --rate 4,8 ...
CI and release engineer
Deterministic diagnostics and machine-readable reports you can archive with a build.
gpumemprof diagnose --duration 0 --output ./diag # exit 3 on risk
Open-source contributor
Collectors, telemetry schemas, analyzers, visualization, runtime support, inference observability.
stormlog/device_collectors.py # DeviceMemoryCollector
Profiler or tracker#
Stormlog gives you two kinds of instrument, and choosing between them is mostly a matter of which question you’re asking. A profiler answers a bounded question: what happened inside this operation? A tracker answers a time-based one: what happened over time?
A profiler wraps one bounded operation, such as a function or a context block, and returns snapshots and a summary. A tracker samples a whole process on a timer and produces a stream of telemetry events, with alerts when thresholds are crossed.
bounded question
What happened inside this operation?
A function, a training step, a context block
Profiler
Snapshots and a summary
time-based question
What happened over time?
A training run, an evaluation, a long-running process
Tracker
Periodic telemetry events
In practice you often start with a tracker, because you don’t yet know where to look. Once the timeline shows when something changes, a profiler around the suspect step tells you more about that step. The two are different code paths in the package, not one tool with two modes.
import torch
from stormlog import GPUMemoryProfiler
profiler = GPUMemoryProfiler()
device = profiler.device
model = torch.nn.Linear(1024, 128).to(device)
def train_step() -> torch.Tensor:
x = torch.randn(64, 1024, device=device)
return model(x).sum()
profile = profiler.profile_function(train_step)
summary = profiler.get_summary()
print(profile.function_name)
print(f"Peak memory: {summary['peak_memory_usage'] / (1024**3):.2f} GB")from stormlog import MemoryTracker
tracker = MemoryTracker(sampling_interval=0.5, enable_alerts=True)
tracker.start_tracking()
with tracker.phase("train"):
run_training() # your workload
tracker.stop_tracking()
stats = tracker.get_statistics()
peak = stats.get("peak_memory") # None where the backend has no allocator counters
print("Peak allocated:", "unavailable" if peak is None else f"{peak / 1024**3:.2f} GB")
print(f"Events: {stats.get('total_events', 0)}")On a machine without a supported GPU, GPUMemoryProfiler and MemoryTracker raise a RuntimeError. Swap in CPUMemoryProfiler or CPUMemoryTracker to check your instrumentation on a laptop. They offer the same methods (profile_function, get_summary, phase), but the CPU profile result names the function name rather than function_name.
Both examples follow the usage guide in the Stormlog documentation. See also the Stormlog project page.
The Stormlog mental model#
Every Stormlog workflow follows one path. A workload produces measurements. A profiler, tracker, or collector reads them. They’re normalized into events that carry an identity. Those events are written as artifacts. Analysis, plots, and the terminal interface read the artifacts.
Workloads in PyTorch, TensorFlow, JAX, or behind an OpenAI-compatible endpoint are observed by profilers, trackers, and collectors. Memory and runtime measurements become canonical telemetry events with a session identity. Inference measurements become inference JSONL records with their own run identity. Both are written as local artifacts and read by the command line, the terminal interface, reports, visualizations, CI, and later investigation. The terminal interface reads the memory telemetry lane.
Your workload
- PyTorch
- TensorFlow
- JAX
- Inference endpoint
Memory and runtime
Profilers, trackers, collectors
GPUMemoryProfiler · MemoryTracker · stormlog.tensorflow · stormlog.jax
Canonical telemetry and session identity
TelemetryEventV4 · session_id · lifecycle state
Local artifacts
track exports · sink segments · diagnose bundles · OOM bundles
Inference
Workload generator and host collector
stormlog infer profile · stormlog infer collect-server
Inference records and run identity
infer.* JSONL records · run ID · workload digest
Local artifacts
inference JSONL · optional server telemetry JSONL
Where you look at it
- CLI analyze
- TUI
- stormlog query
- reports
- PNG / HTML plots
- CI exit codes
- a teammate, later
There are two lanes because inference measures different things. A training job reports memory counters from inside the process. An inference run mostly observes a server from the outside, as a client, and records requests. The lanes share the ideas that matter (local files, identity, analysis later) but not the record format. We keep that distinction visible on purpose: inference artifacts are infer.* JSONL records, not TelemetryEventV4.
Two things the diagram can’t show. The TUI isn’t a separate analysis engine: it reuses tracker sessions for live data and the same event model for artifacts. And nothing in the path needs a hosted service.
Architecture#
The code is organized as four packages, and the boundaries follow what you’d guess from the names.
stormlogholds the PyTorch profiler and tracker, the CPU fallbacks, telemetry normalization and the session contract, the analyzers, the visualizer, the device collectors, local query, the TUI, and inference.stormlog.tensorflowholds the TensorFlow profiler, tracker, analyzer, visualizer, and runtime diagnostics. It has no TUI of its own.stormlog.jaxholds the JAX profiler, tracker, diagnostics, analyzer, visualizer, and runtime helpers.stormlog.inferprofiles OpenAI-compatible endpoints. It’s separate from the framework tools because the endpoint might be backed by PyTorch, vLLM, SGLang, TensorRT-LLM, MLX-LM, or a hosted gateway.
From top to bottom: user surfaces, profiling and tracking, analysis and sessions, canonical telemetry and artifacts, and backend collectors. Each layer reads from the one below it. Open a layer for more detail.
User surfacesWhere you or a script asks a question.Python APIsgpumemproftfmemprofjaxmemprofstormlog (TUI)stormlog querystormlog infer
Four console scripts: gpumemprof, tfmemprof, jaxmemprof, and stormlog. The stormlog script opens the TUI by default and dispatches query and infer without importing Textual.
Profiling, tracking, workload generationWhat actually measures something.GPUMemoryProfilerMemoryTrackerCPUMemoryProfilerTFMemoryProfilerJAXMemoryProfilerstormlog.infer
Bounded profilers expose profile_function and profile_context. Trackers sample in the background and emit events. stormlog.infer builds a workload matrix and sends controlled traffic to an endpoint.
Analysis, sessions, correlationWhat turns samples into findings.MemoryAnalyzergap_analysisdistributed_analysissessionquerycorrelationissues
Leak and growth heuristics, hidden-memory gap analysis, cross-rank first-cause suspects, session lifecycle, and local query over artifact directories. Shared metric formulas live in derived_fields.
Canonical telemetry and artifactsThe shared shape on disk.TelemetryEventV4telemetry_sinkrollups.jsondiagnose bundleOOM bundleinference JSONLreport.json
stormlog.telemetry normalizes v2, v3 and recognized legacy records into TelemetryEventV4. Sinks write append-only JSONL segments with a manifest. Inference artifacts use their own infer.* records.
Backend collectors and framework runtimesWhere the numbers come from.CUDAROCmMPSCPU fallbackTensorFlow runtimeJAX memory_stats()NVMLpsutil
PyTorch-side device collectors implement one contract: sample(), capabilities(), name(). TensorFlow and JAX read their own runtimes. The inference host collector reads process memory with psutil and device memory with NVML.
The TUI reads the same model
The stormlog command opens a Textual app. It adapts live tracker data through a tracker session and loads saved artifacts as TelemetryEventV4 records, the same records the CLI reads. Plot export reuses the visualizer. That’s why a diagnose bundle written by a CI job can be opened in the TUI Diagnostics tab by someone else later, and why a bug fix in an analyzer shows up in both places.
Collectors declare what they can measure
On the PyTorch side, device memory comes from a collector, and every collector answers three questions: what is the current sample, what can you measure, and what are you called. CUDA, ROCm, and MPS each have one. The tracker checks every sample against the collector’s declaration. A populated field the collector said it can’t provide counts as a collector failure.
Declaring capabilities lets Stormlog leave a gap as a gap. The easy move for a profiler is to fill missing counters with zeros so charts keep drawing, but a zero is a claim: it says the allocator holds nothing. Stormlog has moved away from that more than once. Always-on tracking exports health events instead of synthetic zero samples (0.3.0). JAX statistics the runtime can’t provide are marked unavailable instead of reported as zero (0.3.8). Device-only tracking never fabricates allocator, fragmentation, history, or attribution findings (0.3.10).
See the collector contracttechnical detail
sample()returns a normalizedDeviceMemorySample. Allocator and device counters may beNone.capabilities()returns a frozenDeviceMemoryCapabilitiesdescribing each supported counter and allocator-native feature.name()identifies the backend:cuda,rocm, ormps.- A supported field may be missing from a sample only when collector diagnostics name it as partial.
- Third-party runtimes can pass a collector to
MemoryTracker(collector=...)without a torch device. No global registry is involved. Device-only collectors keep sessions, distributed identity, sinks, query, and TUI behavior. Allocator events, fragmentation, attribution, native history, and bounded profiling stay unavailable.
Source of truth: the architecture guide, which describes the code as it is, not a roadmap.
What Stormlog can see#
“Memory” is several measurements at different layers, and a backend may expose some of them and not others. It helps to keep the layers apart.
From the application at the top to the system at the bottom: application, framework, allocator, device, and system. Each layer exposes different measurements, and not every backend exposes every layer.
Application
model, phase, request
Phases you mark with tracker.phase(); request records in inference runs. Stormlog doesn't guess these.
Framework
PyTorch, TensorFlow, JAX
Snapshots around a function or context. Per-tensor tracking is opt-in (track_tensors=True on PyTorch).
Allocator
allocated, reserved, active, inactive
Counters from the framework's allocator, where the backend exposes them. Allocator history is CUDA-only and opt-in.
Device
used, free, total
Whole-device numbers from the backend collector. These include memory that isn't this process's.
System
host, process, runtime
Process memory (psutil), host and PID identity, and on inference hosts an optional collector for server process and GPU memory.
| Runtime | Allocator counters | Whole-device used, free, total | Native allocator history | Bounded profiler |
|---|---|---|---|---|
| CUDA | Yes | Yes | Yes, opt-in | GPUMemoryProfiler |
| ROCm | Yes | Yes | No, CUDA only today | GPUMemoryProfiler |
| Apple MPS | Allocated and reserved | Used; free and total when the runtime reports a maximum | No | No, use MemoryTracker |
| CPU only | Not applicable | Process and system memory | No | CPUMemoryProfiler |
| Injected device-only collector | Null, with a declared reason | As the collector declares | No | No |
That table is why “supports CUDA, ROCm, MPS, TensorFlow, and JAX” is an incomplete sentence. Support means different things on each. TensorFlow and JAX sit outside the PyTorch collector contract and read their own runtimes:
- TensorFlow runtimes can be CUDA, ROCm, Metal, or CPU. Counters can depend on the runtime on Metal.
tfmemprof infoprints build and runtime diagnostics. - JAX device memory is read through
jax.Device.memory_stats()after an XLA sync. It appears only when the runtime reportsbytes_in_use. Process memory is reported separately.
The point of the layering isn’t to apologize for gaps. It’s that an investigation can use the layers that exist. A device-only runtime still gives you a device-used timeline, sessions, and query. It just can’t give you fragmentation, and it says so.
From a running job to evidence#
A Stormlog investigation tends to have five stages. They aren’t rigid, and you’ll often loop between the middle ones.
1Instrument
Start a tracker in your script, or capture with
gpumemprof track. Mark phases if you want spikes attributed to them.2Observe
Alerts fire when a warning or critical threshold is crossed (
--warning-threshold,--critical-threshold). The TUI’s Monitoring tab shows the same tracker data live.3Diagnose
gpumemprof analyzereads the saved telemetry.gpumemprof diagnosewrites a bundle with a verdict and exits 3 when it raises a risk finding.4Preserve
Everything above already wrote files. For crashes,
--oom-flight-recorderkeeps a rolling buffer and dumps it when an OOM is recognized.5Compare or fix
Query sessions side by side with
stormlog query, or change the code and capture again. The first capture is still there to compare against.
The thing to notice is the line between stages two and three. Until the process exits, the evidence is a set of numbers in memory. After, it’s files.
A training script runs with a MemoryTracker. The tracker turns samples into canonical telemetry events that each carry a session identity. When the run ends, those events exist as files on disk: exports, sink segments, or a diagnose bundle. An analyzer can reload those files later, and the terminal interface, a report, a CI job, or a teammate can use them.
In the process
training.py
Your workload, unchanged except for one tracker.
MemoryTracker
Samples on a timer in a background thread.
TelemetryEventV4
Each sample is normalized into one canonical event.
session_id
Every event carries the identity of this capture.
On disk
track.json · sink segments · diagnose bundle
Files in a directory you chose. They survive the process.
Analyzer
Reloads the files and looks for growth, gaps, and rank drift.
TUI · report · CI · a teammate
Anyone with the files can ask the same question again.
gpumemprof track --duration 30 --interval 0.5 --output track.json --format json
gpumemprof analyze track.json --format txt --output analysis.txt
gpumemprof diagnose --duration 5 --interval 0.5 --output ./diag_bundle
echo $?Those commands run on a CPU-only host too, using the CPU fallback, which makes them a reasonable first check of your setup before you point them at a GPU.
Sessions and canonical telemetry#
Telemetry is only useful across tools if they agree on what an event is. Stormlog has one canonical event and one session contract, and the tracker, the CLI, diagnose bundles, OOM bundles, and the TUI all use them.
- session_id
- A unique ID for one capture. Every event, diagnose bundle, and OOM bundle from that capture carries or references it.
- Lifecycle state
- running, completed, interrupted, or incomplete. A clean stop marks a session completed. A process that died while running is recovered as interrupted on the next start.
- Backend identity
- Which collector produced the event and which runtime it came from, such as cuda, rocm, or mps.
- Distributed identity
- job_id, rank, local_rank, and world_size, inferred from common launcher environment variables or set explicitly.
- Normalized counters
- Allocator and device memory in bytes, with null where a backend can’t provide a counter.
- Capability metadata
- A typed object saying which counters and analyses this backend supports. It travels with the data.
The current event is TelemetryEventV4. Version 2, version 3, and recognized legacy records are upgraded to it on load, so older artifacts still open.
If you reuse one sink directory across runs, captures stay separate by session_id. Analysis defaults to the newest completed session, then the newest interrupted one, then the newest incomplete one, and you can pick another with --session-id.
See the event modeltechnical detail
{
"schema_version": 4,
"session_id": "2b30f4a4-7d2d-48f7-a9f6-7d40c14eb95e",
"timestamp_ns": 1800000000000000000,
"event_type": "sample",
"collector": "stormlog.cuda_tracker",
"sampling_interval_ms": 500,
"pid": 41873,
"host": "gpu-node-03",
"job_id": "train-42",
"rank": 2,
"local_rank": 2,
"world_size": 4,
"device_id": 2,
"allocator_allocated_bytes": 6442450944,
"allocator_reserved_bytes": 8589934592,
"device_used_bytes": 10737418240,
"device_free_bytes": 14495514624,
"device_total_bytes": 25769803776,
"context": null,
"metadata": {
"memory_capabilities": {
"backend": "cuda",
"supports_allocator_reserved": true,
"supports_device_used": true,
"supports_native_allocator_history": true
}
}
}On a device-only runtime the allocator fields are null and the capability object says why. The published JSON Schema rejects a counter that’s declared unsupported but filled in. The loader also checks what a schema can’t say, such as rank being below world_size and used plus free not exceeding total.
Artifacts that outlive the process#
Artifacts are what make the rest of this useful after the job is gone. In plain terms, an artifact is a file or directory holding what Stormlog saw, in a form something else can read.
| Artifact | What it holds | Produced by |
|---|---|---|
| Telemetry exports and sink segments | Canonical events as JSON or CSV, or append-only JSONL segments with a manifest, rollover, and retention limits. | track, trackers, TUI exports |
| Diagnose bundle | environment.json, telemetry_timeline.json, diagnostic_summary.json, manifest.json, and report.json. | diagnose |
| OOM flight-recorder bundle | Recent events from before an out-of-memory error, with a manifest, metadata, and environment. On CUDA, optional native allocator snapshots. | --oom-flight-recorder, capture_oom() |
| Visual exports | PNG and HTML timelines, heatmaps, and dashboards. | The visualizer, the TUI, analyze with --visualization |
| Inference JSONL | A session record, the workload record, request traces, phase windows, cache-state records, optional system samples. | stormlog infer profile |
| Reports | Text or JSON analysis, and a versioned report envelope with a verdict and findings. | analyze, diagnose |
| Session metadata | A session ledger in sink manifests. Bundles record the session that owns them. | Trackers and diagnose |
That gives a failure five properties it usually lacks:
- Reloadable.
gpumemprof analyze ./live_sinkreads a whole directory, not just the last run. - Shareable. Send the directory. The other person sees the same events.
- Comparable.
stormlog query summarygroups the same metric by session or by rank. - Suitable for CI. A fixed exit-code table and a
report.jsonverdict, so a job can branch without parsing text. See the report contract. - Reviewable after the process exits. An OOM bundle holds the events that led to the failure, not only the failure.
Session identity is what holds this together. Without it, a directory with five runs in it is one pile of events. With it, a diagnose bundle’s manifest names its session, an OOM bundle points back at the run that produced it, and the TUI can switch between captures instead of merging them. The artifacts article goes through the file layouts in detail, so I won’t repeat them here.
Local-first
Everything in the core workflow works on files you choose, in directories you choose. stormlog query reads them with no database. The CLI and the TUI run locally. Nothing in the path above sends profiling data to a hosted service.
Some things do leave the machine, and they’re opt-in. Weights & Biases and MLflow exports exist as optional extras. Inference profiling sends requests to the endpoint you point it at, which is the point of it. Stormlog doesn’t record the API key for an inference run, and it stores a cache-reset URL without credentials or query string.
A leak that looks harmless#
Here’s a loop where the bug is one line that looks like bookkeeping.
saved_outputs = []
for step, (x, y) in enumerate(loader):
output = model(x)
loss = criterion(output, y)
loss.backward()
optimizer.step()
optimizer.zero_grad()
saved_outputs.append(output) # looks harmlessThe fix is usually small, such as appending output.detach().cpu(), or not keeping the output at all. The hard part is noticing. A tracker running next to this loop records allocated memory on every interval, and the timeline shows the difference between the two runs immediately.
Illustrative data. The healthy run climbs for the first few steps and then stays flat near 3 gigabytes. The retention run climbs steadily every step and reaches the 10 gigabyte device capacity at step 40, where an out-of-memory error would occur. The table below the chart lists the values.
- healthy run (dashed)
- retention run (solid, heavier)
Show the chart values as a table
| Step | Healthy run (GB) | Retention run (GB) |
|---|---|---|
| 0 | 1.20 | 1.20 |
| 5 | 2.69 | 2.31 |
| 10 | 2.90 | 3.41 |
| 15 | 3.00 | 4.52 |
| 20 | 3.02 | 5.62 |
| 25 | 2.96 | 6.73 |
| 30 | 3.03 | 7.83 |
| 35 | 3.01 | 8.93 |
| 40 | 2.96 | 10.03 |
gpumemprof analyze looks at that timeline for growth. Its leak heuristic reports only consistently positive growth, which is why a healthy run that climbs and then levels off isn’t flagged.
Stormlog can tell you that memory is climbing and when it started. It can’t tell you which line holds the reference. For allocator-level evidence on CUDA there’s an opt-in native history mode that records allocator history and attaches best-effort pointer-to-tensor attribution to a diagnose or OOM bundle. The leak walkthrough runs a measured investigation end to end, with the hardware and numbers.
The rank an average hides#
One GPU number for a whole job is a sum or an average, and both can hide the interesting rank. If one of four ranks drifts, the average moves by a quarter of what that rank did. The chart below makes the point visible: toggle the average and compare it with rank 2.
Illustrative data. Ranks 0, 1 and 3 stay near 5.8 gigabytes. Rank 2 begins to climb around sample 22 and ends near 11.6 gigabytes. The average across all four ranks ends near 7.5 gigabytes, so the average shows a much smaller rise than rank 2 does. A table below lists the values.
rank 0 · rank 1 · rank 2 · rank 3
- ranks 0, 1, 3
- rank 2 (heavier)
Show the chart values as a table
| Sample | Rank 0 | Rank 1 | Rank 2 | Rank 3 | Average |
|---|---|---|---|---|---|
| 0 | 5.20 | 5.52 | 5.81 | 6.05 | 5.64 |
| 10 | 5.68 | 5.90 | 6.18 | 6.51 | 6.07 |
| 20 | 5.73 | 5.96 | 6.26 | 6.56 | 6.13 |
| 25 | 5.74 | 5.96 | 7.11 | 6.56 | 6.34 |
| 30 | 5.75 | 6.03 | 8.70 | 6.57 | 6.76 |
| 35 | 5.75 | 5.94 | 10.17 | 6.57 | 7.11 |
| 40 | 5.75 | 5.97 | 11.61 | 6.57 | 7.47 |
Four pieces make this investigation possible.
- Rank identity. Telemetry carries
job_id,rank,local_rank, andworld_size. They’re inferred from common launcher environment variables, or set with options such as--job-idand--rankongpumemprof track. - Aligned telemetry.
gpumemprof analyzemerges per-rank timelines. With--visualizationit writes a cross-rank timeline plot. - Loading all ranks at once. The TUI Diagnostics tab loads rank artifacts, keeps ranks separate, and renders per-rank timelines and a rank table.
- First-cause suspects. The analyzer ranks which rank and phase spiked first. These are ranked heuristics. When phases overlap across threads, Stormlog marks the attribution ambiguous instead of guessing.
The distributed diagnostics article walks through the rank-aware workflow in detail.
Inference: from endpoint numbers to engine evidence#
Stormlog started with training memory, and it isn’t only that anymore. Inference raises a different class of questions, and most of them are about behavior under load rather than a single peak.
- Time to first token (TTFT) and end-to-end latency
- Throughput, and how it changes with concurrency
- Token counts, and whether they came from the server or an estimate
- Failures, timeouts, and rejected requests
- Workload shape: arrival pattern, prompt lengths, shared prefixes
- Cache state, and whether the cache was cold
- Device memory while all of that happens
- What the serving engine itself was doing
The tool for it is stormlog infer. It sends controlled traffic to any endpoint that accepts an OpenAI-style Chat Completions request, which covers PyTorch servers, vLLM, SGLang, TensorRT-LLM, MLX-LM, and hosted gateways, and it records what a client can observe.
A request travels from the Stormlog client through an OpenAI-compatible endpoint to a serving engine, a server process, and finally the GPU and runtime. Stormlog v0.4.0 observes the client-visible hops and, optionally, host-side process and device memory. Serving-engine scheduler and cache signals exist for vLLM on the development branch and are not released. Other engines are planned.
Stormlog client
A controlled workload
- In v0.4.0Concurrency, input and output lengths, arrival schedule (closed, fixed-rate, Poisson, burst, replay), prompt sharing, warmup. Recorded with a workload digest.
OpenAI-compatible endpoint
What a client can observe
- In v0.4.0End-to-end latency, time to first streamed content (streaming only), token counts with their source, request outcome (ok, timeout, rejected, error, dropped, cancelled).
Serving engine
Queues, scheduler, cache
- Development branch, unreleasedvLLM Prometheus metrics and request spans, scraped during a run. Merged for the next release, vLLM only.
- PlannedSGLang, TensorRT-LLM, and TensorRT, each with its own capability matrix.
Server process
Host-side memory and identity
- In v0.4.0, optionalA collector on the serving host records process memory. The analyzer joins it to the client's case windows only when identity and clocks line up.
GPU and runtime
Device memory, later GPU activity
- In v0.4.0, optionalWhole-device or MIG memory through NVML, with the same join rules. Whole-device numbers aren't a per-request cost.
- Development branch, unreleasedRequest to scheduler-iteration to GPU-activity records, imported from a vLLM hook log. Exact attribution only where verified.
Available today
These are in v0.4.0 and documented in the inference guide:
- A workload matrix over concurrency, input length, and output length, streaming or not, with warmup requests recorded but excluded from analysis.
- Controlled arrivals: closed loop, fixed rate, seeded Poisson, bursts, and replay of a recorded trace. Open-loop runs also report latency from each request’s intended arrival, so time spent waiting for a slot stays visible.
- Prompt control: repeat, unique, or shared-prefix prompts, and a way to ask for a cold prefix cache by calling a reset URL before each case.
- A workload record with a seed and a digest, so the same workload sent to two engine configurations can be recognized as the same.
- Client-observed latency percentiles, TTFT for streaming responses, throughput, failure rate, and token counts with their source recorded on every request.
- An optional collector,
stormlog infer collect-server, that runs on the serving host and records server process memory and whole-device or MIG memory. The analyzer joins it to a case only when server identity and clocks line up, and it never sums values from different collectors. - A fixed exit-code table, so a CI job can tell “no request succeeded” apart from a usage error.
stormlog infer profile \
--base-url http://localhost:8000/v1 \
--model Qwen/Qwen2.5-7B-Instruct \
--arrival poisson --rate 2,4,8 --duration 60 \
--prompt-mode unique --seed 7 \
--input-tokens 512 --output-tokens 128 \
--output artifacts/infer_poisson.jsonl
stormlog infer analyze artifacts/infer_poisson.jsonlFrom endpoint measurements to serving-engine evidence
A client can see that latency rose. It can’t see why. The active roadmap (#210) is about closing that gap by joining five kinds of evidence:
1a request
what the client sent and saw
released
2server, process, and GPU identity
which machine and device the numbers belong to
released, optional collector
3scheduler and cache behavior
queueing, KV-cache pressure, prefix-cache hits
vLLM on the development branch
4bounded GPU activity
a trace window around the incident
in development
5an evidence-backed explanation
findings with their evidence and limits
planned
Two caveats keep this honest. A batch is shared by many requests, so Stormlog’s correlation records represent shared execution through membership. They keep measured batch duration separate from any estimated per-request cost, and don’t assign a shared kernel to one request. And an aggregate metrics scrape stays aggregate evidence. Exact request-to-iteration attribution needs worker instrumentation, which is a separate piece of work.
What exists beyond v0.4.0 is on the development branch and listed under “Unreleased” in the changelog: scraping vLLM’s Prometheus metrics during a run, receiving vLLM request spans, and importing a vLLM execution hook’s log into request, iteration, and membership records. That’s vLLM only, and it isn’t in a release yet.
Native tools, and where Stormlog stops#
Native tools are good, and you should keep using them. Different tools answer different layers of the problem.
| Tool | Strong at | How Stormlog relates |
|---|---|---|
nvidia-smi | Whole-device memory and utilization, and which processes use the GPU, right now. | Stormlog reads device counters through collectors too, but samples them over time next to allocator counters and attaches session identity. |
| PyTorch memory APIs and snapshots | Exact allocator counters and, on CUDA, allocator history. | Stormlog builds on those counters and adds sampling, alerts, artifacts, and analysis. Its native-history mode writes PyTorch’s allocator snapshots into a bundle. |
| TensorFlow and JAX profilers | Deep framework-specific inspection: ops, compilation, device memory profiles. | Stormlog gives one workflow and artifact shape across frameworks. On JAX, the OOM recorder attaches the runtime’s own device memory profile. |
| Nsight and ROCm tools | Kernel timelines and GPU-level detail. | A different layer. Stormlog doesn’t replace it. The inference roadmap looks at linking bounded traces to incidents. |
| System monitors and exporters | Host and fleet-wide metrics. | Stormlog is per-workload evidence. Ingesting DCGM readings as optional server telemetry is an open issue (#247). |
What Stormlog adds is mostly about everything around the measurement: a common workflow across frameworks, normalized telemetry, artifacts that survive the process, session-aware investigation, Python, CLI, and TUI paths to the same data, automated diagnostics, repeatable runs, CI-friendly output, and inference workload profiling. That’s a different job from reading a kernel timeline.
What Stormlog is not
- A replacement for every vendor GPU profiler.
- A hosted observability platform.
- A root-cause oracle. Its analyzers are heuristics that point at evidence, and findings carry confidence and limits.
- A guarantee that every backend exposes the same counters.
- A reason to skip the native tools for your framework or runtime.
- An LLM that guesses what went wrong. The measurements and labels don’t depend on a model.
Shipped, building, researching#
It’s easy to blur “works today”, “being built”, and “we’d like to find out”. Here they’re kept apart. The three columns are the same split the project uses, and the inference roadmap itself says its milestones describe outcomes, not release versions or dates. I’m following that: nothing below has a ship date.
Shipped
In v0.4.0 or an earlier release.
- PyTorch, TensorFlow, and JAX profilers and trackersJAX since 0.3.5.
- Sessions, TelemetryEventV4, append-only sinks, rollups
- Diagnose bundles with report.json and a fixed exit-code tableNew in 0.4.0.
- OOM flight recorder, and opt-in CUDA allocator history
- Rank-aware analysis and TUI Diagnostics
- Local query layer, including correlate and run catalogs
- Controlled arrivals, prompt modes, and cold-cache requests for inference#212New in 0.4.0.
- Host server telemetry with clock alignment and tensor-parallel groups#2140.3.10.
Building
Concrete items on the inference roadmap (#210). Some are merged but unreleased.
- vLLM metrics and request spans during a run#215On the development branch. Not in a release.
- GPU execution capture, and linking requests to scheduler iterations and GPU activity#216, #217The vLLM importer is listed as Unreleased. Capture work is on feature branches.
- SLO goodput and repeatable baseline comparisons#213
- Evidence-backed explanations of incidents#218
- Incident capture with bounded history#219, #220, #221
- SGLang, TensorRT-LLM, and TensorRT, each with its own capability matrix#222, #223, #224
- A local Web UI for inference profiling#227Design and a prototype with synthetic fixtures.
Researching
Experiments. No implementation is promised, and “reject” is an acceptable result.
- Optional native probes: CUPTI, USDT or eBPF, programmable GPU probes#118, #235
- Workload-aware incident detection against simpler baselines#225
- Per-request GPU cost estimates under shared batching#226
- A scrubbing policy for shareable artifacts#111
- Native allocator debugging on MPS and ROCm#97
- TorchTPU support, once its public runtime exists#125
Execution correlation, without double counting
The contract for relating a request to shared server iterations and GPU activity shipped in 0.3.10 (#211). It reports iteration elapsed time, summed activity duration, and merged GPU interval time as three different numbers. The rule behind it: shared execution isn’t falsely assigned to one request. Only the collector side is still being built, and the first target is vLLM.
A local Web UI
#227 describes a local UI for investigating a slow or failed inference run through its raw requests and metrics, with every displayed value traceable to a field or a calculation. It’s designed to sit on Stormlog’s existing artifacts and report contracts behind a small service boundary, not to reimplement analysis in the frontend. A prototype exists (PR #228) with explicitly synthetic fixtures. It’s a design and data-contract exercise, not something you can install.
Native probes are research
#118 asks whether an optional native collector can provide evidence, or lower collection cost, that the existing PyTorch and Nsight route can’t. Candidates include CUPTI activity collection, USDT or eBPF probes, and programmable GPU probes. None is a dependency of Stormlog, none is committed, and a negative result is a valid way for that issue to end.
Try it#
Install the package, then add extras for the runtime you use.
pip install stormlog
# Pick the extras you need:
pip install "stormlog[torch]" # PyTorch
pip install "stormlog[tf]" # TensorFlow
pip install "stormlog[jax]" # JAX
pip install "stormlog[viz]" # PNG and HTML plots
pip install "stormlog[tui,torch]" # terminal UI
pip install "stormlog[infer-tokenizers]" # better token counts for inference
pip install "stormlog[all]" # everythingThen a ten-second capture, an analysis, a diagnose bundle, and a look at what was saved. This runs on a CPU-only laptop too.
pip install "stormlog[torch]"
gpumemprof info
gpumemprof track \
--duration 10 \
--interval 0.5 \
--output run.json \
--format json
gpumemprof analyze run.json --format txt --output analysis.txt
gpumemprof diagnose --duration 0 --output ./diag
stormlog query sessions ./diag --tableFor inference, point the profiler at any OpenAI-compatible endpoint you’re allowed to send traffic to:
stormlog infer profile \
--base-url http://localhost:8000/v1 \
--model Qwen/Qwen2.5-7B-Instruct \
--concurrency 1,4,8 \
--requests 20 \
--output infer.jsonl
stormlog infer analyze infer.jsonlTo see everything in one place, pip install "stormlog[tui,torch]" and run stormlog. Docs, source, and the package page:
- Documentation, including the production cookbook and the CPU compatibility guide
- GitHub, where issues and the roadmap live
- stormlog on PyPI
- Getting started, the step-by-step version of this section
How I checked these commandstechnical detail
I ran the CPU-only sequence above, and the query command, against a fresh install of stormlog 0.4.0 from PyPI on a machine without a GPU. The inference command ran against a local stub server, so I checked its options and artifact output, not its behavior against a deployed model. I ran the tracker example through MemoryTracker with an injected device-only collector, which is how I found that peak_memory is None when a backend has no allocator counters. GPUMemoryProfiler refuses to start without a supported accelerator, so I checked its calls against the 0.4.0 source instead of running it, and I haven’t run either class on a GPU.