Andrey's Blog

Linux: Processes and Resource Utilization

Table of Contents

  1. Tracking Processes
  2. Open Files: lsof
  3. Tracing: strace and ltrace
  4. Threads
  5. CPU: Time, Priority, Load
  6. Memory
  7. I/O Monitoring
  8. Per-Process Monitoring: pidstat
  9. Control Groups (cgroups)
  10. Cheat Sheet

1. Tracking Processes

ps and process states

ps aux                         # all processes, BSD style
ps -ef                         # all processes, System V style
ps -l                          # long format: includes PRI (priority) and NI (nice)
ps -o pid,ppid,stat,ni,rss,cmd -p 1234   # pick columns
pstree -p                      # process tree with PIDs

The STAT / S column shows the process state:

StateMeaning
RRunning or runnable (waiting for CPU)
SSleeping, interruptible (waiting for an event)
DUninterruptible sleep (usually waiting on disk/NFS I/O)
TStopped (e.g. Ctrl-Z, or being traced)
ZZombie: finished, but the parent hasn’t collected its exit status
IIdle kernel thread

Everything ps shows comes from /proc/<PID>/. For example, /proc/1234/status and /proc/1234/cmdline contain the process’s details.

top

top shows a live, refreshing view of the system and its busiest processes.

top - 10:42:01 up 3 days,  2:14,  1 user,  load average: 0.52, 0.61, 0.70  ◄─ uptime + load
Tasks: 312 total,   1 running, 311 sleeping,   0 stopped,   0 zombie         ◄─ process states
%Cpu(s):  4.1 us,  1.2 sy,  0.0 ni, 94.3 id,  0.3 wa,  0.0 hi,  0.1 si,  0.0 st
MiB Mem :  31842.1 total,  18210.4 free,   7120.3 used,   6511.4 buff/cache
MiB Swap:   8192.0 total,   8192.0 free,      0.0 used.  24721.8 avail Mem

    PID USER      PR  NI    VIRT    RES    SHR S  %CPU  %MEM     TIME+ COMMAND
   2817 juser     20   0 3401236 412800 118400 S   6.0   1.3  12:03.44 firefox
    │     │       │   │     │       │      │   │    │     │        │
    │     │       │   │     │       │      │   │    │     │        └─ Total CPU time used
    │     │       │   │     │       │      │   │    │     └────────── % of RAM (resident)
    │     │       │   │     │       │      │   │    └──────────────── % of one CPU since last update
    │     │       │   │     │       │      │   └───────────────────── State
    │     │       │   │     │       │      └───────────────────────── Shared memory (KiB)
    │     │       │   │     │       └──────────────────────────────── Resident memory in RAM (KiB)
    │     │       │   │     └──────────────────────────────────────── Virtual memory size (KiB)
    │     │       │   └────────────────────────────────────────────── Nice value
    │     │       └────────────────────────────────────────────────── Kernel priority
    │     └────────────────────────────────────────────────────────── Owner
    └──────────────────────────────────────────────────────────────── Process ID

CPU line fields: us user, sy kernel (system), ni niced processes, id idle, wa waiting for I/O, hi/si hardware/software interrupts, st stolen by the hypervisor (in VMs).

Interactive keys

KeyAction
SpacebarUpdate the display immediately
PSort by current CPU usage (default)
MSort by resident memory usage
TSort by total (cumulative) CPU time
uShow only one user’s processes
fChoose which fields (columns) to display
1Toggle per-CPU lines
HToggle thread view
cToggle full command line
kKill a process (asks for PID and signal)
rRenice a process
? / hHelp for all commands
qQuit
top -p 1234 -p 5678        # watch only specific PIDs
top -u juser               # only one user's processes
top -b -n 1 > snapshot.txt # batch mode: one snapshot, good for scripts

Friendlier alternatives: htop and btop.


2. Open Files: lsof

lsof lists open files and the processes using them. “Files” includes regular files, directories, libraries, devices, pipes and network sockets.

lsof                   # everything (very long, run as root to see all)
lsof +D /usr           # open files under /usr, recursively (can be slow)
lsof -p 1234           # files opened by PID 1234
lsof /var/log/messages # who has this file open?
lsof -u juser          # files opened by a user
lsof -i :22            # who is using port 22
lsof -i TCP -sTCP:LISTEN   # all listening TCP sockets
lsof +L1               # deleted files still held open (disk space not freed!)

Reading the output:

COMMAND  PID  USER   FD   TYPE DEVICE SIZE/OFF    NODE NAME
vim     4321 juser  cwd    DIR  259,2     4096  131074 /home/juser
vim     4321 juser  txt    REG  259,2  4029120  524301 /usr/bin/vim
vim     4321 juser    3u   REG  259,2    12288  131090 /home/juser/.notes.md.swp
                      │
                      └── FD: file descriptor or special role
FD valueMeaning
cwdCurrent working directory
rtdRoot directory
txtProgram executable
memMemory-mapped file (e.g. shared libraries)
0, 1, 2stdin, stdout, stderr
suffix r / w / uOpened for read / write / both

Classic use case: df says the disk is full but du can’t find the files. A process is still holding a deleted file open. Find it with lsof +L1, then restart that process.


3. Tracing: strace and ltrace

┌───────────────────────────────────────┐
│ Program                               │
│      │ calls                          │
│      ▼                                │
│ Shared libraries (libc, libssl, …)    │ ◄── ltrace watches these calls
│      │ calls                          │
└──────┼────────────────────────────────┘
       ▼  system calls (openat, read, write, …)
┌───────────────────────────────────────┐ ◄── strace watches these calls
│ Linux kernel                          │
└───────────────────────────────────────┘

strace — system calls

strace prints every system call a process makes, with arguments and return values, to stderr.

strace cat /dev/null                 # trace a new command
strace -o trace.txt cat /dev/null    # write the trace to a file
strace -f ./script.sh                # also follow child processes/threads
strace -p 1234                       # attach to a running process
strace -e trace=openat,read ls       # only certain syscalls
strace -e trace=file ls              # only file-related syscalls
strace -c ls                         # summary: count and time per syscall

Typical output line:

openat(AT_FDCWD, "/etc/ld.so.cache", O_RDONLY|O_CLOEXEC) = 3
└─┬──┘ └───────────────┬─────────────────────────────┘   └┬┘
syscall            arguments                       return value
                                              (fd 3, or -1 ENOENT on error)

Great for answering “why does this program fail?” Look for the last failing openat or connect before the error, e.g. = -1 ENOENT (No such file or directory).

ltrace — library calls

ltrace traces calls to shared library functions (e.g. malloc, strlen, printf).

ltrace ls
ltrace -c ls        # summary

4. Threads

A process can contain several threads. They share memory, open files and the PID, but each has its own thread ID (TID). In a single-threaded process, TID = PID. The first thread of any process (the main thread) also has TID = PID.

┌─────────────── Process (PID 12345) ───────────────┐
│  Shared: memory, open files, PID, user            │
│                                                   │
│  ┌──────────────┐ ┌──────────────┐ ┌────────────┐ │
│  │ Thread       │ │ Thread       │ │ Thread     │ │
│  │ TID 12345    │ │ TID 12346    │ │ TID 12347  │ │
│  │ (main)       │ │              │ │            │ │
│  │ own stack,   │ │ own stack,   │ │ own stack, │ │
│  │ registers    │ │ registers    │ │ registers  │ │
│  └──────────────┘ └──────────────┘ └────────────┘ │
└───────────────────────────────────────────────────┘
ps m                             # show threads under each process
ps m -o pid,tid,command          # with PID and TID columns
ps -eLf                          # all threads, one per line (LWP = TID)
ls /proc/12345/task/             # one directory per thread
top -H                           # threads in top

5. CPU: Time, Priority, Load

Measuring CPU time

time ls
real    0m0.004s    ◄─ wall-clock time from start to end
user    0m0.001s    ◄─ CPU time running the program's own code
sys     0m0.003s    ◄─ CPU time the kernel spent on its behalf (syscalls)

Priority and nice

The scheduler uses a nice value from -20 (highest priority, “least nice”) to 19 (lowest priority, “nicest”). The default is 0.

 -20 ◄──────────────────────── 0 ────────────────────────► 19
 most CPU, highest priority  default    least CPU, lowest priority
 (root only)                         (any user can move a process this way)
nice -n 10 ./long-build.sh     # start a command with nice 10
renice -n 19 -p 1234           # lowest priority for a running process
sudo renice -n -5 -p 1234      # raising priority requires root
ps -l                          # NI column = nice, PRI = kernel priority

Load average

uptime
#  10:42:01 up 3 days,  2:14,  1 user,  load average: 0.52, 0.61, 0.70
#                                                     └1min┘└5min┘└15min┘

The load average is the average number of processes that are runnable (state R) plus, on Linux, those in uninterruptible sleep (D, usually waiting on disk).


6. Memory

Pages and virtual memory

The kernel manages memory in pages, usually 4 KiB.

getconf PAGE_SIZE     # → 4096

Each process sees its own virtual memory. The MMU translates virtual pages to physical page frames in RAM. Pages are only loaded when first touched (demand paging).

 Process virtual memory            Physical RAM               Disk
 ┌──────────────────┐            ┌──────────────┐       ┌──────────────┐
 │ page 0 ──────────┼───────────►│ frame 7      │       │ program file │
 │ page 1 ──────────┼───────────►│ frame 2      │       │ / swap       │
 │ page 2 (not yet  │ page fault │              │ load  │              │
 │   loaded) ───────┼──────────► │ frame 9 ◄────┼───────┤              │
 └──────────────────┘            └──────────────┘       └──────────────┘

A page fault happens when a process touches a page that isn’t mapped yet:

ps -o pid,min_flt,maj_flt,cmd -p 1234
/usr/bin/time -v ls     # shows major/minor page faults and max RSS

Memory columns in ps/top:

ColumnMeaning
VIRT / VSZAll virtual memory the process has mapped, including memory not actually in RAM. Often huge and not alarming
RES / RSSResident set size: what is actually in RAM now
SHRPart of RES that is shared (e.g. libraries)

free

free -h
               total        used        free      shared  buff/cache   available
Mem:            31Gi       7.0Gi        17Gi       1.2Gi       6.4Gi        24Gi
Swap:          8.0Gi          0B       8.0Gi

vmstat — system-wide overview

vmstat 2        # a new line every 2 seconds (the first line is the average since boot)
procs -----------memory---------- ---swap-- -----io---- -system-- -------cpu-------
 r  b   swpd   free   buff  cache   si   so    bi    bo   in   cs us sy id wa st
 1  0      0 1820412  4096 6511400   0    0     5    40  812 1650  4  1 94  1  0
GroupColumnsWatch for
procsr runnable, b blocked (D state)r consistently above CPU count; b > 0 for long
memoryswpd, free, buff, cache (KiB)
swapsi swap in, so swap outConstant non-zero values = memory pressure (thrashing)
iobi blocks in, bo blocks out
systemin interrupts/s, cs context switches/s
cpuus, sy, id, wa, st (%)High wa = waiting on disk

7. I/O Monitoring

iostat — per-device statistics

iostat              # CPU + device summary since boot
iostat -x 2         # extended stats every 2 seconds
iostat -p ALL       # include partitions
iostat -d nvme0n1   # one device only

Key extended columns: r/s, w/s (requests per second), rkB/s, wkB/s (throughput), r_await/w_await (average ms per request), %util (how busy the device is).

iotop — per-process I/O

sudo iotop          # like top, but for disk I/O
sudo iotop -o       # only processes currently doing I/O
sudo iotop -a       # accumulated I/O since start

Answers “which process is hammering the disk?” when iostat shows a busy device.

Pressure Stall Information (PSI)

Modern kernels report how much time tasks were stalled waiting for a resource:

cat /proc/pressure/cpu
cat /proc/pressure/memory
cat /proc/pressure/io
# some avg10=0.00 avg60=0.12 avg300=0.08 total=123456

some is the % of time at least one task was stalled. full (for memory/io) is the % of time all non-idle tasks were stalled. This is often a better health signal than load average.


8. Per-Process Monitoring: pidstat

pidstat (from sysstat) shows statistics per process over time, like vmstat for individual processes.

pidstat 2              # CPU usage of active processes every 2 s
pidstat -p 1234 1      # one process, every second
pidstat -r -p 1234 1   # memory (faults, RSS)
pidstat -d -p 1234 1   # disk I/O
pidstat -t -p 1234 1   # per thread

Unlike top, its output is a scrolling log, which makes it easy to spot changes or save to a file.


9. Control Groups (cgroups)

What they are

cgroups let the kernel group processes and then limit, account for, and isolate their resources (CPU, memory, I/O, number of processes). systemd puts every service, user session and container into its own cgroup.

Modern systems use cgroups v2: a single unified tree mounted at /sys/fs/cgroup.

/sys/fs/cgroup                          (root cgroup)
├── system.slice                        system services
│   ├── sshd.service
│   ├── chronyd.service
│   └── k3s.service
├── user.slice                          logged-in users
│   └── user-1000.slice
│       ├── session-2.scope             your login session
│       └── user@1000.service           your user services (systemd --user)
└── machine.slice                       VMs and containers

Finding a process’s cgroup

cat /proc/self/cgroup        # cgroup of the current process (the cat itself)
# v2 output:   0::/user.slice/user-1000.slice/session-2.scope
# v1 output:   many lines, one per controller (e.g. 4:memory:/user.slice)
cat /proc/1234/cgroup        # any process

systemd-cgls                 # the whole cgroup tree with processes
systemd-cgtop                # top-like resource usage per cgroup

Controllers and interface files

Each cgroup is a directory. Its files are the interface:

FilePurpose
cgroup.procsPIDs in this cgroup. Write a PID here to move a process in
cgroup.controllersControllers available here (cpu memory io pids …)
cgroup.subtree_controlControllers enabled for child cgroups
memory.current / memory.maxCurrent usage / hard limit (OOM-kill above it)
cpu.maxCPU limit as quota period, e.g. 50000 100000 = 50% of one CPU
cpu.weightRelative CPU share when CPUs are contended (default 100)
io.maxI/O bandwidth/IOPS limits per device
pids.maxMaximum number of processes/threads
cat /sys/fs/cgroup/system.slice/sshd.service/memory.current
cat /sys/fs/cgroup/system.slice/sshd.service/cgroup.procs

Setting limits through systemd

Rather than writing the files by hand, let systemd manage them:

# Run a command in a temporary cgroup with limits
systemd-run --user --scope -p MemoryMax=500M -p CPUQuota=50% ./heavy-task

# Limit an existing service (persists as a drop-in)
sudo systemctl set-property nginx.service MemoryMax=1G CPUWeight=50

# Or in the unit file
[Service]
MemoryMax=1G
CPUQuota=200%      # 2 full CPUs
TasksMax=100

Why it matters for containers

Containers are ordinary processes placed in namespaces (what they can see) and cgroups (what they can use). For example, Kubernetes pod resources translate directly into cgroup settings:

Kubernetescgroup v2 file
resources.limits.memorymemory.max
resources.limits.cpucpu.max
resources.requests.cpucpu.weight

A container that gets OOMKilled has hit its memory.max.


10. Cheat Sheet

TaskCommand
Live process viewtop (P CPU, M memory, T total time)
Watch specific processestop -p pid1 -p pid2
Process list / treeps aux, pstree -p
Show threadsps m -o pid,tid,command, top -H
Files opened by a processlsof -p pid
Who uses a file or directorylsof /path, lsof +D /dir
Who listens on a portlsof -i :port
Deleted but open fileslsof +L1
Trace system callsstrace cmd, strace -p pid, strace -c cmd
Trace library callsltrace cmd
Time a commandtime cmd, /usr/bin/time -v cmd
Lower a process’s priorityrenice -n 19 -p pid
Start with lower prioritynice -n 10 cmd
Load averageuptime (compare with nproc)
Memory overviewfree -h (watch available)
Page sizegetconf PAGE_SIZE
System-wide CPU/memory/swap/IOvmstat 2
Disk device statsiostat -x 2
Per-process disk I/Osudo iotop -o
Per-process stats over timepidstat -p pid 1
Resource pressurecat /proc/pressure/{cpu,memory,io}
cgroup of a processcat /proc/pid/cgroup
cgroup tree / usagesystemd-cgls, systemd-cgtop
Run with resource limitssystemd-run --user --scope -p MemoryMax=500M cmd