Read

Module 6 · Processes and jobs

top: watching the system live

Read the three header lines of top, sort the process list with single keys, and capture a snapshot in batch mode for logs and scripts.

What you will learn

  • Interpret the memory, CPU and load-average header of `top`.
  • Sort by CPU, memory, PID or time with the interactive keys, and quit cleanly.
  • Use `-b -n` for a non-interactive snapshot and relate it to `uptime` and `/proc/loadavg`.

ps, repeated

ps is a photograph; top is a film. It redraws the screen every few seconds with the processes using the most CPU at the top, plus a header summarising memory, CPU and load. It is the first thing to open when a machine feels slow: in a few seconds you know whether the problem is one runaway process, memory pressure or something waiting on disk. It is also interactive and takes over the terminal, so remember how to leave: q, or Ctrl-C.

~% top -b -n 1 | head -8
Mem: 12244K used, 234648K free, 4K shrd, 0K buff, 2808K cached
CPU:   8% usr   8% sys   0% nic  83% idle   0% io   0% irq   0% sirq
Load average: 0.08 0.02 0.01 1/48 1001
  PID  PPID USER     STAT   VSZ %VSZ %CPU COMMAND
  999   997 root     R     1352   1%  13% top -b -n 1
    8     2 root     IW       0   0%   3% [rcu_sched]
  859     1 root     S     1356   1%   0% -/bin/sh
    1     0 root     S     1352   1%   0% init

The three header lines

Mem is the whole machine's RAM: used, free, shared, buffers and cache. Remember that *cached* memory is file data the kernel keeps around because it can, and gives back instantly; a Linux box with little "free" memory and a lot of cache is healthy, not full. CPU splits the processor's time into user code (usr), kernel code (sys), niced processes (nic), idle, waiting for I/O (io) and interrupt handling. A high io with a low usr means the disk, not the CPU, is the bottleneck.

Load average is the number of processes that wanted to run (running, or waiting for CPU or disk), averaged over the last 1, 5 and 15 minutes. Compare it with the number of CPUs: a load of 1.0 on a one-core machine like this VM means it is exactly busy; 4.0 means three processes are queueing at any moment. The last two fields, 1/48 1001, are *runnable/total processes* and the PID most recently handed out. The same numbers come from uptime and, raw, from cat /proc/loadavg.

Keys and options

Key / optionEffect
PSort by CPU usage (the default)
MSort by memory (%VSZ)
NSort by PID
TSort by CPU time consumed
RReverse the sort
q, Ctrl-CQuit
-d NRefresh every N seconds (default 5)
-n NExit after N refreshes
-bBatch mode: plain text, no screen control, for files and pipes

top -b -n 1 is the combination to remember: one refresh, plain text, back to the prompt. Pipe it into head or redirect it into a file and you have a timestamped record of what the machine was doing, which is how monitoring scripts use top. Without -b, top emits terminal escape codes to clear and redraw, which look like garbage in a file. Note that the %CPU column is computed between two refreshes, so the first screen of an interactive top can only show rough figures; wait one cycle before trusting it.

Commands in this lesson

CommandWhat it does
topLive view, refreshed every 5 s; `q` quits.
top -d 1Refresh every second.
top -b -n 1One plain-text snapshot, for files and pipes.
top -b -n 1 | head -12Header plus the top few processes.
uptimeTime, uptime and the three load averages.
cat /proc/loadavgThe raw load figures top reads.

Quiz

  1. How do you leave an interactive `top`?

    • Press `q`
    • Type `exit`
    • Press Esc
  2. What does a load average of 2.0 mean on a single-core machine?

    • The CPU is at 20%
    • On average two processes wanted to run: one running, one waiting
    • Two users are logged in
  3. Why use `top -b -n 1` instead of plain `top` in a script?

    • It is faster
    • It prints plain text once and exits, with no screen-control codes
    • It shows more processes
  4. A machine shows 200 MB free and 6 GB cached. Is it running out of memory?

    • Yes, 200 MB is almost nothing
    • No: cache is reclaimable file data the kernel keeps because the RAM was otherwise unused
    • Only if swap is enabled

Practice

  1. Capture one plain-text snapshot of `top` into `/root/lab/l53/top.txt`.

  2. Save the output of `uptime` (which includes the load averages) to `/root/lab/l53/uptime.txt`.

  3. Start an interactive `top` that refreshes every 2 seconds (then press `q` to leave it).

Open this lesson in the app to do the tasks in a real Linux machine and have them checked.