awk: columns and simple reports
Splitting lines into fields, picking columns, filtering rows by value and adding up totals with awk, then a review of the whole Finding things module.
What you will learn
- Print chosen fields with `$1`, `$NF` and `-F` for other separators.
- Filter rows with patterns and comparisons, and use `NR`, `NF`, `BEGIN` and `END`.
- Choose between grep, find, sed and awk for a given job.
grep picks lines and sed rewrites them; awk understands that lines often have columns. It reads each line, splits it into fields on runs of spaces or tabs, and names them $1, $2, and so on; $0 is the whole line, NF is the number of fields and $NF the last one. An awk program is a list of pattern { action } pairs: for every line that matches the pattern, the action runs. Leave out the pattern and the action runs for every line; leave out the action and matching lines are printed.
~% cd /root/lab/l40
l40% cat sales.txt
ana madrid 1200
luis sevilla 950
maria bilbao 2100
pepe madrid 600
l40% awk '{print $1, $3}' sales.txt
ana 1200
luis 950
maria 2100
pepe 600
l40% awk '$3 > 1000' sales.txt
ana madrid 1200
maria bilbao 2100
l40% awk '$2 == "madrid" {print $1}' sales.txt
ana
pepe
l40% awk '{sum += $3} END {print "total:", sum}' sales.txt
total: 4850
Fields and separators
print $1, $3 prints two fields separated by a space (the comma inserts the output separator; without it the values are glued together). For files that use another separator, -F sets it: awk -F: '{print $1}' /etc/passwd lists user names, and -F, reads simple CSV. Always wrap the program in single quotes, because $1 inside double quotes would be expanded by the shell before awk ever saw it.
Patterns, variables, BEGIN and END
Patterns can be regular expressions, /madrid/, or comparisons on fields: $3 > 1000, $2 == "madrid", $1 ~ /^m/ (field matches a regex), joined with && and ||. NR is the current line number, so NR > 1 skips a header. Variables need no declaration and start at zero, which makes totals one line long: {sum += $3} END {print sum}. BEGIN { } runs before the first line, handy for printing a heading, and END { } after the last, where totals and averages (sum/NR) belong. printf aligns columns: printf "%-8s %6d\n", $1, $3.
Module review: which tool?
| Question | Tool |
|---|---|
| Which lines contain X? | grep (+ -i -n -v -c -r, regex, -E) |
| Which files are called X, are big, old or writable? | find (-name -type -size -mtime -perm) |
| Do something to each of those files | find -exec, xargs |
| Where does command X come from? | which, type |
| Change X into Y, drop or pick lines | sed |
| Columns, filters by value, totals | awk |
Commands in this lesson
| Command | What it does |
|---|---|
awk '{print $1, $3}' FILE | Print fields 1 and 3. |
awk '{print $NF}' FILE | Print the last field. |
awk -F: '{print $1}' FILE | Use `:` as the field separator. |
awk '$3 > 1000' FILE | Rows where field 3 exceeds 1000. |
awk '/re/ {print $2}' FILE | Field 2 of lines matching a regex. |
awk '{s += $3} END {print s}' FILE | Sum a column. |
awk 'NR > 1' FILE | Skip the first line. |
Quiz
In awk, what is `$NF`?
- The number of fields
- The last field of the line
- The file name
How do you print the user names (first `:`-separated field) of /etc/passwd?
- `awk '{print $1}' /etc/passwd`
- `awk -F: '{print $1}' /etc/passwd`
- `awk ':{print 1}' /etc/passwd`
What does `awk '{s += $2} END {print s}' nums.txt` print?
- The second column, line by line
- The sum of the second column
- The number of lines
Why must an awk program be in single quotes?
- So the shell does not expand `$1` and friends before awk runs
- awk does not accept double quotes
- Single quotes make it run faster
You need the names of all `.log` files over 1 MiB under /var. Which tool is the right starting point?
- `grep -r`
- `find` with `-name` and `-size`
- `awk`
Which command replaces `localhost` with `127.0.0.1` everywhere in `hosts.conf`, editing the file?
- `awk -i 's/localhost/127.0.0.1/' hosts.conf`
- `sed -i 's/localhost/127.0.0.1/g' hosts.conf`
- `grep -r localhost 127.0.0.1 hosts.conf`
Practice
/root/lab/l40/sales.txt has three columns: seller, city, amount. Save just the seller and the amount of every line, separated by a space, into /root/lab/l40/cols.txt.
Compute the total of the amount column of /root/lab/l40/sales.txt and save just the number into /root/lab/l40/total.txt.
Module challenge: from /etc/passwd, list the user names of the accounts whose shell (the 7th `:`-separated field) is exactly `/bin/false`, sorted alphabetically, into /root/lab/l40/nologin.txt.
Open this lesson in the app to do the tasks in a real Linux machine and have them checked.