Read

Module 3 · Text, redirection and pipes

sort and uniq

Put lines in order, numerically or by a column, then collapse and count the duplicates: the classic recipe for a quick frequency report.

What you will learn

  • Sort text and numbers correctly with sort, sort -n and sort -r.
  • Sort by a specific field with -t and -k.
  • Remove or count duplicates with sort -u and uniq -c, and understand why uniq needs sorted input.

sort reads lines and writes them back in order; uniq reads lines and merges adjacent duplicates. Together, usually with a | between them, they answer questions like "which values appear in this column?" and "how often?". The two most important things to learn are that plain sort orders text, not numbers, and that uniq only sees duplicates that are next to each other.

Text order versus numeric order

By default sort compares character by character, so 100 comes before 9 because 1 is smaller than 9. Add -n and the lines are compared as numbers. -r reverses whatever order is in use, so sort -nr gives biggest first. -f ignores case, and -u keeps one copy of each distinct line, saving you a trip through uniq when you do not need counts.

~% printf '10\n9\n100\n' | sort
10
100
9
~% printf '10\n9\n100\n' | sort -n
9
10
100
~% printf '10\n9\n100\n' | sort -nr
100
10
9
~% printf 'b\na\nc\na\n' | sort -u
a
b
c

Sorting by a column

-k N sorts by the Nth field, where fields are separated by runs of blanks unless you choose another separator with -t. sort -t: -k3 -n /etc/passwd orders accounts by their numeric user id, the third colon-separated field. Write -k2,2 to compare only field 2; a bare -k2 means "from field 2 to the end of the line", which can give surprising ties.

uniq: adjacent duplicates only

uniq compares each line with the one before it and prints a line only when it differs. If the same value appears in lines 1 and 5 with something else in between, both survive. That is why the idiom is always sort file | uniq, never uniq file on unsorted data. uniq -c prefixes each surviving line with how many times it occurred, -d prints only the lines that were repeated, and -u only the ones that appeared exactly once.

~% cd /root/lab/l26
l26% cat names.txt
maria
luis
ana
luis
maria
maria
l26% uniq names.txt
maria
luis
ana
luis
maria
l26% sort names.txt | uniq -c
      1 ana
      2 luis
      3 maria
l26% sort names.txt | uniq -c | sort -rn
      3 maria
      2 luis
      1 ana
l26% sort names.txt | uniq -d
luis
maria

Writing the result back

Remember from lesson 22 that sort file > file destroys the file. sort has a safe alternative built in: sort -o file file reads everything first and then writes the output, so sorting a file in place is sort -o names.txt names.txt. For anything else, write to a new name.

Commands in this lesson

CommandWhat it does
sort fileSort lines as text.
sort -n fileSort numerically.
sort -nr fileNumeric, largest first.
sort -u fileSort and drop duplicate lines.
sort -t: -k3 -n fileSort by field 3 using : as separator.
sort -o file fileSort a file in place safely.
sort file | uniq -cCount how many times each line appears.
sort file | uniq -dShow only the lines that repeat.

Quiz

  1. What is the output of `printf '2\n10\n1\n' | sort`?

    • 1, 2, 10
    • 1, 10, 2
    • 10, 2, 1
  2. Why is `uniq` almost always used after `sort`?

    • Because uniq only merges duplicates that are on adjacent lines
    • Because uniq cannot read files directly
    • Because sort removes blank lines first
  3. Which command lists the most frequent lines of log.txt first, with their counts?

    • `sort log.txt | uniq -c | sort -rn`
    • `uniq -c log.txt | sort`
    • `sort -u log.txt | wc -l`
  4. What does `sort -t: -k3 -n /etc/passwd` sort by?

    • The user name
    • The third colon-separated field, numerically (the UID)
    • The third character of each line
  5. Which is the safe way to sort names.txt in place?

    • `sort names.txt > names.txt`
    • `sort -o names.txt names.txt`
    • `sort -i names.txt`

Practice

  1. /root/lab/l26/names.txt lists names with repetitions. Create /root/lab/l26/unique.txt with the names sorted alphabetically and each one appearing only once.

  2. /root/lab/l26/numbers.txt contains 10, 9, 100 and 25, one per line. Write them sorted numerically from largest to smallest into /root/lab/l26/desc.txt.

  3. Using sort and uniq, write into /root/lab/l26/counts.txt how many times each name appears in /root/lab/l26/names.txt (the uniq -c format).

Open this lesson in the app to do the tasks in a real Linux machine and have them checked.