Regular expressions, part 1: ., *, ^, $
The four characters that turn a grep pattern from a fixed string into a description: any character, repetition, start of line and end of line.
What you will learn
- Explain what ., *, ^ and $ mean inside a pattern.
- Anchor a pattern to the start or end of a line and find empty lines.
- Escape a metacharacter to match it literally, or use grep -F.
So far grep patterns were plain words. A regular expression (regex) is a pattern in a small language where a few characters have special meaning. With just four of them you can say "lines that start with a hash", "lines that end in a digit", "empty lines" or "the word color, however it is spelled". The same syntax works in grep, sed, awk, vi and most programming languages, so what you learn here travels well.
The four characters
| Char | Meaning | Example |
|---|---|---|
. | any single character (except newline) | c.t matches cat, cot, cut, c9t |
* | zero or more of the preceding item | ab*c matches ac, abc, abbbc |
^ | start of line | ^# lines beginning with # |
$ | end of line | ;$ lines ending in ; |
* is the one that trips people up, because it does not mean what it means in file names. In a shell wildcard * is "anything"; in a regex it is "repeat the previous thing zero or more times". So .* is "anything" (any character, repeated), and a* matches the empty string too, which is why grep 'a*' file prints every line. The very common .* lets you say "this, then anything, then that": grep 'start.*end'.
Anchors
Without anchors a pattern matches anywhere in the line. ^ pins it to the start, $ to the end, and both together make it match the whole line: grep '^root