Text Tools: grep, sed, awk
Find lines, cut columns, count and rank, rewrite text and add up numbers: a web server's log and a config file, investigated from the command line.
An interactive Linux lesson: 25 steps, about 35 minutes, on a live simulation in your browser.
Customers report that checkout failed this morning. On web01, in /srv/shop, you have the web server's access log: one line per request. wc -l says 36 lines. That is small enough to check every answer by eye, and every command in this lesson works unchanged on 36 million.
Before searching a log, read one line and name its parts. Split on spaces, the first line is: the client IP; two dashes (unused identity fields); the timestamp in brackets; the request in quotes, which is a method (GET), a URL (/) and a protocol; the status code (200); and the bytes sent (5124).
What you will learn
Find the lines: grep
- Read one line before you search
- grep: keep the lines that match: grep is a filter on whole lines: a line that contains the pattern is printed, every other line is dropped.
- Find the server errors: grep matches characters, not columns. Put enough context in the pattern to say which 500 you mean.
- Patterns: anchors and classes: A regular expression describes a shape: ^ and $ pin it to the ends of the line, [ ] is one character from a set, . is any one character.
- Repeat, or, and only the match: * + ? say how many, | says or. With -E they work unescaped; with -o grep prints the match instead of the line.
- Break it: an unquoted pattern: The shell reads your line first. Put every pattern in single quotes so that grep, not the shell, interprets it.
- Around the match, across the tree
Slice, sort, count
- cut: keep the columns you name: cut keeps the columns you name: -d is the separator, -f the field numbers. It splits at every single separator character.
- Count the clients: uniq only merges lines that are next to each other. Sort first, or it counts runs instead of totals.
- The top-N pipeline: Top N of anything: cut the column | sort | uniq -c | sort -rn | head.
- sort compares text unless told: sort compares text by default: -n for numbers, -r to reverse, -k N to sort by column N, -u to drop duplicates.
- Which URLs returned 500?
Rewrite text: sed
- sed: substitute on the way through: sed edits the stream, not the file: every line goes in, is changed or not, and comes out on stdout.
- Addresses: which lines to act on: A sed command is [address]action: which lines, then what to do with them. No address means every line.
- Break it: sed -i without looking: sed -i has no undo. Run the command without -i and read the output first; then add -i, or -i.bak to keep the original.
- Drill: change one setting
Columns with a brain: awk
- awk: fields without counting spaces: awk splits every line into fields on runs of whitespace: $1 is the first, $NF the last, $0 the whole line.
- pattern { action }: An awk program is pattern { action }: for each line where the pattern is true, run the action. Patterns can compare a field as a number.
- Add things up: END: awk variables survive from line to line. Accumulate in the main block, print in END.
- Drill: who saw an error?
JSON, and choosing a tool
- JSON is not lines and columns: JSON is a tree, not lines and columns. jq walks the tree by path; line tools can only guess.
- jq: ask by path
- Pick the smallest tool: Use the smallest tool that answers the question: grep for lines, cut for a column, sort | uniq -c to count, sed to rewrite, awk when columns need logic, jq for JSON.
Recap & playground
- Cheat sheet
- Playground