learninfra · Linux · Networking · Kubernetes · System Design · AI Infrastructure · Exam blueprints · Drills

Text Tools: grep, sed, awk

Find lines, cut columns, count and rank, rewrite text and add up numbers: a web server's log and a config file, investigated from the command line.

An interactive Linux lesson: 25 steps, about 35 minutes, on a live simulation in your browser.

Customers report that checkout failed this morning. On web01, in /srv/shop, you have the web server's access log: one line per request. wc -l says 36 lines. That is small enough to check every answer by eye, and every command in this lesson works unchanged on 36 million.

Before searching a log, read one line and name its parts. Split on spaces, the first line is: the client IP; two dashes (unused identity fields); the timestamp in brackets; the request in quotes, which is a method (GET), a URL (/) and a protocol; the status code (200); and the bytes sent (5124).

What you will learn

  1. Find the lines: grep

    • Read one line before you search
    • grep: keep the lines that match: grep is a filter on whole lines: a line that contains the pattern is printed, every other line is dropped.
    • Find the server errors: grep matches characters, not columns. Put enough context in the pattern to say which 500 you mean.
    • Patterns: anchors and classes: A regular expression describes a shape: ^ and $ pin it to the ends of the line, [ ] is one character from a set, . is any one character.
    • Repeat, or, and only the match: * + ? say how many, | says or. With -E they work unescaped; with -o grep prints the match instead of the line.
    • Break it: an unquoted pattern: The shell reads your line first. Put every pattern in single quotes so that grep, not the shell, interprets it.
    • Around the match, across the tree
  2. Slice, sort, count

    • cut: keep the columns you name: cut keeps the columns you name: -d is the separator, -f the field numbers. It splits at every single separator character.
    • Count the clients: uniq only merges lines that are next to each other. Sort first, or it counts runs instead of totals.
    • The top-N pipeline: Top N of anything: cut the column | sort | uniq -c | sort -rn | head.
    • sort compares text unless told: sort compares text by default: -n for numbers, -r to reverse, -k N to sort by column N, -u to drop duplicates.
    • Which URLs returned 500?
  3. Rewrite text: sed

    • sed: substitute on the way through: sed edits the stream, not the file: every line goes in, is changed or not, and comes out on stdout.
    • Addresses: which lines to act on: A sed command is [address]action: which lines, then what to do with them. No address means every line.
    • Break it: sed -i without looking: sed -i has no undo. Run the command without -i and read the output first; then add -i, or -i.bak to keep the original.
    • Drill: change one setting
  4. Columns with a brain: awk

    • awk: fields without counting spaces: awk splits every line into fields on runs of whitespace: $1 is the first, $NF the last, $0 the whole line.
    • pattern { action }: An awk program is pattern { action }: for each line where the pattern is true, run the action. Patterns can compare a field as a number.
    • Add things up: END: awk variables survive from line to line. Accumulate in the main block, print in END.
    • Drill: who saw an error?
  5. JSON, and choosing a tool

    • JSON is not lines and columns: JSON is a tree, not lines and columns. jq walks the tree by path; line tools can only guess.
    • jq: ask by path
    • Pick the smallest tool: Use the smallest tool that answers the question: grep for lines, cut for a column, sort | uniq -c to count, sed to rewrite, awk when columns need logic, jq for JSON.
  6. Recap & playground

    • Cheat sheet
    • Playground