Read

Module 8 · Archives, compression and the network

Downloading with wget

Fetch files and web pages from the command line with wget: choose the file name or directory, print to stdout, look at the HTTP headers, and understand what the sandbox lets through.

What you will learn

  • Download to a chosen file (`-O`), directory (`-P`) or stdout (`-O -`).
  • Inspect status codes and headers with `-S`, and check URLs with `--spider`.
  • Recognise the sandbox's limits: http:// only, port 80, CORS-friendly sites, 502 Fetch Error.

wget speaks HTTP: give it a URL and it downloads what the server sends back. It needs no browser and no mouse, so it is how servers fetch software, scripts grab data and admins test that a web service answers. All the live examples in this lesson need the machine's network switched on (Built-in mode) and an address (udhcpc -i eth0).

~% cd /tmp
~% wget http://httpbin.org/robots.txt
Connecting to httpbin.org (192.168.87.1:80)
robots.txt           100% |*******************************|    30   0:00:00 ETA
~% cat robots.txt
User-agent: *
Disallow: /deny

By default the file lands in the current directory, named after the last part of the URL. The progress bar and the *Connecting to* line go to stderr, so they do not pollute redirected output. Choose where the data goes with these options:

OptionMeaning
-O FILESave under this name
-O -Write the body to stdout, for pipes
-P DIRSave into DIR (it must exist here)
-qQuiet: no progress output
-SShow the server's response headers
-T SECGive up after SEC seconds without data
--spiderOnly check the URL exists; exit status says
-U AGENTSend a custom User-Agent
~% wget -q -O - http://httpbin.org/user-agent
{
  "user-agent": "Wget"
}
~% wget -S -O /dev/null http://httpbin.org/get
Connecting to httpbin.org (192.168.87.1:80)
  HTTP/1.1 200 OK
  access-control-allow-origin: *
  content-type: application/json
~% wget -q --spider http://httpbin.org/status/404; echo $?
1

The first line of a response is the status code: 200 OK, 301/302 redirect, 404 not found, 500 server error. wget exits with 0 only on success, which makes it a good health check in scripts: wget -q --spider URL && echo alive.

What the sandbox lets through

Your VM runs inside a browser tab, and a web page cannot open raw network connections. TempMV's virtual network accepts the guest's TCP connection to port 80, reads the HTTP request, and replays it as an ordinary browser fetch(). That has honest consequences, and learning them teaches you how the web works:

  • Write http://. This BusyBox wget has no TLS (not an http or ftp url: https://…); the browser upgrades the request to HTTPS for you when needed.
  • Only port 80: http://host:8080/ gets Connection refused. SSH, mail or database ports cannot work.
  • Only sites that allow cross-origin requests (CORS) answer. Others produce server returned error: HTTP/1.1 502 Fetch Error, written by the emulator when the browser blocked the response. Sites that work: http://httpbin.org/get, http://api.github.com/zen, files on http://cdn.jsdelivr.net/.

Commands in this lesson

CommandWhat it does
wget URLDownload into the current directory.
wget -O file URLDownload under a chosen name.
wget -q -O - URLPrint the body to stdout, quietly.
wget -P dir URLDownload into an existing directory.
wget -S -O /dev/null URLShow the response headers only.
wget -q --spider URLCheck a URL; exit status 0 if it exists.

Quiz

  1. Which command saves http://httpbin.org/get as data.json?

    • wget -P data.json http://httpbin.org/get
    • wget -O data.json http://httpbin.org/get
    • wget http://httpbin.org/get > data.json
    • wget -S data.json http://httpbin.org/get
  2. What does `-O -` do?

    • Deletes the downloaded file
    • Writes the body to standard output
    • Disables output entirely
    • Overwrites an existing file
  3. Inside TempMV, `wget http://example.com` returns '502 Fetch Error'. What is the most likely cause?

    • example.com is down
    • The browser was not allowed to read the site (no CORS), so the emulator reported a 502
    • wget needs -S
    • The DNS server is missing
  4. Why does `wget https://httpbin.org/get` fail in this VM while the http:// URL works?

    • httpbin.org has no HTTPS
    • This BusyBox wget has no TLS support; the browser does the HTTPS part for http:// URLs
    • Port 443 is blocked by your router
    • HTTPS needs root
  5. Which command tells a script whether a URL answers, without saving anything?

    • wget -q --spider URL
    • wget -P /dev URL
    • ping URL
    • nslookup URL

Practice

  1. Network needed. Download `http://httpbin.org/get` and save it as `/root/lab/l78/me.json`.

  2. Network needed. Download `http://httpbin.org/robots.txt` into the directory `/root/lab/l78/dl`, keeping its original file name. Remember that this wget does not create the directory.

  3. Network needed. Request `http://httpbin.org/get` with wget showing the server's response headers, throw the body away (`-O /dev/null`), and save what wget prints on stderr into `/root/lab/l78/headers.txt`.

Open this lesson in the app to do the tasks in a real Linux machine and have them checked.