Module 8 · Archives, compression and the network
Downloading with wget
Fetch files and web pages from the command line with wget: choose the file name or directory, print to stdout, look at the HTTP headers, and understand what the sandbox lets through.
What you will learn
- Download to a chosen file (`-O`), directory (`-P`) or stdout (`-O -`).
- Inspect status codes and headers with `-S`, and check URLs with `--spider`.
- Recognise the sandbox's limits: http:// only, port 80, CORS-friendly sites, 502 Fetch Error.
wget speaks HTTP: give it a URL and it downloads what the server sends back. It needs no browser and no mouse, so it is how servers fetch software, scripts grab data and admins test that a web service answers. All the live examples in this lesson need the machine's network switched on (Built-in mode) and an address (udhcpc -i eth0).
~% cd /tmp
~% wget http://httpbin.org/robots.txt
Connecting to httpbin.org (192.168.87.1:80)
robots.txt 100% |*******************************| 30 0:00:00 ETA
~% cat robots.txt
User-agent: *
Disallow: /deny
By default the file lands in the current directory, named after the last part of the URL. The progress bar and the *Connecting to* line go to stderr, so they do not pollute redirected output. Choose where the data goes with these options:
| Option | Meaning |
|---|---|
-O FILE | Save under this name |
-O - | Write the body to stdout, for pipes |
-P DIR | Save into DIR (it must exist here) |
-q | Quiet: no progress output |
-S | Show the server's response headers |
-T SEC | Give up after SEC seconds without data |
--spider | Only check the URL exists; exit status says |
-U AGENT | Send a custom User-Agent |
~% wget -q -O - http://httpbin.org/user-agent
{
"user-agent": "Wget"
}
~% wget -S -O /dev/null http://httpbin.org/get
Connecting to httpbin.org (192.168.87.1:80)
HTTP/1.1 200 OK
access-control-allow-origin: *
content-type: application/json
~% wget -q --spider http://httpbin.org/status/404; echo $?
1
The first line of a response is the status code: 200 OK, 301/302 redirect, 404 not found, 500 server error. wget exits with 0 only on success, which makes it a good health check in scripts: wget -q --spider URL && echo alive.
What the sandbox lets through
Your VM runs inside a browser tab, and a web page cannot open raw network connections. TempMV's virtual network accepts the guest's TCP connection to port 80, reads the HTTP request, and replays it as an ordinary browser fetch(). That has honest consequences, and learning them teaches you how the web works:
- Write
http://. This BusyBox wget has no TLS (not an http or ftp url: https://…); the browser upgrades the request to HTTPS for you when needed. - Only port 80:
http://host:8080/getsConnection refused. SSH, mail or database ports cannot work. - Only sites that allow cross-origin requests (CORS) answer. Others produce
server returned error: HTTP/1.1 502 Fetch Error, written by the emulator when the browser blocked the response. Sites that work:http://httpbin.org/get,http://api.github.com/zen, files onhttp://cdn.jsdelivr.net/.
Commands in this lesson
| Command | What it does |
|---|---|
wget URL | Download into the current directory. |
wget -O file URL | Download under a chosen name. |
wget -q -O - URL | Print the body to stdout, quietly. |
wget -P dir URL | Download into an existing directory. |
wget -S -O /dev/null URL | Show the response headers only. |
wget -q --spider URL | Check a URL; exit status 0 if it exists. |
Quiz
Which command saves http://httpbin.org/get as data.json?
- wget -P data.json http://httpbin.org/get
- wget -O data.json http://httpbin.org/get
- wget http://httpbin.org/get > data.json
- wget -S data.json http://httpbin.org/get
What does `-O -` do?
- Deletes the downloaded file
- Writes the body to standard output
- Disables output entirely
- Overwrites an existing file
Inside TempMV, `wget http://example.com` returns '502 Fetch Error'. What is the most likely cause?
- example.com is down
- The browser was not allowed to read the site (no CORS), so the emulator reported a 502
- wget needs -S
- The DNS server is missing
Why does `wget https://httpbin.org/get` fail in this VM while the http:// URL works?
- httpbin.org has no HTTPS
- This BusyBox wget has no TLS support; the browser does the HTTPS part for http:// URLs
- Port 443 is blocked by your router
- HTTPS needs root
Which command tells a script whether a URL answers, without saving anything?
- wget -q --spider URL
- wget -P /dev URL
- ping URL
- nslookup URL
Practice
Network needed. Download `http://httpbin.org/get` and save it as `/root/lab/l78/me.json`.
Network needed. Download `http://httpbin.org/robots.txt` into the directory `/root/lab/l78/dl`, keeping its original file name. Remember that this wget does not create the directory.
Network needed. Request `http://httpbin.org/get` with wget showing the server's response headers, throw the body away (`-O /dev/null`), and save what wget prints on stderr into `/root/lab/l78/headers.txt`.
Open this lesson in the app to do the tasks in a real Linux machine and have them checked.