Common causes
- The download saved an HTML page (404, login, redirect) instead of the archive
- The archive is an uncompressed tar named .tar.gz
- It is compressed with xz, bzip2 or zstd rather than gzip
- It is actually a ZIP file with the wrong extension
- A web server or browser already decompressed a Content-Encoding: gzip download
- FTP ASCII mode or a text editor altered the binary data
How to fix it
- Check what the file is. Run file backup.tar.gz. A gzip file reports 'gzip compressed data'; other results like 'POSIX tar archive', 'XZ compressed data', 'Zip archive' or 'HTML document' tell you what to do next.
- Let tar detect the compression. GNU tar detects compression on extract, so tar -xf backup.tar.gz works for gzip, bzip2, xz and plain tar. Only add -z when you are sure it is gzip.
- Use the matching tool. Use tar -xJf for .tar.xz, tar -xjf for .tar.bz2, tar --zstd -xf for .tar.zst, and unzip for a Zip archive.
- Re-download if it is HTML. Look inside with head -c 300 backup.tar.gz. If you see HTML, download again with curl -fL -o file URL and include any required authentication.
- Re-transfer in binary mode. If the file reports as gzip on the server but not locally, the transfer changed it. Copy it again with scp, rsync or FTP binary mode and compare sha256sum.
Terminal
file backup.tar.gz
# gzip compressed data -> tar -xzf backup.tar.gz
# POSIX tar archive -> tar -xf backup.tar.gz
# XZ compressed data -> tar -xJf backup.tar.gz
# Zip archive data -> unzip backup.tar.gz
# HTML document -> download it again with curl -fL How to stop it happening again
- Use curl -fL so error pages are not saved as archives
- Name archives by their real compression (.tar, .tar.gz, .tar.xz, .zip)
- Use tar -xf without a compression flag and let GNU tar detect it
- Publish and verify checksums for downloadable archives