1. urllib 是什么? #
urllib 是 Python 标准库,用于发送 HTTP 请求、解析 URL、处理响应,无需额外安装。
- 适合脚本、工具类项目等不想引入第三方依赖的场景。
- 日常开发中
requests 更简洁,但 urllib 有助于理解 HTTP 底层机制。
2. 模块结构 #
urllib.request:发送 HTTP 请求、下载文件,最常用。
urllib.parse:编码/解码 URL、拼接查询参数。
urllib.error:捕获 HTTPError(4xx/5xx)和 URLError(网络/超时)。
urllib.response 通常不需要直接导入,由 urlopen() 返回响应对象。
import urllib.request
import urllib.parse
import urllib.error
3. 前置知识:HTTP 与 URL #
- GET 用于获取数据,POST 用于提交数据(如调用 API)。
- 常见状态码:
200 成功、404 资源不存在、500 服务器错误。
- URL 由协议、域名、路径、查询参数组成:
https://api.example.com/users?page=1。
urlopen() 返回的响应内容是字节,需 .decode('utf-8') 转为字符串。
4. GET 请求 #
urllib.request.urlopen(url) 发送 GET 请求,返回响应对象。
- 推荐用
with 语句,请求结束后自动关闭连接。
response.read() 读取全部内容(字节),response.status 获取状态码。
response.getheader('Content-Type') 可读取单个响应头。
import urllib.request
with urllib.request.urlopen('https://httpbin.org/get') as response:
html = response.read().decode('utf-8')
print(f"状态码:{response.status}")
print(html[:200])
5. 带查询参数的 GET #
- 查询参数用字典表示,通过
urllib.parse.urlencode() 编码后拼到 URL。
urlencode() 会自动处理中文等特殊字符,无需手动 quote。
- 适合搜索、分页、筛选等需要传参的 GET 请求。
import urllib.request
import urllib.parse
params = {'q': 'Python', 'page': 1, 'limit': 10}
query = urllib.parse.urlencode(params)
url = f'https://httpbin.org/get?{query}'
with urllib.request.urlopen(url) as response:
print(response.read().decode('utf-8'))
6. POST 请求(JSON) #
- 现代 API 普遍使用 JSON,需设置
Content-Type: application/json。
- POST 的
data 参数必须是字节,用 json.dumps().encode('utf-8') 转换。
- 用
urllib.request.Request() 创建请求,可指定 method、data、headers。
import urllib.request
import json
data = json.dumps({'name': '张三', 'age': 25}).encode('utf-8')
req = urllib.request.Request(
'https://httpbin.org/post',
data=data,
headers={'Content-Type': 'application/json'},
method='POST'
)
with urllib.request.urlopen(req) as response:
print(response.read().decode('utf-8'))
7. 设置请求头与超时 #
- 常用请求头:
User-Agent(标识客户端)、Accept(期望的响应格式)、Authorization(Bearer Token 认证)。
timeout=秒 防止请求无限挂起,超时抛出 urllib.error.URLError。
- 部分网站会校验
User-Agent,脚本请求时建议设置,避免被拦截。
import urllib.request
import urllib.error
headers = {
'User-Agent': 'MyApp/1.0',
'Accept': 'application/json',
'Authorization': 'Bearer your_token'
}
req = urllib.request.Request('https://httpbin.org/headers', headers=headers)
try:
with urllib.request.urlopen(req, timeout=10) as response:
print(response.read().decode('utf-8'))
except urllib.error.URLError as e:
print(f"请求失败:{e.reason}")
8. urllib.parse:URL 处理 #
urlencode(dict):字典 → 查询字符串,GET/POST 传参最常用。
urlparse(url):拆分 URL 为协议、域名、路径、查询参数等部分。
quote(text) / unquote(text):编码/解码 URL 中的中文、空格等特殊字符。
parse_qs(query_string):查询字符串 → 字典(值为列表,支持同名参数)。
from urllib.parse import urlparse, urlencode, quote, unquote, parse_qs
result = urlparse('https://api.example.com/users?page=1')
print(result.scheme, result.hostname, result.path, result.query)
print(urlencode({'q': 'Python', 'page': 1}))
encoded = quote('特殊<字符>测试')
print(encoded)
print(unquote(encoded))
print(parse_qs('name=张三&hobbies=编程&hobbies=读书'))
9. urllib.error:异常处理 #
HTTPError:服务器返回 4xx/5xx 状态码,e.code 为状态码,e.reason 为原因。
URLError:网络不可达、DNS 失败、超时等,e.reason 为错误描述。
HTTPError 是 URLError 的子类,先捕获 HTTPError,再捕获 URLError。
- 生产代码中应对每种异常分别处理,避免一个请求失败导致程序崩溃。
import urllib.request
import urllib.error
try:
with urllib.request.urlopen('https://httpbin.org/status/404') as response:
print(response.read())
except urllib.error.HTTPError as e:
print(f"HTTP 错误:{e.code} - {e.reason}")
except urllib.error.URLError as e:
print(f"URL 错误:{e.reason}")
10. 下载文件 #
urllib.request.urlretrieve(url, save_path) 直接将远程文件保存到本地。
- 下载前用
os.makedirs() 确保目标目录存在。
- 适合下载图片、安装包、数据文件等静态资源。
import urllib.request
import os
def download_file(url, save_path):
os.makedirs(os.path.dirname(save_path), exist_ok=True)
try:
urllib.request.urlretrieve(url, save_path)
print(f'下载成功:{save_path}')
except Exception as e:
print(f'下载失败:{e}')
download_file('https://www.baidu.com', 'html/baidu.html')
11. 常见错误与注意事项 #
- 忘记
.decode('utf-8'),直接打印字节会得到 b'...' 而非可读文本。
- POST 的
data 传了字符串而非字节,会报 TypeError。
- 未设置
timeout,网络异常时程序可能长时间阻塞。
- 未设置
User-Agent,部分网站会拒绝脚本请求。
- SSL 证书问题可用
ssl.create_default_context() 处理,但生产环境不要跳过证书验证。
content = response.read().decode('utf-8')
data = json.dumps({'key': 'value'}).encode('utf-8')
urllib.request.urlopen(url, timeout=10)
11. 总结 #
urllib 是标准库网络方案:request 发请求、parse 处理 URL、error 捕异常。
- 核心流程:
Request 构造请求 → urlopen 发送 → read().decode() 读响应 → json.loads 解析。
- GET 用
urlencode 拼参数,POST JSON 设 Content-Type 并编码 body。
- 批量请求配合
14.concurrent.md 的 ThreadPoolExecutor 可显著提速。
11.1 速查 #
import urllib.request
import urllib.parse
import urllib.error
import json
with urllib.request.urlopen(url, timeout=10) as resp:
text = resp.read().decode('utf-8')
url = f'{base}?{urllib.parse.urlencode(params)}'
data = json.dumps(payload).encode('utf-8')
req = urllib.request.Request(url, data=data,
headers={'Content-Type': 'application/json'}, method='POST')
with urllib.request.urlopen(req) as resp:
result = json.loads(resp.read())
urllib.request.urlretrieve(url, 'local_file.zip')
13.2 最佳实践 #
- 始终用
with urlopen(...) 和 timeout,避免连接泄漏和无限等待
- POST JSON 记得设
Content-Type: application/json,body 编码为字节
- 用
try/except 分别处理 HTTPError 和 URLError
- 需要更简洁的 API 时优先考虑
requests,urllib 留给零依赖场景