1. 生成器是什么? #

方式 写法 典型场景
生成器函数 def f(): yield x 自定义分批、过滤、读文件
生成器表达式 (x for x in items) 简单映射/过滤

2. 生成器函数:yield #

# 生成器函数:逐行读取文件
def read_lines(path):
    with open(path, "r", encoding="utf-8") as f:
        for line in f:
            # yield 逐行产出,不占满内存
            yield line.strip()

# 消费生成器
for line in read_lines("access.log"):
    print(line)
# 按指定大小分批产出元素
def batch_items(items, size=100):
    """把可迭代对象按批产出,便于分批写库、调接口"""
    batch = []
    for item in items:
        batch.append(item)
        if len(batch) >= size:
            yield batch
            batch = []
    if batch:
        yield batch

def bulk_update(chunk):
    """批量更新用户"""
    print(f"Bulk updating {len(chunk)} users: {chunk}")

# 示例:2100 个用户 ID
user_ids = list(range(1, 2100))
# 每 500 个一批处理
for chunk in batch_items(user_ids, size=500):
    bulk_update(chunk)
普通函数 return 生成器 yield
结果 一次返回全部 多次逐个返回
内存 常需先构造完整列表 惰性,省内存
复用 可多次调用 生成器对象只能遍历一次

3. 生成器表达式 #

# 列表推导:立即占内存
squares_list = [x ** 2 for x in range(1000000)]

# 生成器表达式:惰性求值
squares_gen = (x ** 2 for x in range(1000000))
# 直接传给 sum,无需先转列表
total = sum(x ** 2 for x in range(1000000))

4. yield from #

def chain_sources(*sources):
    """
    接收多个可迭代对象,依次遍历所有元素。
    """
    for src in sources:
        # 委托给子可迭代对象
        yield from src

# 依次产出两个列表的所有元素
for item in chain_sources([1, 2, 3], [4, 5, 6]):
    print(item)

5. 项目中的典型场景 #

import json

# 解析 JSON 行
def parse_rows(lines):
    for line in lines:
        if not line:
            continue
        yield json.loads(line)

# 过滤活跃用户
def valid_users(rows):
    for row in rows:
        if row.get("active"):
            yield row

# 管道式处理:解析 → 过滤 → 输出
for user in valid_users(parse_rows(["{\"name\": \"Alice\", \"active\": true}", "{\"name\": \"Bob\", \"active\": false}"])):
    print(user)

6. 项目开发要点 #

7. 总结 #

概念 要点
yield 暂停并返回值,函数变生成器
惰性 省内存,适合大文件、大批量
表达式 (… for …) 简单转换
yield from 委托子迭代器
分批 yield batch 写库、调 API 常用模式