1. 大模型与 GPT 是什么 #

2.1 大模型(LLM)是什么 #

2.2 GPT 与 ChatGPT 的关系 #

2.3 发展时间线 #

3. 大模型是如何工作的 #

3.1.文本分词 #

transformer.py

# 定义语料,每个句子中的词和标点由空格分隔
+corpus = [
+   "天气 预报 说 长沙 明天 有 暴雨 。",
+   "出门 请 携带 雨具 。",
+   "这里 是 一座 历史悠久 的 美丽 城市 。",
+   "大雨 在 下午 停了 , 太阳 出来 了 。",
+   "后天 大部分 地区 晴朗 温暖 。",
+   "强降雨 可能 造成 低洼 地区 洪涝 。",
+   "国王 的 女儿 善良 勇敢 。",
+   "公主 住 在 森林 附近 的 城堡 里 。",
+   "很久 以前 一位 国王 住 在 遥远 的 国度 。",
+   "她 每天 喜欢 读书 和 学习 新 知识 。",
+]
# 定义分词函数,用于将句子按空格分割为词列表
+def split_words(sentence):
    # 按空格分割句子,返回词列表
+   return sentence.split()

# 对每个句子调用分词函数,得到所有句子的分词结果
+words_per_sentence = [split_words(s) for s in corpus]
+print(words_per_sentence)
# 构建一个集合,用于收集语料中出现过的所有词(去除重复)
+all_words = set()
# 遍历每个分词后的句子
+for words in words_per_sentence:
    # 遍历句子中的每个词
+   for word in words:
        # 将词加入集合
+       all_words.add(word)

3.2 概率统计 #

在分词之后,大模型的下一个重要步骤是理解词与词之间的关系。最简单的方式,是统计每一个词后面会接什么词,以及这种组合出现的概率。这可以帮助模型“预测”一句话下一个最有可能出现的词,这就是最初的语言模型思想。

例如,在前面分词的基础上,我们遍历所有句子,对于每一个词,记录它后面紧跟着哪些词(以及次数),得到相邻词对的统计表。进而,可以计算条件概率——比如“天气”后接“预报”的概率是多少、“国王”后面最常出现哪些词等等。

transformer.py

# 导入 defaultdict,用于实现嵌套的字典统计结构
+from collections import defaultdict
# 定义语料,每个句子中的词和标点由空格分隔
corpus = [
    "天气 预报 说 长沙 明天 有 暴雨 。",
    "出门 请 携带 雨具 。",
    "这里 是 一座 历史悠久 的 美丽 城市 。",
    "大雨 在 下午 停了 , 太阳 出来 了 。",
    "后天 大部分 地区 晴朗 温暖 。",
    "强降雨 可能 造成 低洼 地区 洪涝 。",
    "国王 的 女儿 善良 勇敢 。",
    "公主 住 在 森林 附近 的 城堡 里 。",
    "很久 以前 一位 国王 住 在 遥远 的 国度 。",
    "她 每天 喜欢 读书 和 学习 新 知识 。",
]
# 定义分词函数,用于将句子按空格分割为词列表
def split_words(sentence):
    # 按空格分割句子,返回词列表
    return sentence.split()

# 对每个句子调用分词函数,得到所有句子的分词结果
words_per_sentence = [split_words(s) for s in corpus]
# 构建一个集合,用于收集语料中出现过的所有词(去除重复)
all_words = set()
# 遍历每个分词后的句子
for words in words_per_sentence:
    # 遍历句子中的每个词
    for word in words:
        # 将词加入集合
        all_words.add(word)

# 构建相邻词出现次数统计:pair_count[当前词][下一个词] = 出现次数
+pair_count = defaultdict(lambda: defaultdict(int))
# 遍历所有分词句子
+for words in words_per_sentence:
    # 遍历句子中的每对相邻词
+   for i in range(len(words) - 1):
        # 取出当前词
+       current_word = words[i]
        # 取出下一个词
+       next_word = words[i + 1]
        # 当前词到下一个词的转移次数加一
+       pair_count[current_word][next_word] += 1
+print(pair_count)
# 构建下一个词的条件概率表:next_word_prob[当前词][下一个词]=概率
+next_word_prob = {}
# 遍历每个当前词及其统计映射
+for current_word, count_map in pair_count.items():
    # 统计所有下一个词的出现次数之和
+   total = sum(count_map.values())
    # 计算各下一个词的概率,组成一个新的字典
+   next_word_prob[current_word] = {
+       word: count / total for word, count in count_map.items()
+   }
+print(next_word_prob)

3.3 生成文本 #

transformer.py

# 导入 defaultdict,用于实现嵌套的字典统计结构
from collections import defaultdict
# 定义语料,每个句子中的词和标点由空格分隔
corpus = [
   "天气 预报 说 长沙 明天 有 暴雨",
   "出门 请 携带 雨具",
   "这里 是 一座 历史悠久 的 美丽 城市",
   "大雨 在 下午 停了 , 太阳 出来 了",
   "后天 大部分 地区 晴朗 温暖",
   "强降雨 可能 造成 低洼 地区 洪涝",
   "国王 的 女儿 善良 勇敢",
   "公主 住 在 森林 附近 的 城堡 里",
   "很久 以前 一位 国王 住 在 遥远 的 国度",
   "她 每天 喜欢 读书 和 学习 新 知识",
]
# 定义分词函数,用于将句子按空格分割为词列表
def split_words(sentence):
    # 按空格分割句子,返回词列表
    return sentence.split()

# 对每个句子调用分词函数,得到所有句子的分词结果
words_per_sentence = [split_words(s) for s in corpus]
# 构建一个集合,用于收集语料中出现过的所有词(去除重复)
all_words = set()
# 遍历每个分词后的句子
for words in words_per_sentence:
    # 遍历句子中的每个词
    for word in words:
        # 将词加入集合
        all_words.add(word)

# 构建相邻词出现次数统计:pair_count[当前词][下一个词] = 出现次数
pair_count = defaultdict(lambda: defaultdict(int))
# 遍历所有分词句子
for words in words_per_sentence:
    # 遍历句子中的每对相邻词
    for i in range(len(words) - 1):
        # 取出当前词
        current_word = words[i]
        # 取出下一个词
        next_word = words[i + 1]
        # 当前词到下一个词的转移次数加一
        pair_count[current_word][next_word] += 1
# 构建下一个词的条件概率表:next_word_prob[当前词][下一个词]=概率
next_word_prob = {}
# 遍历每个当前词及其统计映射
for current_word, count_map in pair_count.items():
    # 统计所有下一个词的出现次数之和
    total = sum(count_map.values())
    # 计算各下一个词的概率,组成一个新的字典
    next_word_prob[current_word] = {
        word: count / total for word, count in count_map.items()
    }

# 预测下一个词的函数,根据当前词和概率表返回概率最高的下一个词
+def predict_next_word(current_word):
    # 获取当前词对应的下一个词概率映射
+   prob_map = next_word_prob.get(current_word)
    # 如果没有找到,则返回None
+   if not prob_map:
+       return None
    # 返回概率最大的下一个词,若概率相同优先选句末标点,其次按词典序
+   return max(
+       prob_map.items(),
+       key=lambda item: -item[1],
+   )[0]

# 定义一个函数,用于将词列表拼接为字符串
+def join_words(words):
    # 使用空字符串作为分隔符,将词列表连接成一个字符串
+   return "".join(words)

# 定义一个函数,用于根据首词自动补全生成句子
+def complete_sentence(first_word):
    # 检查首词是否在所有词的集合中
+   if first_word not in all_words:
        # 如果不在词表中,则抛出异常提示
+       raise ValueError(f"首词「{first_word}」不在语料词表中")

    # 初始化生成的词列表,将首词加入列表
+   generated_words = [first_word]
    # 初始化当前词为首词
+   current_word = first_word
    # 进入生成循环,直到无法继续
+   while True:
        # 预测下一个词
+       next_word = predict_next_word(current_word)
        # 如果无法预测下一个词,则跳出循环
+       if next_word is None:
+           break
        # 将下一个词加入已生成词列表
+       generated_words.append(next_word)
        # 更新当前词为刚刚得到的下一个词
+       current_word = next_word
    # 返回拼接成字符串后的完整句子
+   return join_words(generated_words)

# 定义补全句子的起始词,这里选用“后天”
+first_word = "后天"
# 从首词开始调用补全函数,返回生成的句子
+result_sentence = complete_sentence(first_word)
# 打印补全后的句子结果
+print(f"生成文本: {result_sentence}")

4.执行流程 #

4.1. 程序做什么? #

一句话概括:

输入:

后天

输出:

补全结果: 后天大部分地区晴朗温暖

4.2. 整体流程 #

flowchart TD A[corpus 训练语料] --> B[split_words 分词] B --> C[words_per_sentence 每句词列表] C --> D[all_words 全部词集合] C --> E[pair_count 相邻词次数统计] E --> F[next_word_prob 下一个词概率表] G[first_word 首词] --> H[complete_sentence 补全句子] H --> I[predict_next_word 查概率选词] F --> I I --> H H --> J[join_words 拼成句子] J --> K[result_sentence 输出结果]

4.3. 逐步执行 #

4.3.1 训练语料 #

corpus = [
    "天气 预报 说 长沙 明天 有 暴雨",
    "出门 请 携带 雨具",
    ...
]

4.3.2 分词 #

def split_words(sentence):
    return sentence.split()

4.3.3 收集全部词 #

words_per_sentence = [split_words(s) for s in corpus]
all_words = set()
for words in words_per_sentence:
    for word in words:
        all_words.add(word)

4.3.4 统计相邻词 #

pair_count[current_word][next_word] += 1

4.3.5 计算概率 #

next_word_prob[current_word] = {
    word: count / total for word, count in count_map.items()
}

$$ P(\text{下一个词} \mid \text{当前词}) = \frac{\text{该组合出现次数}}{\text{当前词后出现总次数}} $$

$$ P(\text{大部分} \mid \text{后天}) = 100\% $$

4.3.6 预测下一个词 #

def predict_next_word(current_word):
    prob_map = next_word_prob.get(current_word)
    if not prob_map:
        return None
    return max(prob_map.items(), key=lambda item: -item[1])[0]

4.3.7 补全句子 #

def complete_sentence(first_word):
    generated_words = [first_word]
    current_word = first_word
    while True:
        next_word = predict_next_word(current_word)
        if next_word is None:
            break
        generated_words.append(next_word)
        current_word = next_word
    return join_words(generated_words)

4.4. 程序启动与数据准备 #

描述脚本导入后、生成前的初始化过程(统计阶段)。

sequenceDiagram autonumber participant 主程序 participant corpus as 语料 corpus participant split_words as 分词 split_words participant words_per_sentence as 分词结果 participant all_words as 词表 all_words participant pair_count as 相邻词统计 participant next_word_prob as 概率表 主程序->>corpus: 读取 10 条训练句子 loop 遍历每条句子 主程序->>split_words: split_words(句子) split_words-->>主程序: 返回词列表 主程序->>words_per_sentence: 追加到分词结果 end loop 遍历所有词 主程序->>all_words: 去重收集全部词 end loop 遍历相邻词对 主程序->>pair_count: current_word → next_word 次数 +1 end loop 遍历每个当前词 主程序->>next_word_prob: 次数 ÷ 总次数 = 概率 end Note over next_word_prob: 数据准备完成,等待生成

4.5. 从首词补全句子 #

描述 complete_sentence("后天") 的生成阶段(核心循环)。

sequenceDiagram autonumber participant 主程序 participant complete_sentence as 补全 complete_sentence participant predict_next_word as 预测 predict_next_word participant next_word_prob as 概率表 participant join_words as 拼接 join_words 主程序->>complete_sentence: complete_sentence("后天") complete_sentence->>complete_sentence: generated_words = ["后天"]<br/>current_word = "后天" loop 逐词生成(直到无法预测) complete_sentence->>predict_next_word: predict_next_word(current_word) predict_next_word->>next_word_prob: 查询 current_word 的概率分布 next_word_prob-->>predict_next_word: 返回 {下一个词: 概率} predict_next_word-->>complete_sentence: 返回概率最高的 next_word alt next_word 为 None complete_sentence->>complete_sentence: break 结束循环 else next_word 有值 complete_sentence->>complete_sentence: 追加 next_word<br/>current_word = next_word end end complete_sentence->>join_words: join_words(generated_words) join_words-->>complete_sentence: 拼接后的字符串 complete_sentence-->>主程序: 返回完整句子 主程序->>主程序: print 补全结果

4.6. 首词「后天」的生成过程 #

语料中有一条:后天 大部分 地区 晴朗 温暖

因此生成路径唯一且确定:

步骤 当前词 下一个词 概率
0 (起始) 后天 —
1 后天 大部分 100%
2 大部分 地区 100%
3 地区 晴朗 50%(贪心选中)
4 晴朗 温暖 100%
5 温暖 — 无后继,结束

最终输出:后天大部分地区晴朗温暖