Tags

C++14 编程语言5 BDB7 database16 transaction5 groovy1 DSL1 Spring MVC5 安全9 Java67 shell11 架构20 设计模式3 linux15 网络编程7 Python38 运维8 jekyll4 博客3 vim1 productivity4 高并发7 分布式10 存储6 HTTP11 编码5 Web12 JavaScript2 spring21 maven7 quartz7 junit3 ant2 spring-security1 DNS3 tomcat4 Debug2 高可用5 nginx9 JVM9 开放平台3 过载保护3 Concurrency2 aop1 proxy1 cglib1 生活5 性能优化5 Mina1 NIO1 序列化6 BTrace1 rpc6 Thrift1 architecture2 Distributed1 classloader1 zookeeper3 分布式锁1 缓存3 Observability14 log4j4 redis6 消息队列5 mysql7 elasticsearch15 移动互联网1 uuid1 Troubleshooting3 工作1 分享1 主从复制1 TDD2 hadoop4 spark7 大数据1 kafka3 广告1 Protobuf3 搜索1 git3 crawler2 图数据库10 neo4j5 aerospike3 Titan1 Bloom Filter1 markdown1 kramdown1 瑜伽1 呼吸1 生活的艺术1 antlr1 parser1 感恩节1 puppeteer1 chrome-headless1 cluster1 微服务2 AI231 RDD1 shuffle1 data skew1 敏捷3 gitlab1 机器学习2 特征工程1 Parameter Server1 埋点1 Interview21 Algorithms15 Data Structures3 LeetCode15 Hash Table1 Prefix Sum1 Two Pointers1 Sliding Window1 Stack1 Monotonic Stack1 Linked List1 Binary Tree1 Recursion2 DFS2 BFS2 Graph1 Union Find1 Dijkstra1 Binary Search1 Heap1 Priority Queue1 Intervals1 Greedy1 Backtracking1 String1 KMP1 Dynamic Programming2 Knapsack1 Design1 LRU1 Trie1 Attention2 Transformer24 KV Cache4 RoPE1 FlashAttention2 NumPy5 PyTorch34 LayerNorm1 Backpropagation1 Autograd1 Tokenizer1 BPE1 Sampling2 Beam Search1 Speculative Decoding2 Loss Function1 DPO2 PPO1 GRPO4 AdamW1 LoRA10 Machine Learning14 k-means1 PCA1 AUC1 NDCG1 Convolution1 NMS1 Thread Pool1 Memory Pool1 GEMM1 Allreduce1 Rate Limiting1 AI-Infra129 LLM101 Agent7 Roadmap4 Math10 Deep Learning8 Pretraining7 Post-Training15 RLHF8 Reasoning2 Distillation2 Evaluation5 Methodology1 Inference16 Quantization5 Decoding1 Long Context1 Pruning1 Small Models1 Multimodal11 Diffusion16 Vision1 Contrastive Learning1 VLM2 Training1 Speech2 Audio2 Generative Models2 Text-to-Image1 Video Generation6 Image Generation1 Unified Models1 CUDA14 Triton12 GPU32 Hugging Face6 transformers4 peft8 trl3 tokenizers1 datasets1 SFT1 QLoRA1 vLLM22 NCCL12 RDMA11 Megatron13 DeepSpeed7 Distributed Training16 torchtitan7 MFU2 Parallelism1 Checkpoint1 torchft1 Fault Tolerance1 Data Pipeline1 Blog2 GitHub1 大模型推理17 SGLang5 RL10 verl10 Ray2 FSDP2 AReaL1 Kubernetes11 Source Code2 DiT7 xDiT4 Roofline1 torch.compile1 SVDQuant1 TeaCache1 Cache-DiT1 Sparse Attention1 Demo1 Slides1 Sequence Parallelism1 PipeFusion1 Autoregressive1 Real-time1 Serving1 ControlNet1 diffusers1 Benchmarking1 MLOps3 DRA1 Kueue1 Volcano1 Scheduling1 MIG1 HAMi1 Storage1 KServe1 llm-d2 Gateway API1 Multi-Tenancy1 FinOps1 Open Source6 CI1 AI-Application9 API5 Reliability2 Tool Calling1 Cost1 Latency1 Benchmark1 Prompt Engineering1 Context Engineering1
C++

C++ 在 AI-Infra(09):系列总结与通关自测

C++ for AI-Infra: Series Recap and Final Self-Test


C++ 在 AI-Infra(08):构建、调试与测试工具链

Build, Debug and Test Toolchain


C++ 在 AI-Infra(07):与 Python 之间——pybind11、Python C API 与 ABI

pybind11, the Python C API and ABI


C++ 在 AI-Infra(06):并发、内存模型、TLS 与守卫

Concurrency, Memory Model, TLS and Guards


C++ 在 AI-Infra(05):宏、静态注册与代码生成

Macros, Static Registration and Code Generation


C++ 在 AI-Infra(04):多态与类型擦除——运行时如何选择实现

Polymorphism and Type Erasure


C++ 在 AI-Infra(03):模板与泛型编程

Templates and Generic Programming


C++ 在 AI-Infra(02):值、引用与所有权——对象模型与 RAII

Value Semantics, Ownership and RAII


C++ 在 AI-Infra(01 下):工程布局——命名空间、库的分层与 CMake

Project Layout: Namespaces, Library Layering and CMake


C++ 在 AI-Infra(01 上):编译模型——从一个 .cpp 到可加载的 .so

The Compilation Model: From a .cpp to a Loadable .so


C++ 在 AI-Infra:从对象模型到算子扩展(总纲)

C++ for AI-Infra, from the Object Model to Operator Extensions


面试手撕代码(19):Infra 岗手撕——并发与系统

Infra Interviews: Thread-Safe LRU, Bounded Queues, Thread Pools, Memory Pools, Blocked GEMM, Ring Allreduce, Paged KV and Token Buckets


记一个诡异的C++问题


如何确保C库可以正确被C++客户端程序调用


Java

面试手撕代码(20):系列总结与通关自测

Coding Interviews: Series Recap and Final Self-Test


面试手撕代码(13):设计题与数据结构实现

Design Problems: LRU, LFU, Trie, O(1) Random Set and Fenwick Tree


面试手撕代码(12):动态规划(二)——背包、区间、状态机、树形

Dynamic Programming II: Knapsack, Interval, State Machine and Tree DP


面试手撕代码(11):动态规划(一)——线性与二维

Dynamic Programming I: Define the State, Then the Rest Follows


面试手撕代码(10):字符串

Strings: Palindromes, KMP, Parsing, Big Numbers and Custom Ordering


面试手撕代码(09):回溯

Backtracking: Choose, Recurse, Undo — and Prune


面试手撕代码(08):堆、Top-K、区间与贪心

Heaps, Top-K, Intervals and Greedy: Sort by the Right Key


面试手撕代码(07):二分——只有一个模板

Binary Search: One Template, First Position Where the Predicate Holds


面试手撕代码(06):图——BFS / DFS / 拓扑排序 / 并查集 / 最短路

Graphs: BFS, DFS, Topological Sort, Union-Find and Dijkstra


面试手撕代码(05):二叉树

Binary Trees: Three Questions Every Recursion Must Answer


面试手撕代码(04):链表

Linked Lists: Dummy Heads, Three Pointers and the Tortoise–Hare Proof


面试手撕代码(03):栈、单调栈与单调队列

Stacks, Monotonic Stacks and Monotonic Queues: Settle When You Pop


面试手撕代码(02):双指针与滑动窗口

Two Pointers and Sliding Windows: Monotonicity Turns O(n²) into O(n)


面试手撕代码(01):数组、哈希与前缀和

Arrays, Hashing and Prefix Sums: Trade Space for a Loop


面试手撕代码:从 LeetCode 中等题到 Transformer 组件(总纲)

Coding Interviews: From LeetCode Mediums to Hand-Written Transformer Components


一个java大堆引发的『血案』


使用CompletableFuture异步编程


关于编码规范的一些建议


Java各种锁介绍


如何单元测试二方库


Java8时间处理


input too large for RSA cipher


Quartz的misfire机制


如何让一个Quartz实例不执行任务


怎样获取form-data方式POST的数据


Tomcat调优


java服务端监控平台设计


Java Attach API


JMX学习


java standalone模板


Java中如何正确的加载配置文件


Java DNS查询内部实现


如何让java程序优先使用自定义的DNS nameserver


应用如何记录集中日志


如何防止表单重复提交


配置tomcat的access_log


Quartz突然停止执行问题


Java文件读取支持timeout


JPA的事务管理器配置


Quartz工作机制


Java NIO.2


Java NIO


如何监控线上应用的运行状态


log4j详细介绍


log4j日志路径问题


使用拦截器做简单的性能监控


一个简单分页查询组件实现


优雅的Builder模式


Java虚拟机学习笔记


巧用TheadLocal


BTrace实战


循环引用序列化问题


mina学习笔记


一个简单的性能优化和防止DOS攻击的示例


return async result in java


Java并发学习笔记


高并发下额度限制问题


JVM编码


如何在系统启动时完成资源加载


Java Heap OOM问题


配置文件串串SHOW


如何使用tomcat高效调试


创建可执行的jar包


Quartz与Spring的整合-使用Spring的FactoryBean实现动态Properties


maven的resources插件


如何往HttpServletRequest中塞请求参数


Java网络IO编程


Python

Python 在 AI-Infra(08):系列总结与通关自测

Python for AI-Infra: Series Recap and Final Self-Test


Python 在 AI-Infra(07):项目工程化与生产交付

Python Project Engineering and Production Delivery


Python 在 AI-Infra(06):单元测试、问题定位与调试实践

Python Unit Testing, Troubleshooting, and Debugging


Python 在 AI-Infra(05):内存管理与优化

Python Memory Management and Optimization


Python 在 AI-Infra(04):Python的动态机制及工程实践

Python Dynamic Mechanisms and Practice


Python 在 AI-Infra(03):并发、异步与任务协作

Python Concurrency, Asynchrony, and Task Collaboration in AI Systems


Python 在 AI-Infra(02 下):类型系统——数据契约设计

Type System III: Data Contract Design


Python 在 AI-Infra(02 中):类型系统——类型信息的分发与消费

Type System II: Distributing and Consuming Type Information


Python 在 AI-Infra(02 上):类型系统——类型表达与 typing 工具箱

Type System I: Type Expression and the typing Toolbox


Python 在 AI-Infra(01 下):对象如何工作——对象模型、协议、装饰器与生成器

How Objects Work — Object Model, Protocols, Decorators and Generators


Python 在 AI-Infra(01 上):代码如何被执行——执行模型、作用域、导入与异常

How Code Runs — Execution Model, Scopes, Imports and Exceptions


Python 在 AI-Infra:从语言机制到生产交付(总纲)

Python for AI-Infra, from Language Mechanisms to Production Delivery (Overview)


算法工程师的工具箱(07):系列总结与通关自测

Tooling for AI Algorithm Engineers: Series Recap and Final Self-Test


算法工程师的工具箱(06):GPU 直觉与实验管理——两个上限、四块显存、能复现

GPU Intuition and Experiment Management: Two Ceilings, Four Memory Buckets, and Reproducibility


算法工程师的工具箱(05):Hugging Face 生态——六个库与一次 LoRA SFT 的组装

The Hugging Face Ecosystem: Six Libraries, a Six-Line LoRA SFT, and Why Reading the Source Is the Fastest Way to Learn


算法工程师的工具箱(04):PyTorch 使用层(下)——混合精度、显存的账与多卡启用

PyTorch in Use, Part 2: Mixed Precision, the Memory Ledger and Turning On Multi-GPU


算法工程师的工具箱(03):PyTorch 使用层(上)——五个对象与二十行训练循环

PyTorch in Use, Part 1: Five Objects and a Twenty-Line Training Loop


算法工程师的工具箱(02):数据科学三剑客——NumPy 的形状直觉、Pandas 的错误分析、Matplotlib 的曲线

The Scientific Python Stack: Shapes and Broadcasting in NumPy, Error Analysis in Pandas, Reading Curves in Matplotlib


算法工程师的工具箱(01):Python 使用层——读懂训练代码的语法、流式过一遍语料、把实验写成脚本

Python in Use for Algorithm Engineers: The Syntax Behind Training Code, Streaming a Corpus, and Scripting an Experiment


算法工程师的工具箱:从一个想法到一次能跑的实验(总纲)

Tooling for AI Algorithm Engineers: NumPy, PyTorch, Hugging Face and the GPU in Your Head


面试手撕代码(20):系列总结与通关自测

Coding Interviews: Series Recap and Final Self-Test


面试手撕代码(19):Infra 岗手撕——并发与系统

Infra Interviews: Thread-Safe LRU, Bounded Queues, Thread Pools, Memory Pools, Blocked GEMM, Ring Allreduce, Paged KV and Token Buckets


面试手撕代码(13):设计题与数据结构实现

Design Problems: LRU, LFU, Trie, O(1) Random Set and Fenwick Tree


面试手撕代码(12):动态规划(二)——背包、区间、状态机、树形

Dynamic Programming II: Knapsack, Interval, State Machine and Tree DP


面试手撕代码(11):动态规划(一)——线性与二维

Dynamic Programming I: Define the State, Then the Rest Follows


面试手撕代码(10):字符串

Strings: Palindromes, KMP, Parsing, Big Numbers and Custom Ordering


面试手撕代码(09):回溯

Backtracking: Choose, Recurse, Undo — and Prune


面试手撕代码(08):堆、Top-K、区间与贪心

Heaps, Top-K, Intervals and Greedy: Sort by the Right Key


面试手撕代码(07):二分——只有一个模板

Binary Search: One Template, First Position Where the Predicate Holds


面试手撕代码(06):图——BFS / DFS / 拓扑排序 / 并查集 / 最短路

Graphs: BFS, DFS, Topological Sort, Union-Find and Dijkstra


面试手撕代码(05):二叉树

Binary Trees: Three Questions Every Recursion Must Answer


面试手撕代码(04):链表

Linked Lists: Dummy Heads, Three Pointers and the Tortoise–Hare Proof


面试手撕代码(03):栈、单调栈与单调队列

Stacks, Monotonic Stacks and Monotonic Queues: Settle When You Pop


面试手撕代码(02):双指针与滑动窗口

Two Pointers and Sliding Windows: Monotonicity Turns O(n²) into O(n)


面试手撕代码(01):数组、哈希与前缀和

Arrays, Hashing and Prefix Sums: Trade Space for a Loop


面试手撕代码:从 LeetCode 中等题到 Transformer 组件(总纲)

Coding Interviews: From LeetCode Mediums to Hand-Written Transformer Components


Python 中如何定义数据类并做校验:从 __init__、dataclass 到 Pydantic


python2.x的一个需要注意的地方


BTrace

BTrace实战


AI

Prompt 与上下文工程:模型这一步该看到什么(总纲)

Prompt and Context Engineering: What the Model Should See at This Step


模型作为组件(07):系列总结与通关自测

The Model as a Component: Series Recap and Final Self-Test


模型作为组件(06):客户端工程——重试、超时、幂等、限流与流式解析

Client Engineering for LLM APIs: Retries, Timeouts, Idempotency, Rate Limits and Stream Parsing


模型作为组件(05):选型——榜单的失真、自己的评测集与供应商的弃用周期

Model Selection Beyond Leaderboards: Your Own Evals, Open vs Closed, Routing and Deprecation Cycles


模型作为组件(04):成本与延迟的账——一次调用花多少钱、慢在哪一段

The Token Cost and Latency Ledger for LLM Applications


模型作为组件(03):API 契约(二)——推理模型:thinking、effort 与跨轮的推理状态

The LLM API Contract, Part 2: Reasoning Models — Thinking, Effort and Cross-turn State


模型作为组件(02):API 契约(一)——消息、工具调用、结构化输出与流式,四家 API 的共同骨架

The LLM API Contract, Part 1: Messages, Tool Calls, Structured Output and Streaming


模型作为组件(01):失效模式——非确定性、幻觉、上下文与越界,把「模型会出错」拆成七条可检测的性质

Failure Modes of LLMs as Components: Nondeterminism, Hallucination, Context and Overreach


模型作为组件:契约、失效模式与选型(总纲)

The Model as a Component: Contract, Failure Modes and Selection


AI-Infra 开源贡献指南(05):系列总结与通关自测

Contributing to AI-Infra Open Source: Series Recap and Final Self-Test


AI-Infra 开源贡献指南(04):两个真实 PR 的完整走读——PyTorch 与 vLLM

Two Real Pull Requests, End to End: One in PyTorch, One in vLLM


AI-Infra 开源贡献指南(03):做出一个能被合入的改动

Landing a Mergeable Change: Diff, Tests, Benchmarks, PR, CI and Review


AI-Infra 开源贡献指南(02):找到切入点——从 issue、RFC 到性能回归

Finding Your Entry Point: Issues, RFCs, Roadmaps, CI Failures and Regressions


AI-Infra 开源贡献指南(01):读懂一个百万行的代码库

Reading a Million-Line Codebase: Maps, Entry Points, Symbols, Tests and History


AI-Infra 开源贡献指南(总纲)

A Guide to Contributing to AI-Infra Open Source Projects


AI 平台工程(09):系列总结与通关自测

AI Platform Engineering: Series Recap and Final Self-Test


AI 平台工程(08):可观测、成本与 FinOps

Observability, Cost and FinOps: from DCGM to the Token Bill


AI 平台工程(07):模型网关与多租户——路由、配额与灰度

The Model Gateway: KV-Aware Routing, Multi-Tenancy, Quotas and Canaries


AI 平台工程(06):Serving 平台——从 InferenceService 到 llm-d

Serving Platforms: KServe, Triton, Ray Serve, LeaderWorkerSet and llm-d


AI 平台工程(05):网络与存储——RDMA 进容器、并行文件系统与 checkpoint I/O

Networking and Storage: RDMA in Containers, Parallel File Systems and Checkpoint I/O


AI 平台工程(04):GPU 共享与切分——MIG、时间片、MPS 与 HAMi

Sharing and Partitioning GPUs: MIG, Time-Slicing, MPS and HAMi


AI 平台工程(03):AI 任务调度——gang scheduling、队列与拓扑感知

Scheduling AI Jobs: Gang Scheduling, Queues, Quotas and Topology Awareness


AI 平台工程(02):容器里的 GPU——驱动、CUDA、device plugin 与镜像

GPUs in Containers: Driver, CUDA, Device Plugin, DRA and Images


AI 平台工程(01):引擎的需求清单与平台的整体架构

What Engines Demand from the Platform, and the Platform's Two Layers


AI 平台工程:资源层与交付层(总纲)

AI Platform Engineering: the Resource Layer and the Delivery Layer


扩散模型推理基础设施(10):系列总结与通关自测

Diffusion Model Inference Infrastructure: Series Recap and Final Self-Test


扩散模型推理基础设施(09):配置、评测与排障——从一张卡的推导到一条伪影的排查

Configuration, Evaluation and Troubleshooting for Diffusion Inference (with Series Summary)


扩散模型推理基础设施(08):三个引擎的对照导读——同一张图的请求在 SGLang Diffusion、vLLM-Omni 与 xDiT 里各走过什么

Three Engines Compared: One Request Through SGLang Diffusion, vLLM-Omni and xDiT


扩散模型推理基础设施(07):serving 形态——请求、批、三段分离、附件、异步任务与成本

Serving Diffusion Models: Request Shapes, Batching, Disaggregation, Add-ons, Async Jobs and Cost


扩散模型推理基础设施(06):少步与自回归——把步数变成系统参数

Few-Step and Autoregressive Generation: When the Step Count Becomes a System Parameter


扩散模型推理基础设施(05):多卡并行——序列并行、CFG 并行与 PipeFusion,为什么不是张量并行

Multi-GPU Diffusion Inference: Sequence Parallelism, CFG Parallelism and PipeFusion


扩散模型推理基础设施(04):视频——长序列 attention 的账与稀疏化

Video Diffusion: The Long-Sequence Attention Bill and How Sparsity Pays It


扩散模型推理基础设施(03):跨步冗余——TeaCache、First-Block Cache 一族的缓存与跳步

Temporal Redundancy Across Denoising Steps: TeaCache, First-Block Cache and Friends


扩散模型推理基础设施(02):单卡执行——attention 后端、编译、FP8 / INT4 与 offload

Single-GPU Execution: Attention Backends, Compilation, FP8 / INT4 and Offloading


扩散模型推理基础设施(01):负载画像——一次生成在 GPU 上发生什么

Workload Anatomy: FLOPs, Bytes and Seconds of One Diffusion Generation


扩散模型推理基础设施:图像与视频生成的 serving(总纲)

Diffusion Model Inference Infrastructure: Serving Image and Video Generation


RL 后训练基础设施(09):系列总结与通关自测

RL Post-Training Infrastructure: Series Recap and Final Self-Test


RL 后训练基础设施(08):配置、可观测与排障——从一张卡的比例到一条 hang 的排查

Configuration, Observability and Troubleshooting for RL Post-Training Systems


RL 后训练基础设施(07):verl 源码导读——从一个 GRPO 配置追到每个 worker

Reading verl: From One GRPO Config to Every Worker


RL 后训练基础设施(06):Agentic rollout——多轮、工具、沙箱与环境服务

Agentic Rollout: Multi-Turn Trajectories, Tools, Sandboxes and Environment Services


RL 后训练基础设施(05):异步与 off-policy——把同步的墙拆掉之后要补什么

Asynchrony and Off-Policy: What You Owe After Tearing Down the Synchronization Wall


RL 后训练基础设施(04):权重同步——从训练分片到推理分片

Weight Synchronization: From Training Shards to Inference Shards


RL 后训练基础设施(03):共置——训练器与推理引擎在同一组 GPU 上共存

Colocation: Handing GPU Memory Back and Forth Between Trainer and Rollout Engine


RL 后训练基础设施(02):系统形态——共置、分离与异步

RL System Topologies: Colocated, Disaggregated and Asynchronous


RL 后训练基础设施(01):负载画像——一步 RL 里发生什么

Anatomy of an RL Step: Rollout, Reward and Train


RL 后训练基础设施:rollout 与训练如何共享一组 GPU(总纲)

RL Post-Training Infrastructure: How Rollout and Training Share the Same GPUs


大模型推理系统揭秘(16):系列总结与通关自测

Deep Dive into vLLM: Series Recap and Final Self-Test


大模型推理系统揭秘(15):vLLM 与 SGLang:同一个请求穿过两套 Serving 系统

vLLM vs SGLang: One Request Through Two LLM Serving Systems


大模型推理系统揭秘(14):回到源码:一次请求在 vLLM 内部的真实旅程


大模型推理系统揭秘(13):Serving Infra 的下一站:从模型执行器到分布式智能操作系统


大模型推理系统揭秘(12):硬件解耦:如何不让芯片差异污染 Serving 核心?


大模型推理系统揭秘(11):请求形态的扩展:multi-LoRA 与多模态


大模型推理系统揭秘(10):模型适配:如何跟上变化极快的模型世界?


大模型推理系统揭秘(09):PD 分离:从资源混部走向计算解耦


大模型推理系统揭秘(08):Multi-GPU:一张卡不够时如何扩展?


大模型推理系统揭秘(07):解码的扩展:采样、投机解码与结构化输出


大模型推理系统揭秘(06):GPU 执行:如何让每个 Token 算得更快?


大模型推理系统揭秘(05):KV Cache:LLM Serving 的第一号内存问题


大模型推理系统揭秘(04):Scheduler:GPU 这一轮到底给谁用?


大模型推理系统揭秘(03):鸟瞰 vLLM:一个请求如何穿过整个推理系统?


大模型推理系统揭秘(02):如何衡量一个 LLM Serving 系统?


大模型推理系统揭秘(01):为什么 LLM Serving 比传统 DL 推理难?


大模型推理系统揭秘:从 vLLM 看 LLM Serving Infra 核心技术(总纲)


大规模训练工程(09):系列总结与通关自测

Large-Scale Training Engineering: Series Recap and Final Self-Test


大规模训练工程(08):长时训练的可观测与运维——从指标到 hang 排查

Observability and Operations for Long-Running Training


大规模训练工程(07):训练稳定性与数据管线——loss spike、梯度范数、数据混合与流式加载

Training Stability and the Data Pipeline


大规模训练工程(06):容错与弹性——故障率数学、straggler、SDC 与弹性训练

Fault Tolerance and Elasticity: Failure Math, Stragglers, SDC and Elastic Training


大规模训练工程(05):分布式 checkpoint——格式、异步保存与重分片恢复

Distributed Checkpoint: Format, Async Save and Resharding


大规模训练工程(04):千卡配置实战——并行搭配、micro-batch、激活重计算与 MFU 调优

Configuring a Thousand-GPU Job: Parallelism, Micro-batch, Recompute and MFU


大规模训练工程(03):三个框架——Megatron-LM、DeepSpeed 与 torchtitan 的架构对比与源码导读

Megatron-LM, DeepSpeed and torchtitan: Architecture and Source Guide


大规模训练工程(02):并行策略全景——每种并行切的是哪种状态

A Map of Parallelism: Which State Does Each Strategy Shard


大规模训练工程(01):训练任务的状态解剖——显存账与 MFU

Anatomy of Training State: Memory Accounting and MFU


大规模训练工程:从并行策略到容错恢复(总纲)

Large-Scale Training Engineering, from Parallelism to Fault Tolerance


通信与互联(09):系列总结与通关自测

Communication and Interconnect: Series Recap and Final Self-Test


通信与互联(08):MoE 的通信——all-to-all、DeepEP 与 GPU 发起的通信

Communication for MoE: All-to-All, DeepEP and GPU-Initiated Networking


通信与互联(07):推理侧的通信——custom all-reduce 与 KV 传输

Communication on the Inference Side: Custom All-Reduce and KV Cache Transfer


通信与互联(06):nccl-tests、调优与排障——从带宽曲线到 hang

nccl-tests, Tuning and Debugging Hangs: From Bandwidth Curves to Flight Recorder


通信与互联(05):PyTorch 的通信栈——ProcessGroupNCCL、stream 语义与计算通信重叠

PyTorch's Communication Stack: ProcessGroupNCCL, Stream Semantics, and Compute-Communication Overlap


通信与互联(04):NCCL 架构——拓扑探测、channel、算法与协议

NCCL Architecture: Topology Detection, Channels, Algorithms and Protocols


通信与互联(03):RDMA 与 GPUDirect——绕过 CPU 和主机内存的数据通路

RDMA and GPUDirect: Bypassing the CPU and Host Memory


通信与互联(02):硬件互联——PCIe、NVLink、NVSwitch 与网络拓扑

Hardware Interconnect: PCIe, NVLink, NVSwitch and Network Topology


通信与互联(01):集合通信原语与代价模型——α-β 模型与 ring all-reduce

Collective Communication Primitives and the Alpha-Beta Cost Model: Deriving Ring All-Reduce


通信与互联:从 NCCL 到 RDMA(总纲)

Communication and Interconnect for AI-Infra, from NCCL to RDMA


LoRA 专题(04):系列总结与通关自测

LoRA Series: Recap and Final Self-Test


LoRA 专题(03):工程:adapter 文件、合并、多 LoRA 服务与参考模型

LoRA 03: In Production — Adapter Files, Merging, Multi-LoRA Serving and the Free Reference Model


LoRA 专题(02):选参:r、target_modules、alpha、lr 与 QLoRA / DoRA / PiSSA 各让模型变成什么

LoRA 02: Choosing r, target_modules, alpha and lr — and What QLoRA, DoRA and PiSSA Each Change


LoRA 专题(01):低秩假设:为什么两个瘦矩阵够用,以及它省了哪几本账

LoRA 01: The Low-Rank Hypothesis — Why Two Thin Matrices Suffice, and Exactly What They Save


LoRA 专题:SFT 的默认微调方式,从低秩假设到多租户服务(总纲)

LoRA, the Default Way to Fine-Tune: From the Low-Rank Hypothesis to Multi-Tenant Serving — Series Overview


读 Hugging Face 源码(05):系列总结与通关自测

Reading the Hugging Face Source: Series Recap and Final Self-Test


读 Hugging Face 源码(04):peft 与 trl——LoRA 怎么挂上去,SFT / DPO / GRPO 的 loss 各在哪一行

Inside peft and trl: get_peft_model, the LoRA Linear, SFT Labels and Packing, dpo_loss and GRPO Advantages


读 Hugging Face 源码(03):tokenizers 与 datasets——从 messages 到 input_ids,从 Arrow 文件到 collate_fn

Inside tokenizers and datasets: the Rust Pipeline, Chat Templates, Arrow Tables, map and Fingerprints


读 Hugging Face 源码(02):generate——一次采样的完整调用链

Inside transformers, Part 2: generate, GenerationConfig, LogitsProcessors, StoppingCriteria and the Decode Loop


读 Hugging Face 源码(01):transformers 模型侧——from_pretrained 怎么把三个文件变成 nn.Module,forward 怎么走到 loss

Inside transformers, Part 1: from_pretrained, the Decoder Stack, Attention Dispatch, KV Cache and the Loss


读 Hugging Face 源码:从 from_pretrained 到 GRPO 的 loss(总纲)

Reading the Hugging Face Source: transformers, tokenizers, datasets, peft and trl, One Call Chain at a Time


GPU Kernel 工程(11):系列总结与通关自测

GPU Kernel Engineering: Series Recap and Final Self-Test


GPU Kernel 工程(10):剖析、测试与贡献——把 kernel 做成产品

Profiling, Testing and Contributing: Turning a Kernel into a Product


GPU Kernel 工程(09):量化与融合 kernel——推理系统的其余部分

Quantized and Fused Kernels: The Rest of the Inference Stack


GPU Kernel 工程(08):Attention Kernel——FlashAttention 与 PagedAttention

Attention Kernels: FlashAttention and PagedAttention from Derivation to Code


GPU Kernel 工程(07):Triton——块级编程与编译器的边界

Triton: Block-Level Programming and Where the Compiler Stops


GPU Kernel 工程(06):Tensor Core、CUTLASS 与 CuTe

Tensor Cores, CUTLASS and CuTe: Programming the Matrix Units


GPU Kernel 工程(05):GEMM——从 naive 到分块

GEMM from Naive to Tiled: Reaching the Compute Ceiling on CUDA Cores


GPU Kernel 工程(04):共享内存与 reduction——softmax、LayerNorm 与 online softmax

Shared Memory and Reductions: Softmax, LayerNorm and Online Softmax


GPU Kernel 工程(03):访存合并与 elementwise kernel

Memory Coalescing and Elementwise Kernels: Hitting the Bandwidth Ceiling


GPU Kernel 工程(02):CUDA 编程模型与第一个 kernel

The CUDA Programming Model and Your First Kernel, Measured


GPU Kernel 工程(01):GPU 为什么这样设计——硬件结构与 Roofline

Why GPUs Look the Way They Do: Architecture and the Roofline Model


GPU Kernel 工程:从 CUDA 执行模型到 FlashAttention(总纲)

GPU Kernel Engineering, from the CUDA Execution Model to FlashAttention


多模态(10):系列总结与通关自测

Multimodal Models: Series Recap and Final Self-Test


多模态(09):自回归图像生成与统一模型

Autoregressive Image Generation and Unified Understanding-Generation Models


多模态(08):Latent diffusion、DiT 与文生图配方

Latent Diffusion, DiT and How Text-to-Image Models Are Built


多模态(07):扩散模型(下):score matching、flow matching 与 classifier-free guidance

Diffusion II: Score Matching, Flow Matching, Noise Schedules, and Classifier-Free Guidance


多模态(06):扩散模型(上):DDPM——加噪、去噪与「预测噪声」

Diffusion I: DDPM — Forward Noising, Reverse Denoising, the ELBO, and DDIM


多模态(05):语音(下):语音理解、语音生成与全双工

Speech II: Speech Understanding, Speech Generation (TTS), Omni Models and Full-Duplex Dialogue


多模态(04):语音(上):从波形到 token——mel 谱、Whisper 与神经 codec

Speech I: From Waveform to Tokens — Mel Spectrograms, Whisper, and Neural Codecs with RVQ


多模态(03):VLM 的训练:数据、阶段与评测

Training a VLM: Data, Stages, Evaluation and Hallucination


多模态(02):VLM 的结构:connector、注入方式与动态分辨率

VLM Architecture: Connectors, Injection Methods and Dynamic Resolution


多模态(01):视觉编码器:CLIP、SigLIP 与自监督 ViT

Vision Encoders: CLIP, SigLIP and Self-Supervised ViTs


多模态:从视觉编码器到扩散模型(总纲)

Multimodal Models: From Vision Encoders to Diffusion


高效推理与压缩(07):系列总结与通关自测

Efficient Inference and Model Compression: Series Recap and Final Self-Test


高效推理与压缩(06):剪枝、深度缩放与小模型配方

Pruning, Depth Scaling and How Small Models Are Made


高效推理与压缩(05):KV cache 压缩:量化、驱逐与稀疏 attention

KV Cache Compression: Quantization, Eviction and Sparse Attention


高效推理与压缩(04):量化感知训练、低比特与量化模型的评测

Quantization-Aware Training, Extreme Low-Bit and How to Evaluate a Quantized Model


高效推理与压缩(03):训练后量化:误差模型、GPTQ、AWQ 与旋转

Post-Training Quantization: Error Models, GPTQ, AWQ and Rotation


高效推理与压缩(02):投机解码:草稿、接受率与树

Speculative Decoding: Drafters, Acceptance Rates and Draft Trees


高效推理与压缩(01):解码策略、采样与约束生成

Decoding Strategies: Sampling, Truncation, Penalties and Constrained Generation


高效推理与压缩(算法侧):解码、投机、量化与 KV(总纲)

Efficient Inference and Model Compression: The Algorithm Side


算法工程师的实验方法论:用有限的算力得出可信的结论

Experimental Methodology for AI Algorithm Engineers: Credible Conclusions on a Finite Compute Budget


后训练(09):系列总结与通关自测

Post-Training: Series Recap and Final Self-Test


后训练(08):评测:benchmark、LLM-as-judge、Arena 与污染

Evaluating LLMs: Benchmarks, LLM-as-Judge, Arenas and Contamination


后训练(07):蒸馏:logits 级、序列级与 on-policy

Knowledge Distillation for LLMs: Logit-Level, Sequence-Level and On-Policy


后训练(06):Agent 与工具调用的 RL:多轮环境、轨迹数据与延后的奖励

Agentic RL: Multi-Turn Environments, Trajectory Data and Delayed Rewards


后训练(05):推理模型与可验证奖励:R1 的配方、PRM 与 test-time compute

Reasoning Models and Verifiable Rewards: The R1 Recipe, Process Reward Models and Test-Time Compute


后训练(04):离线 RL:从 RLHF 目标推出 DPO 及其变体

Offline Preference Optimization: Deriving DPO from the RLHF Objective, and Its Family


后训练(03):在线 RL:PPO、GRPO 与 RLHF 三件套

Online RL for LLMs: PPO, GRPO and the Policy–Reward–Reference Trio


后训练(02):偏好数据与奖励模型:Bradley-Terry、pairwise loss 与 reward hacking

Preference Data and Reward Models: Bradley-Terry, Pairwise Loss and Reward Hacking


后训练(01):SFT:指令数据、chat template、loss mask 与参数高效微调

Supervised Fine-Tuning: Instruction Data, Chat Templates, Loss Masking and Parameter-Efficient Fine-Tuning


后训练:从 SFT 到可验证奖励(总纲)

Post-Training: From Supervised Fine-Tuning to Verifiable Rewards


预训练(06):系列总结与通关自测

Pretraining: Series Recap and Final Self-Test


预训练(05):训练配方与稳定性:学习率、batch、调度与 loss spike

Pretraining Recipes and Training Stability: Learning Rate, Batch Size, Schedules and Loss Spikes


预训练(04):预训练数据工程:从 Common Crawl 到 15T token,去重、过滤与配比的账

Pretraining Data Engineering: From Common Crawl to 15T Tokens, the Arithmetic of Deduplication, Filtering and Mixing


预训练(03):Scaling law:从 Chinchilla 到"过训练",算力怎么分给参数与数据

Scaling Laws: From Chinchilla to Over-Training, Splitting Compute between Parameters and Data


预训练(02):分词与词表:BPE、词表大小与 token 效率

Tokenizers and Vocabulary: BPE, Vocabulary Size and Token Efficiency


预训练(01):一次预训练是怎么跑起来的:从两个网页文件到一个会续写英文的模型

Pretraining End to End: From Two Common Crawl Files to a Model That Writes English, on a Laptop


预训练:从 tokenizer 到训练配方(总纲)

Pretraining: Tokenizers, Scaling Laws, Data Pipelines and Training Recipes


Transformer 与 LLM:系列总结与通关自测

Transformers and LLMs: Series Recap and Final Self-Test


Transformer 与 LLM(13):多模态:vision encoder 的算量与 image token 的 KV 代价

Multimodal LLMs: The Cost of Vision Encoders and Image Tokens


Transformer 与 LLM(12):量化、投机解码与 LoRA

Quantization, Speculative Decoding and LoRA: Three Ways to Reshape the Computation


Transformer 与 LLM(11):浮点格式、数值稳定性与混合精度

Floating-Point Formats, Numerical Stability and Mixed Precision


Transformer 与 LLM(10):前向的算量与访存量

FLOPs, Bytes and Roofline: Prefill versus Decode


Transformer 与 LLM(09):MTP——改训练目标而不改主干的多 token 预测

Multi-Token Prediction: Denser Supervision from the Same Data, and a Free Speculative Draft


Transformer 与 LLM(08):MoE 的路由、激活参数量与通信形态

Mixture of Experts: Routing, Active Parameters and Communication Patterns


Transformer 与 LLM(07):位置编码与长上下文

Positional Encoding and Long Context: RoPE Wavelengths, Extrapolation and Cost


Transformer 与 LLM(06):Attention 变体与 KV cache

Attention Variants and the KV Cache: Deriving MHA, GQA, MQA and MLA


Transformer 与 LLM(05):Transformer 解剖与参数量

Transformer Anatomy and Parameter Count: From config.json to 8.03B


Transformer 与 LLM(04):手搓 GPT(下)——nanoGPT train.py 与训一个会续写的模型

Building GPT by Hand, Part 2: train.py Line by Line, Then Train One on Shakespeare


Transformer 与 LLM(03):手搓 GPT(上)——nanoGPT model.py 逐行解析

Building GPT by Hand, Part 1: Every Line of nanoGPT's model.py


Transformer 与 LLM(02):一个 token 的旅程——训练侧与推理侧

The Dynamic View: What Happens to a Token During Training and During Inference


Transformer 与 LLM(01):Transformer 长什么样——从一句话到下一个 token

The Static View: Every Box in a Decoder-only Transformer, Computed by Hand


Transformer 与 LLM:结构、实现与算量(总纲)

Transformers and LLMs: Architecture, Implementation and Arithmetic


深度学习基础(07):系列总结与通关自测

Deep Learning Foundations: Series Recap and Final Self-Test


深度学习基础(06):RNN——从 LSTM 到 attention 的诞生

RNN: Backpropagation Through Time, the LSTM Gate as a Residual Path, and How Attention Was Born from the seq2seq Bottleneck


深度学习基础(05):CNN——从 LeNet 到 ResNet,再到 ViT

CNN: Convolution as a Constrained Linear Layer, the ResNet Legacy and How ViT Turns Images into Tokens


深度学习基础(04):正则化与泛化——为什么参数比样本多却不过拟合

Regularization and Generalization: Double Descent, Implicit Bias, Dropout, Weight Decay and When Overfitting Comes Back


深度学习基础(03):优化器——从 SGD 到 AdamW 与学习率调度

Optimizers: From SGD to AdamW, Warmup, Clipping and the Batch-Learning-Rate Scaling Rule


深度学习基础(02):训练为什么不稳定——初始化、归一化与残差

Why Deep Training Is Unstable: Variance Propagation, Initialization, Normalization and Residual Connections


深度学习基础(01):反向传播——手推一个两层网络

Backpropagation by Hand: Shapes, the 2x Rule and Why Activations Must Be Saved


深度学习基础:从反向传播到残差(总纲)

Deep Learning Foundations: From Backpropagation to Residual Connections


LLM 时代的经典机器学习(11):系列总结与通关自测

Classical Machine Learning in the LLM Era: Series Recap and Final Self-Test


LLM 时代的经典机器学习(10):评估——从混淆矩阵到 judge 的一致性

Evaluation: Confusion Matrices, Thresholds, Calibration, Paired Tests, and Every LLM Pitfall They Predict


LLM 时代的经典机器学习(09):去重——MinHash 与 LSH 的概率

Deduplication: Jaccard, MinHash as an Unbiased Estimator, and the LSH S-Curve


LLM 时代的经典机器学习(08):降维——PCA、SVD、t-SNE 与 embedding 的各向异性

Dimensionality Reduction: PCA, SVD, t-SNE / UMAP, and the Anisotropy of Embedding Spaces


LLM 时代的经典机器学习(07):聚类——K-Means、DBSCAN 与「这批语料里有什么」

Clustering: K-Means, DBSCAN, Hierarchical Clustering, and What a Corpus Looks Like in Embedding Space


LLM 时代的经典机器学习(06):集成——随机森林、梯度提升与数据质量分类器的算力账

Ensembles: Random Forests, Gradient Boosting, and Why Data Filters Use Small Models


LLM 时代的经典机器学习(05):SVM 与核方法——最大间隔、核技巧与 attention 的远亲

Support Vector Machines and Kernels: Maximum Margin, the Kernel Trick, and Why Attention Is a Kernel Smoother


LLM 时代的经典机器学习(04):三个基础分类器——朴素贝叶斯、KNN 与决策树

Three Basic Classifiers: Naive Bayes Computes Probabilities, KNN Finds Neighbours, Decision Trees Ask Questions


LLM 时代的经典机器学习(03):逻辑回归与奖励模型——每个分类头的原型

Logistic Regression and Reward Models: The Skeleton of Every Classification Head


LLM 时代的经典机器学习(02):线性回归——最小二乘、梯度下降与 Ridge / Lasso

Linear Regression: Least Squares, Gradient Descent, and Why Ridge Is Weight Decay


LLM 时代的经典机器学习(01):什么是学习——划分、泛化、过拟合与偏差-方差

What Is Learning: Splits, Generalization, Overfitting and the Bias-Variance Decomposition


LLM 时代还要学经典机器学习吗:只讲它在哪里重现(总纲)

Classical Machine Learning in the LLM Era: What Survives and Where It Reappears


PyTorch 深度实践(11):系列总结与通关自测

Deep Dive into PyTorch: Series Recap and Final Self-Test


PyTorch 深度实践(10):PyTorch 的工程体系——一次改动如何安全地到达用户

The Engineering System of PyTorch: How a Change Travels Safely from Commit to Production


PyTorch 深度实践(09):分布式 PyTorch

Distributed Training in PyTorch: Collectives, DDP, FSDP, TP, PP, CP and EP


PyTorch 深度实践(08):性能优化与调试

Performance Optimization and Debugging in PyTorch


PyTorch 深度实践(07):编译执行与图优化

Compilation and Graph Optimization in PyTorch


PyTorch 深度实践(06):C++ 扩展与自定义算子

C++ Extensions and Custom Operators in PyTorch


PyTorch 深度实践(05):Dispatcher 与算子系统

The Dispatcher and Operator System in PyTorch


PyTorch 深度实践(04):nn.Module 与训练系统

nn.Module and Training Systems in PyTorch


PyTorch 深度实践(03):自动求导与动态计算图

Autograd and Dynamic Computation Graphs in PyTorch


PyTorch 深度实践(02):Tensor 与内存布局

Tensor Abstraction and Memory Layout in PyTorch


PyTorch 深度实践(01):PyTorch 整体介绍

PyTorch Overall Introduction


PyTorch 深度实践:从 Tensor 到深度学习运行时(总纲)

Deep Dive into PyTorch, from Tensor to Deep Learning Runtime


C++ 在 AI-Infra(09):系列总结与通关自测

C++ for AI-Infra: Series Recap and Final Self-Test


C++ 在 AI-Infra(08):构建、调试与测试工具链

Build, Debug and Test Toolchain


C++ 在 AI-Infra(07):与 Python 之间——pybind11、Python C API 与 ABI

pybind11, the Python C API and ABI


C++ 在 AI-Infra(06):并发、内存模型、TLS 与守卫

Concurrency, Memory Model, TLS and Guards


C++ 在 AI-Infra(05):宏、静态注册与代码生成

Macros, Static Registration and Code Generation


C++ 在 AI-Infra(04):多态与类型擦除——运行时如何选择实现

Polymorphism and Type Erasure


C++ 在 AI-Infra(03):模板与泛型编程

Templates and Generic Programming


C++ 在 AI-Infra(02):值、引用与所有权——对象模型与 RAII

Value Semantics, Ownership and RAII


C++ 在 AI-Infra(01 下):工程布局——命名空间、库的分层与 CMake

Project Layout: Namespaces, Library Layering and CMake


C++ 在 AI-Infra(01 上):编译模型——从一个 .cpp 到可加载的 .so

The Compilation Model: From a .cpp to a Loadable .so


C++ 在 AI-Infra:从对象模型到算子扩展(总纲)

C++ for AI-Infra, from the Object Model to Operator Extensions


算法工程师的工具箱(07):系列总结与通关自测

Tooling for AI Algorithm Engineers: Series Recap and Final Self-Test


算法工程师的工具箱(06):GPU 直觉与实验管理——两个上限、四块显存、能复现

GPU Intuition and Experiment Management: Two Ceilings, Four Memory Buckets, and Reproducibility


算法工程师的工具箱(05):Hugging Face 生态——六个库与一次 LoRA SFT 的组装

The Hugging Face Ecosystem: Six Libraries, a Six-Line LoRA SFT, and Why Reading the Source Is the Fastest Way to Learn


算法工程师的工具箱(04):PyTorch 使用层(下)——混合精度、显存的账与多卡启用

PyTorch in Use, Part 2: Mixed Precision, the Memory Ledger and Turning On Multi-GPU


算法工程师的工具箱(03):PyTorch 使用层(上)——五个对象与二十行训练循环

PyTorch in Use, Part 1: Five Objects and a Twenty-Line Training Loop


算法工程师的工具箱(02):数据科学三剑客——NumPy 的形状直觉、Pandas 的错误分析、Matplotlib 的曲线

The Scientific Python Stack: Shapes and Broadcasting in NumPy, Error Analysis in Pandas, Reading Curves in Matplotlib


算法工程师的工具箱(01):Python 使用层——读懂训练代码的语法、流式过一遍语料、把实验写成脚本

Python in Use for Algorithm Engineers: The Syntax Behind Training Code, Streaming a Corpus, and Scripting an Experiment


算法工程师的工具箱:从一个想法到一次能跑的实验(总纲)

Tooling for AI Algorithm Engineers: NumPy, PyTorch, Hugging Face and the GPU in Your Head


算法工程师的数学(09):系列总结与通关自测

Mathematics for AI Algorithm Engineers: Series Recap and Final Self-Test


算法工程师的数学(08):统计推断与拟合——评测的置信区间与 scaling law

Statistical Inference and Curve Fitting: Confidence Intervals for Benchmarks and How Scaling Laws Are Fit


算法工程师的数学(07):导数、梯度与链式法则——softmax 的梯度与策略梯度

Derivatives, Gradients and the Chain Rule: The Softmax Gradient and the Policy Gradient


算法工程师的数学(06):熵、交叉熵与 KL——从困惑度到 DPO

Entropy, Cross-Entropy and KL Divergence: From Perplexity to the DPO Loss


算法工程师的数学(05):从最大似然到交叉熵——第一个要会推的 loss

From Maximum Likelihood to Cross-Entropy: The First Loss You Should Be Able to Derive


算法工程师的数学(04):概率入门——语言模型是一个条件分布

Probability Basics: A Language Model Is a Conditional Distribution


算法工程师的数学(03):正交与旋转、特征值与 SVD——从 RoPE 到 LoRA

Orthogonal Matrices, Rotations, Eigenvalues and SVD: Why RoPE Encodes Relative Position and Why LoRA Works


算法工程师的数学(02):内积、范数与余弦相似度

Inner Product, Norms and Cosine Similarity: One Language for Attention, Retrieval, Regularization and Quantization Error


算法工程师的数学(01):向量、矩阵与形状——一个 token 过一层要算多少

Vectors, Matrices and Shapes: How Much Compute Does One Token Through One Layer Take


算法工程师的数学:读公式不卡壳的最小集(总纲)

Mathematics for AI Algorithm Engineers: The Minimal Set to Read Papers and Derive Losses


AI 应用工程师学习地图:在非确定性组件之上做可靠产品

A Learning Roadmap for AI Application Engineers


AI-Infra 工程师学习地图:从后端工程师到基础设施贡献者

A Learning Roadmap for AI Infrastructure Engineers


AI 算法工程师学习地图:从数学基础到大模型训练

A Learning Roadmap for AI Algorithm Engineers in the LLM Era


AI 全栈学习地图:造模型、跑模型、用模型的三张图

One System, Three Roles — an Overview of the Three AI Learning Roadmaps


面试手撕代码(20):系列总结与通关自测

Coding Interviews: Series Recap and Final Self-Test


面试手撕代码(18):手撕经典 ML 与评测指标

Classical ML and Metrics by Hand: k-means, Logistic Regression, KNN, PCA, AUC, NDCG, conv2d, NMS


面试手撕代码(17):手撕损失函数与训练算法

Losses and Training Algorithms by Hand: CE, KL, InfoNCE, DPO, PPO/GAE, GRPO, AdamW, Schedules and LoRA


面试手撕代码(16):手撕 tokenizer 与解码

Tokenizer and Decoding by Hand: BPE, Sampling, Beam Search, Reservoir Sampling and Speculative Acceptance


面试手撕代码(15):手撕 Transformer block 与反向传播

Transformer Block and Backprop by Hand: LayerNorm, SwiGLU, Parameter Counting, Gradients and Micrograd


面试手撕代码(14):手撕 attention 家族

Attention from Scratch: Softmax, SDPA, Multi-Head, GQA, RoPE, KV Cache and Online Softmax


面试手撕代码:从 LeetCode 中等题到 Transformer 组件(总纲)

Coding Interviews: From LeetCode Mediums to Hand-Written Transformer Components


AI基础架构:从大数据到深度学习


RDD

Spark RDD


Parameter Server

参数服务器


Interview

面试手撕代码(20):系列总结与通关自测

Coding Interviews: Series Recap and Final Self-Test


面试手撕代码(19):Infra 岗手撕——并发与系统

Infra Interviews: Thread-Safe LRU, Bounded Queues, Thread Pools, Memory Pools, Blocked GEMM, Ring Allreduce, Paged KV and Token Buckets


面试手撕代码(18):手撕经典 ML 与评测指标

Classical ML and Metrics by Hand: k-means, Logistic Regression, KNN, PCA, AUC, NDCG, conv2d, NMS


面试手撕代码(17):手撕损失函数与训练算法

Losses and Training Algorithms by Hand: CE, KL, InfoNCE, DPO, PPO/GAE, GRPO, AdamW, Schedules and LoRA


面试手撕代码(16):手撕 tokenizer 与解码

Tokenizer and Decoding by Hand: BPE, Sampling, Beam Search, Reservoir Sampling and Speculative Acceptance


面试手撕代码(15):手撕 Transformer block 与反向传播

Transformer Block and Backprop by Hand: LayerNorm, SwiGLU, Parameter Counting, Gradients and Micrograd


面试手撕代码(14):手撕 attention 家族

Attention from Scratch: Softmax, SDPA, Multi-Head, GQA, RoPE, KV Cache and Online Softmax


面试手撕代码(13):设计题与数据结构实现

Design Problems: LRU, LFU, Trie, O(1) Random Set and Fenwick Tree


面试手撕代码(12):动态规划(二)——背包、区间、状态机、树形

Dynamic Programming II: Knapsack, Interval, State Machine and Tree DP


面试手撕代码(11):动态规划(一)——线性与二维

Dynamic Programming I: Define the State, Then the Rest Follows


面试手撕代码(10):字符串

Strings: Palindromes, KMP, Parsing, Big Numbers and Custom Ordering


面试手撕代码(09):回溯

Backtracking: Choose, Recurse, Undo — and Prune


面试手撕代码(08):堆、Top-K、区间与贪心

Heaps, Top-K, Intervals and Greedy: Sort by the Right Key


面试手撕代码(07):二分——只有一个模板

Binary Search: One Template, First Position Where the Predicate Holds


面试手撕代码(06):图——BFS / DFS / 拓扑排序 / 并查集 / 最短路

Graphs: BFS, DFS, Topological Sort, Union-Find and Dijkstra


面试手撕代码(05):二叉树

Binary Trees: Three Questions Every Recursion Must Answer


面试手撕代码(04):链表

Linked Lists: Dummy Heads, Three Pointers and the Tortoise–Hare Proof


面试手撕代码(03):栈、单调栈与单调队列

Stacks, Monotonic Stacks and Monotonic Queues: Settle When You Pop


面试手撕代码(02):双指针与滑动窗口

Two Pointers and Sliding Windows: Monotonicity Turns O(n²) into O(n)


面试手撕代码(01):数组、哈希与前缀和

Arrays, Hashing and Prefix Sums: Trade Space for a Loop


面试手撕代码:从 LeetCode 中等题到 Transformer 组件(总纲)

Coding Interviews: From LeetCode Mediums to Hand-Written Transformer Components


Algorithms

面试手撕代码(20):系列总结与通关自测

Coding Interviews: Series Recap and Final Self-Test


面试手撕代码(13):设计题与数据结构实现

Design Problems: LRU, LFU, Trie, O(1) Random Set and Fenwick Tree


面试手撕代码(12):动态规划(二)——背包、区间、状态机、树形

Dynamic Programming II: Knapsack, Interval, State Machine and Tree DP


面试手撕代码(11):动态规划(一)——线性与二维

Dynamic Programming I: Define the State, Then the Rest Follows


面试手撕代码(10):字符串

Strings: Palindromes, KMP, Parsing, Big Numbers and Custom Ordering


面试手撕代码(09):回溯

Backtracking: Choose, Recurse, Undo — and Prune


面试手撕代码(08):堆、Top-K、区间与贪心

Heaps, Top-K, Intervals and Greedy: Sort by the Right Key


面试手撕代码(07):二分——只有一个模板

Binary Search: One Template, First Position Where the Predicate Holds


面试手撕代码(06):图——BFS / DFS / 拓扑排序 / 并查集 / 最短路

Graphs: BFS, DFS, Topological Sort, Union-Find and Dijkstra


面试手撕代码(05):二叉树

Binary Trees: Three Questions Every Recursion Must Answer


面试手撕代码(04):链表

Linked Lists: Dummy Heads, Three Pointers and the Tortoise–Hare Proof


面试手撕代码(03):栈、单调栈与单调队列

Stacks, Monotonic Stacks and Monotonic Queues: Settle When You Pop


面试手撕代码(02):双指针与滑动窗口

Two Pointers and Sliding Windows: Monotonicity Turns O(n²) into O(n)


面试手撕代码(01):数组、哈希与前缀和

Arrays, Hashing and Prefix Sums: Trade Space for a Loop


面试手撕代码:从 LeetCode 中等题到 Transformer 组件(总纲)

Coding Interviews: From LeetCode Mediums to Hand-Written Transformer Components


LeetCode

面试手撕代码(20):系列总结与通关自测

Coding Interviews: Series Recap and Final Self-Test


面试手撕代码(13):设计题与数据结构实现

Design Problems: LRU, LFU, Trie, O(1) Random Set and Fenwick Tree


面试手撕代码(12):动态规划(二)——背包、区间、状态机、树形

Dynamic Programming II: Knapsack, Interval, State Machine and Tree DP


面试手撕代码(11):动态规划(一)——线性与二维

Dynamic Programming I: Define the State, Then the Rest Follows


面试手撕代码(10):字符串

Strings: Palindromes, KMP, Parsing, Big Numbers and Custom Ordering


面试手撕代码(09):回溯

Backtracking: Choose, Recurse, Undo — and Prune


面试手撕代码(08):堆、Top-K、区间与贪心

Heaps, Top-K, Intervals and Greedy: Sort by the Right Key


面试手撕代码(07):二分——只有一个模板

Binary Search: One Template, First Position Where the Predicate Holds


面试手撕代码(06):图——BFS / DFS / 拓扑排序 / 并查集 / 最短路

Graphs: BFS, DFS, Topological Sort, Union-Find and Dijkstra


面试手撕代码(05):二叉树

Binary Trees: Three Questions Every Recursion Must Answer


面试手撕代码(04):链表

Linked Lists: Dummy Heads, Three Pointers and the Tortoise–Hare Proof


面试手撕代码(03):栈、单调栈与单调队列

Stacks, Monotonic Stacks and Monotonic Queues: Settle When You Pop


面试手撕代码(02):双指针与滑动窗口

Two Pointers and Sliding Windows: Monotonicity Turns O(n²) into O(n)


面试手撕代码(01):数组、哈希与前缀和

Arrays, Hashing and Prefix Sums: Trade Space for a Loop


面试手撕代码:从 LeetCode 中等题到 Transformer 组件(总纲)

Coding Interviews: From LeetCode Mediums to Hand-Written Transformer Components


Transformer

预训练(06):系列总结与通关自测

Pretraining: Series Recap and Final Self-Test


预训练(05):训练配方与稳定性:学习率、batch、调度与 loss spike

Pretraining Recipes and Training Stability: Learning Rate, Batch Size, Schedules and Loss Spikes


预训练(04):预训练数据工程:从 Common Crawl 到 15T token,去重、过滤与配比的账

Pretraining Data Engineering: From Common Crawl to 15T Tokens, the Arithmetic of Deduplication, Filtering and Mixing


预训练(03):Scaling law:从 Chinchilla 到"过训练",算力怎么分给参数与数据

Scaling Laws: From Chinchilla to Over-Training, Splitting Compute between Parameters and Data


预训练(02):分词与词表:BPE、词表大小与 token 效率

Tokenizers and Vocabulary: BPE, Vocabulary Size and Token Efficiency


预训练(01):一次预训练是怎么跑起来的:从两个网页文件到一个会续写英文的模型

Pretraining End to End: From Two Common Crawl Files to a Model That Writes English, on a Laptop


预训练:从 tokenizer 到训练配方(总纲)

Pretraining: Tokenizers, Scaling Laws, Data Pipelines and Training Recipes


Transformer 与 LLM:系列总结与通关自测

Transformers and LLMs: Series Recap and Final Self-Test


Transformer 与 LLM(13):多模态:vision encoder 的算量与 image token 的 KV 代价

Multimodal LLMs: The Cost of Vision Encoders and Image Tokens


Transformer 与 LLM(12):量化、投机解码与 LoRA

Quantization, Speculative Decoding and LoRA: Three Ways to Reshape the Computation


Transformer 与 LLM(11):浮点格式、数值稳定性与混合精度

Floating-Point Formats, Numerical Stability and Mixed Precision


Transformer 与 LLM(10):前向的算量与访存量

FLOPs, Bytes and Roofline: Prefill versus Decode


Transformer 与 LLM(09):MTP——改训练目标而不改主干的多 token 预测

Multi-Token Prediction: Denser Supervision from the Same Data, and a Free Speculative Draft


Transformer 与 LLM(08):MoE 的路由、激活参数量与通信形态

Mixture of Experts: Routing, Active Parameters and Communication Patterns


Transformer 与 LLM(07):位置编码与长上下文

Positional Encoding and Long Context: RoPE Wavelengths, Extrapolation and Cost


Transformer 与 LLM(06):Attention 变体与 KV cache

Attention Variants and the KV Cache: Deriving MHA, GQA, MQA and MLA


Transformer 与 LLM(05):Transformer 解剖与参数量

Transformer Anatomy and Parameter Count: From config.json to 8.03B


Transformer 与 LLM(04):手搓 GPT(下)——nanoGPT train.py 与训一个会续写的模型

Building GPT by Hand, Part 2: train.py Line by Line, Then Train One on Shakespeare


Transformer 与 LLM(03):手搓 GPT(上)——nanoGPT model.py 逐行解析

Building GPT by Hand, Part 1: Every Line of nanoGPT's model.py


Transformer 与 LLM(02):一个 token 的旅程——训练侧与推理侧

The Dynamic View: What Happens to a Token During Training and During Inference


Transformer 与 LLM(01):Transformer 长什么样——从一句话到下一个 token

The Static View: Every Box in a Decoder-only Transformer, Computed by Hand


Transformer 与 LLM:结构、实现与算量(总纲)

Transformers and LLMs: Architecture, Implementation and Arithmetic


面试手撕代码(15):手撕 Transformer block 与反向传播

Transformer Block and Backprop by Hand: LayerNorm, SwiGLU, Parameter Counting, Gradients and Micrograd


面试手撕代码(14):手撕 attention 家族

Attention from Scratch: Softmax, SDPA, Multi-Head, GQA, RoPE, KV Cache and Online Softmax


PyTorch

AI-Infra 开源贡献指南(05):系列总结与通关自测

Contributing to AI-Infra Open Source: Series Recap and Final Self-Test


AI-Infra 开源贡献指南(04):两个真实 PR 的完整走读——PyTorch 与 vLLM

Two Real Pull Requests, End to End: One in PyTorch, One in vLLM


AI-Infra 开源贡献指南(03):做出一个能被合入的改动

Landing a Mergeable Change: Diff, Tests, Benchmarks, PR, CI and Review


AI-Infra 开源贡献指南(02):找到切入点——从 issue、RFC 到性能回归

Finding Your Entry Point: Issues, RFCs, Roadmaps, CI Failures and Regressions


AI-Infra 开源贡献指南(01):读懂一个百万行的代码库

Reading a Million-Line Codebase: Maps, Entry Points, Symbols, Tests and History


AI-Infra 开源贡献指南(总纲)

A Guide to Contributing to AI-Infra Open Source Projects


大规模训练工程(08):长时训练的可观测与运维——从指标到 hang 排查

Observability and Operations for Long-Running Training


大规模训练工程(06):容错与弹性——故障率数学、straggler、SDC 与弹性训练

Fault Tolerance and Elasticity: Failure Math, Stragglers, SDC and Elastic Training


大规模训练工程(05):分布式 checkpoint——格式、异步保存与重分片恢复

Distributed Checkpoint: Format, Async Save and Resharding


读 Hugging Face 源码(01):transformers 模型侧——from_pretrained 怎么把三个文件变成 nn.Module,forward 怎么走到 loss

Inside transformers, Part 1: from_pretrained, the Decoder Stack, Attention Dispatch, KV Cache and the Loss


Transformer 与 LLM(04):手搓 GPT(下)——nanoGPT train.py 与训一个会续写的模型

Building GPT by Hand, Part 2: train.py Line by Line, Then Train One on Shakespeare


Transformer 与 LLM(03):手搓 GPT(上)——nanoGPT model.py 逐行解析

Building GPT by Hand, Part 1: Every Line of nanoGPT's model.py


PyTorch 深度实践(11):系列总结与通关自测

Deep Dive into PyTorch: Series Recap and Final Self-Test


PyTorch 深度实践(10):PyTorch 的工程体系——一次改动如何安全地到达用户

The Engineering System of PyTorch: How a Change Travels Safely from Commit to Production


PyTorch 深度实践(09):分布式 PyTorch

Distributed Training in PyTorch: Collectives, DDP, FSDP, TP, PP, CP and EP


PyTorch 深度实践(08):性能优化与调试

Performance Optimization and Debugging in PyTorch


PyTorch 深度实践(07):编译执行与图优化

Compilation and Graph Optimization in PyTorch


PyTorch 深度实践(06):C++ 扩展与自定义算子

C++ Extensions and Custom Operators in PyTorch


PyTorch 深度实践(05):Dispatcher 与算子系统

The Dispatcher and Operator System in PyTorch


PyTorch 深度实践(04):nn.Module 与训练系统

nn.Module and Training Systems in PyTorch


PyTorch 深度实践(03):自动求导与动态计算图

Autograd and Dynamic Computation Graphs in PyTorch


PyTorch 深度实践(02):Tensor 与内存布局

Tensor Abstraction and Memory Layout in PyTorch


PyTorch 深度实践(01):PyTorch 整体介绍

PyTorch Overall Introduction


PyTorch 深度实践:从 Tensor 到深度学习运行时(总纲)

Deep Dive into PyTorch, from Tensor to Deep Learning Runtime


算法工程师的工具箱(07):系列总结与通关自测

Tooling for AI Algorithm Engineers: Series Recap and Final Self-Test


算法工程师的工具箱(06):GPU 直觉与实验管理——两个上限、四块显存、能复现

GPU Intuition and Experiment Management: Two Ceilings, Four Memory Buckets, and Reproducibility


算法工程师的工具箱(05):Hugging Face 生态——六个库与一次 LoRA SFT 的组装

The Hugging Face Ecosystem: Six Libraries, a Six-Line LoRA SFT, and Why Reading the Source Is the Fastest Way to Learn


算法工程师的工具箱(04):PyTorch 使用层(下)——混合精度、显存的账与多卡启用

PyTorch in Use, Part 2: Mixed Precision, the Memory Ledger and Turning On Multi-GPU


算法工程师的工具箱(03):PyTorch 使用层(上)——五个对象与二十行训练循环

PyTorch in Use, Part 1: Five Objects and a Twenty-Line Training Loop


算法工程师的工具箱(02):数据科学三剑客——NumPy 的形状直觉、Pandas 的错误分析、Matplotlib 的曲线

The Scientific Python Stack: Shapes and Broadcasting in NumPy, Error Analysis in Pandas, Reading Curves in Matplotlib


算法工程师的工具箱:从一个想法到一次能跑的实验(总纲)

Tooling for AI Algorithm Engineers: NumPy, PyTorch, Hugging Face and the GPU in Your Head


面试手撕代码(17):手撕损失函数与训练算法

Losses and Training Algorithms by Hand: CE, KL, InfoNCE, DPO, PPO/GAE, GRPO, AdamW, Schedules and LoRA


面试手撕代码(15):手撕 Transformer block 与反向传播

Transformer Block and Backprop by Hand: LayerNorm, SwiGLU, Parameter Counting, Gradients and Micrograd


面试手撕代码(14):手撕 attention 家族

Attention from Scratch: Softmax, SDPA, Multi-Head, GQA, RoPE, KV Cache and Online Softmax


LoRA

扩散模型推理基础设施(07):serving 形态——请求、批、三段分离、附件、异步任务与成本

Serving Diffusion Models: Request Shapes, Batching, Disaggregation, Add-ons, Async Jobs and Cost


LoRA 专题(04):系列总结与通关自测

LoRA Series: Recap and Final Self-Test


LoRA 专题(03):工程:adapter 文件、合并、多 LoRA 服务与参考模型

LoRA 03: In Production — Adapter Files, Merging, Multi-LoRA Serving and the Free Reference Model


LoRA 专题(02):选参:r、target_modules、alpha、lr 与 QLoRA / DoRA / PiSSA 各让模型变成什么

LoRA 02: Choosing r, target_modules, alpha and lr — and What QLoRA, DoRA and PiSSA Each Change


LoRA 专题(01):低秩假设:为什么两个瘦矩阵够用,以及它省了哪几本账

LoRA 01: The Low-Rank Hypothesis — Why Two Thin Matrices Suffice, and Exactly What They Save


LoRA 专题:SFT 的默认微调方式,从低秩假设到多租户服务(总纲)

LoRA, the Default Way to Fine-Tune: From the Low-Rank Hypothesis to Multi-Tenant Serving — Series Overview


读 Hugging Face 源码(05):系列总结与通关自测

Reading the Hugging Face Source: Series Recap and Final Self-Test


读 Hugging Face 源码(04):peft 与 trl——LoRA 怎么挂上去,SFT / DPO / GRPO 的 loss 各在哪一行

Inside peft and trl: get_peft_model, the LoRA Linear, SFT Labels and Packing, dpo_loss and GRPO Advantages


读 Hugging Face 源码:从 from_pretrained 到 GRPO 的 loss(总纲)

Reading the Hugging Face Source: transformers, tokenizers, datasets, peft and trl, One Call Chain at a Time


面试手撕代码(17):手撕损失函数与训练算法

Losses and Training Algorithms by Hand: CE, KL, InfoNCE, DPO, PPO/GAE, GRPO, AdamW, Schedules and LoRA


Machine Learning

算法工程师的实验方法论:用有限的算力得出可信的结论

Experimental Methodology for AI Algorithm Engineers: Credible Conclusions on a Finite Compute Budget


LLM 时代的经典机器学习(11):系列总结与通关自测

Classical Machine Learning in the LLM Era: Series Recap and Final Self-Test


LLM 时代的经典机器学习(10):评估——从混淆矩阵到 judge 的一致性

Evaluation: Confusion Matrices, Thresholds, Calibration, Paired Tests, and Every LLM Pitfall They Predict


LLM 时代的经典机器学习(09):去重——MinHash 与 LSH 的概率

Deduplication: Jaccard, MinHash as an Unbiased Estimator, and the LSH S-Curve


LLM 时代的经典机器学习(08):降维——PCA、SVD、t-SNE 与 embedding 的各向异性

Dimensionality Reduction: PCA, SVD, t-SNE / UMAP, and the Anisotropy of Embedding Spaces


LLM 时代的经典机器学习(07):聚类——K-Means、DBSCAN 与「这批语料里有什么」

Clustering: K-Means, DBSCAN, Hierarchical Clustering, and What a Corpus Looks Like in Embedding Space


LLM 时代的经典机器学习(06):集成——随机森林、梯度提升与数据质量分类器的算力账

Ensembles: Random Forests, Gradient Boosting, and Why Data Filters Use Small Models


LLM 时代的经典机器学习(05):SVM 与核方法——最大间隔、核技巧与 attention 的远亲

Support Vector Machines and Kernels: Maximum Margin, the Kernel Trick, and Why Attention Is a Kernel Smoother


LLM 时代的经典机器学习(04):三个基础分类器——朴素贝叶斯、KNN 与决策树

Three Basic Classifiers: Naive Bayes Computes Probabilities, KNN Finds Neighbours, Decision Trees Ask Questions


LLM 时代的经典机器学习(03):逻辑回归与奖励模型——每个分类头的原型

Logistic Regression and Reward Models: The Skeleton of Every Classification Head


LLM 时代的经典机器学习(02):线性回归——最小二乘、梯度下降与 Ridge / Lasso

Linear Regression: Least Squares, Gradient Descent, and Why Ridge Is Weight Decay


LLM 时代的经典机器学习(01):什么是学习——划分、泛化、过拟合与偏差-方差

What Is Learning: Splits, Generalization, Overfitting and the Bias-Variance Decomposition


LLM 时代还要学经典机器学习吗:只讲它在哪里重现(总纲)

Classical Machine Learning in the LLM Era: What Survives and Where It Reappears


面试手撕代码(18):手撕经典 ML 与评测指标

Classical ML and Metrics by Hand: k-means, Logistic Regression, KNN, PCA, AUC, NDCG, conv2d, NMS


AI-Infra

AI-Infra 开源贡献指南(05):系列总结与通关自测

Contributing to AI-Infra Open Source: Series Recap and Final Self-Test


AI-Infra 开源贡献指南(04):两个真实 PR 的完整走读——PyTorch 与 vLLM

Two Real Pull Requests, End to End: One in PyTorch, One in vLLM


AI-Infra 开源贡献指南(03):做出一个能被合入的改动

Landing a Mergeable Change: Diff, Tests, Benchmarks, PR, CI and Review


AI-Infra 开源贡献指南(02):找到切入点——从 issue、RFC 到性能回归

Finding Your Entry Point: Issues, RFCs, Roadmaps, CI Failures and Regressions


AI-Infra 开源贡献指南(01):读懂一个百万行的代码库

Reading a Million-Line Codebase: Maps, Entry Points, Symbols, Tests and History


AI-Infra 开源贡献指南(总纲)

A Guide to Contributing to AI-Infra Open Source Projects


AI 平台工程(09):系列总结与通关自测

AI Platform Engineering: Series Recap and Final Self-Test


AI 平台工程(08):可观测、成本与 FinOps

Observability, Cost and FinOps: from DCGM to the Token Bill


AI 平台工程(07):模型网关与多租户——路由、配额与灰度

The Model Gateway: KV-Aware Routing, Multi-Tenancy, Quotas and Canaries


AI 平台工程(06):Serving 平台——从 InferenceService 到 llm-d

Serving Platforms: KServe, Triton, Ray Serve, LeaderWorkerSet and llm-d


AI 平台工程(05):网络与存储——RDMA 进容器、并行文件系统与 checkpoint I/O

Networking and Storage: RDMA in Containers, Parallel File Systems and Checkpoint I/O


AI 平台工程(04):GPU 共享与切分——MIG、时间片、MPS 与 HAMi

Sharing and Partitioning GPUs: MIG, Time-Slicing, MPS and HAMi


AI 平台工程(03):AI 任务调度——gang scheduling、队列与拓扑感知

Scheduling AI Jobs: Gang Scheduling, Queues, Quotas and Topology Awareness


AI 平台工程(02):容器里的 GPU——驱动、CUDA、device plugin 与镜像

GPUs in Containers: Driver, CUDA, Device Plugin, DRA and Images


AI 平台工程(01):引擎的需求清单与平台的整体架构

What Engines Demand from the Platform, and the Platform's Two Layers


AI 平台工程:资源层与交付层(总纲)

AI Platform Engineering: the Resource Layer and the Delivery Layer


扩散模型推理基础设施(10):系列总结与通关自测

Diffusion Model Inference Infrastructure: Series Recap and Final Self-Test


扩散模型推理基础设施(09):配置、评测与排障——从一张卡的推导到一条伪影的排查

Configuration, Evaluation and Troubleshooting for Diffusion Inference (with Series Summary)


扩散模型推理基础设施(08):三个引擎的对照导读——同一张图的请求在 SGLang Diffusion、vLLM-Omni 与 xDiT 里各走过什么

Three Engines Compared: One Request Through SGLang Diffusion, vLLM-Omni and xDiT


扩散模型推理基础设施(07):serving 形态——请求、批、三段分离、附件、异步任务与成本

Serving Diffusion Models: Request Shapes, Batching, Disaggregation, Add-ons, Async Jobs and Cost


扩散模型推理基础设施(06):少步与自回归——把步数变成系统参数

Few-Step and Autoregressive Generation: When the Step Count Becomes a System Parameter


扩散模型推理基础设施(05):多卡并行——序列并行、CFG 并行与 PipeFusion,为什么不是张量并行

Multi-GPU Diffusion Inference: Sequence Parallelism, CFG Parallelism and PipeFusion


扩散模型推理基础设施(04):视频——长序列 attention 的账与稀疏化

Video Diffusion: The Long-Sequence Attention Bill and How Sparsity Pays It


扩散模型推理基础设施(03):跨步冗余——TeaCache、First-Block Cache 一族的缓存与跳步

Temporal Redundancy Across Denoising Steps: TeaCache, First-Block Cache and Friends


扩散模型推理基础设施(02):单卡执行——attention 后端、编译、FP8 / INT4 与 offload

Single-GPU Execution: Attention Backends, Compilation, FP8 / INT4 and Offloading


扩散模型推理基础设施(01):负载画像——一次生成在 GPU 上发生什么

Workload Anatomy: FLOPs, Bytes and Seconds of One Diffusion Generation


扩散模型推理基础设施:图像与视频生成的 serving(总纲)

Diffusion Model Inference Infrastructure: Serving Image and Video Generation


RL 后训练基础设施(09):系列总结与通关自测

RL Post-Training Infrastructure: Series Recap and Final Self-Test


RL 后训练基础设施(08):配置、可观测与排障——从一张卡的比例到一条 hang 的排查

Configuration, Observability and Troubleshooting for RL Post-Training Systems


RL 后训练基础设施(07):verl 源码导读——从一个 GRPO 配置追到每个 worker

Reading verl: From One GRPO Config to Every Worker


RL 后训练基础设施(06):Agentic rollout——多轮、工具、沙箱与环境服务

Agentic Rollout: Multi-Turn Trajectories, Tools, Sandboxes and Environment Services


RL 后训练基础设施(05):异步与 off-policy——把同步的墙拆掉之后要补什么

Asynchrony and Off-Policy: What You Owe After Tearing Down the Synchronization Wall


RL 后训练基础设施(04):权重同步——从训练分片到推理分片

Weight Synchronization: From Training Shards to Inference Shards


RL 后训练基础设施(03):共置——训练器与推理引擎在同一组 GPU 上共存

Colocation: Handing GPU Memory Back and Forth Between Trainer and Rollout Engine


RL 后训练基础设施(02):系统形态——共置、分离与异步

RL System Topologies: Colocated, Disaggregated and Asynchronous


RL 后训练基础设施(01):负载画像——一步 RL 里发生什么

Anatomy of an RL Step: Rollout, Reward and Train


RL 后训练基础设施:rollout 与训练如何共享一组 GPU(总纲)

RL Post-Training Infrastructure: How Rollout and Training Share the Same GPUs


大模型推理系统揭秘(16):系列总结与通关自测

Deep Dive into vLLM: Series Recap and Final Self-Test


大模型推理系统揭秘(15):vLLM 与 SGLang:同一个请求穿过两套 Serving 系统

vLLM vs SGLang: One Request Through Two LLM Serving Systems


大模型推理系统揭秘(14):回到源码:一次请求在 vLLM 内部的真实旅程


大模型推理系统揭秘(13):Serving Infra 的下一站:从模型执行器到分布式智能操作系统


大模型推理系统揭秘(12):硬件解耦:如何不让芯片差异污染 Serving 核心?


大模型推理系统揭秘(11):请求形态的扩展:multi-LoRA 与多模态


大模型推理系统揭秘(10):模型适配:如何跟上变化极快的模型世界?


大模型推理系统揭秘(09):PD 分离:从资源混部走向计算解耦


大模型推理系统揭秘(08):Multi-GPU:一张卡不够时如何扩展?


大模型推理系统揭秘(07):解码的扩展:采样、投机解码与结构化输出


大模型推理系统揭秘(06):GPU 执行:如何让每个 Token 算得更快?


大模型推理系统揭秘(05):KV Cache:LLM Serving 的第一号内存问题


大模型推理系统揭秘(04):Scheduler:GPU 这一轮到底给谁用?


大模型推理系统揭秘(03):鸟瞰 vLLM:一个请求如何穿过整个推理系统?


大模型推理系统揭秘(02):如何衡量一个 LLM Serving 系统?


大模型推理系统揭秘(01):为什么 LLM Serving 比传统 DL 推理难?


大模型推理系统揭秘:从 vLLM 看 LLM Serving Infra 核心技术(总纲)


大规模训练工程(09):系列总结与通关自测

Large-Scale Training Engineering: Series Recap and Final Self-Test


大规模训练工程(08):长时训练的可观测与运维——从指标到 hang 排查

Observability and Operations for Long-Running Training


大规模训练工程(07):训练稳定性与数据管线——loss spike、梯度范数、数据混合与流式加载

Training Stability and the Data Pipeline


大规模训练工程(06):容错与弹性——故障率数学、straggler、SDC 与弹性训练

Fault Tolerance and Elasticity: Failure Math, Stragglers, SDC and Elastic Training


大规模训练工程(05):分布式 checkpoint——格式、异步保存与重分片恢复

Distributed Checkpoint: Format, Async Save and Resharding


大规模训练工程(04):千卡配置实战——并行搭配、micro-batch、激活重计算与 MFU 调优

Configuring a Thousand-GPU Job: Parallelism, Micro-batch, Recompute and MFU


大规模训练工程(03):三个框架——Megatron-LM、DeepSpeed 与 torchtitan 的架构对比与源码导读

Megatron-LM, DeepSpeed and torchtitan: Architecture and Source Guide


大规模训练工程(02):并行策略全景——每种并行切的是哪种状态

A Map of Parallelism: Which State Does Each Strategy Shard


大规模训练工程(01):训练任务的状态解剖——显存账与 MFU

Anatomy of Training State: Memory Accounting and MFU


大规模训练工程:从并行策略到容错恢复(总纲)

Large-Scale Training Engineering, from Parallelism to Fault Tolerance


通信与互联(09):系列总结与通关自测

Communication and Interconnect: Series Recap and Final Self-Test


通信与互联(08):MoE 的通信——all-to-all、DeepEP 与 GPU 发起的通信

Communication for MoE: All-to-All, DeepEP and GPU-Initiated Networking


通信与互联(07):推理侧的通信——custom all-reduce 与 KV 传输

Communication on the Inference Side: Custom All-Reduce and KV Cache Transfer


通信与互联(06):nccl-tests、调优与排障——从带宽曲线到 hang

nccl-tests, Tuning and Debugging Hangs: From Bandwidth Curves to Flight Recorder


通信与互联(05):PyTorch 的通信栈——ProcessGroupNCCL、stream 语义与计算通信重叠

PyTorch's Communication Stack: ProcessGroupNCCL, Stream Semantics, and Compute-Communication Overlap


通信与互联(04):NCCL 架构——拓扑探测、channel、算法与协议

NCCL Architecture: Topology Detection, Channels, Algorithms and Protocols


通信与互联(03):RDMA 与 GPUDirect——绕过 CPU 和主机内存的数据通路

RDMA and GPUDirect: Bypassing the CPU and Host Memory


通信与互联(02):硬件互联——PCIe、NVLink、NVSwitch 与网络拓扑

Hardware Interconnect: PCIe, NVLink, NVSwitch and Network Topology


通信与互联(01):集合通信原语与代价模型——α-β 模型与 ring all-reduce

Collective Communication Primitives and the Alpha-Beta Cost Model: Deriving Ring All-Reduce


通信与互联:从 NCCL 到 RDMA(总纲)

Communication and Interconnect for AI-Infra, from NCCL to RDMA


GPU Kernel 工程(11):系列总结与通关自测

GPU Kernel Engineering: Series Recap and Final Self-Test


GPU Kernel 工程(10):剖析、测试与贡献——把 kernel 做成产品

Profiling, Testing and Contributing: Turning a Kernel into a Product


GPU Kernel 工程(09):量化与融合 kernel——推理系统的其余部分

Quantized and Fused Kernels: The Rest of the Inference Stack


GPU Kernel 工程(08):Attention Kernel——FlashAttention 与 PagedAttention

Attention Kernels: FlashAttention and PagedAttention from Derivation to Code


GPU Kernel 工程(07):Triton——块级编程与编译器的边界

Triton: Block-Level Programming and Where the Compiler Stops


GPU Kernel 工程(06):Tensor Core、CUTLASS 与 CuTe

Tensor Cores, CUTLASS and CuTe: Programming the Matrix Units


GPU Kernel 工程(05):GEMM——从 naive 到分块

GEMM from Naive to Tiled: Reaching the Compute Ceiling on CUDA Cores


GPU Kernel 工程(04):共享内存与 reduction——softmax、LayerNorm 与 online softmax

Shared Memory and Reductions: Softmax, LayerNorm and Online Softmax


GPU Kernel 工程(03):访存合并与 elementwise kernel

Memory Coalescing and Elementwise Kernels: Hitting the Bandwidth Ceiling


GPU Kernel 工程(02):CUDA 编程模型与第一个 kernel

The CUDA Programming Model and Your First Kernel, Measured


GPU Kernel 工程(01):GPU 为什么这样设计——硬件结构与 Roofline

Why GPUs Look the Way They Do: Architecture and the Roofline Model


GPU Kernel 工程:从 CUDA 执行模型到 FlashAttention(总纲)

GPU Kernel Engineering, from the CUDA Execution Model to FlashAttention


Transformer 与 LLM:系列总结与通关自测

Transformers and LLMs: Series Recap and Final Self-Test


Transformer 与 LLM(13):多模态:vision encoder 的算量与 image token 的 KV 代价

Multimodal LLMs: The Cost of Vision Encoders and Image Tokens


Transformer 与 LLM(12):量化、投机解码与 LoRA

Quantization, Speculative Decoding and LoRA: Three Ways to Reshape the Computation


Transformer 与 LLM(11):浮点格式、数值稳定性与混合精度

Floating-Point Formats, Numerical Stability and Mixed Precision


Transformer 与 LLM(10):前向的算量与访存量

FLOPs, Bytes and Roofline: Prefill versus Decode


Transformer 与 LLM(09):MTP——改训练目标而不改主干的多 token 预测

Multi-Token Prediction: Denser Supervision from the Same Data, and a Free Speculative Draft


Transformer 与 LLM(08):MoE 的路由、激活参数量与通信形态

Mixture of Experts: Routing, Active Parameters and Communication Patterns


Transformer 与 LLM(07):位置编码与长上下文

Positional Encoding and Long Context: RoPE Wavelengths, Extrapolation and Cost


Transformer 与 LLM(06):Attention 变体与 KV cache

Attention Variants and the KV Cache: Deriving MHA, GQA, MQA and MLA


Transformer 与 LLM(05):Transformer 解剖与参数量

Transformer Anatomy and Parameter Count: From config.json to 8.03B


Transformer 与 LLM(04):手搓 GPT(下)——nanoGPT train.py 与训一个会续写的模型

Building GPT by Hand, Part 2: train.py Line by Line, Then Train One on Shakespeare


Transformer 与 LLM(03):手搓 GPT(上)——nanoGPT model.py 逐行解析

Building GPT by Hand, Part 1: Every Line of nanoGPT's model.py


Transformer 与 LLM(02):一个 token 的旅程——训练侧与推理侧

The Dynamic View: What Happens to a Token During Training and During Inference


Transformer 与 LLM(01):Transformer 长什么样——从一句话到下一个 token

The Static View: Every Box in a Decoder-only Transformer, Computed by Hand


Transformer 与 LLM:结构、实现与算量(总纲)

Transformers and LLMs: Architecture, Implementation and Arithmetic


PyTorch 深度实践(11):系列总结与通关自测

Deep Dive into PyTorch: Series Recap and Final Self-Test


PyTorch 深度实践(10):PyTorch 的工程体系——一次改动如何安全地到达用户

The Engineering System of PyTorch: How a Change Travels Safely from Commit to Production


PyTorch 深度实践(09):分布式 PyTorch

Distributed Training in PyTorch: Collectives, DDP, FSDP, TP, PP, CP and EP


PyTorch 深度实践(08):性能优化与调试

Performance Optimization and Debugging in PyTorch


PyTorch 深度实践(07):编译执行与图优化

Compilation and Graph Optimization in PyTorch


PyTorch 深度实践(06):C++ 扩展与自定义算子

C++ Extensions and Custom Operators in PyTorch


PyTorch 深度实践(05):Dispatcher 与算子系统

The Dispatcher and Operator System in PyTorch


PyTorch 深度实践(04):nn.Module 与训练系统

nn.Module and Training Systems in PyTorch


PyTorch 深度实践(03):自动求导与动态计算图

Autograd and Dynamic Computation Graphs in PyTorch


PyTorch 深度实践(02):Tensor 与内存布局

Tensor Abstraction and Memory Layout in PyTorch


PyTorch 深度实践(01):PyTorch 整体介绍

PyTorch Overall Introduction


PyTorch 深度实践:从 Tensor 到深度学习运行时(总纲)

Deep Dive into PyTorch, from Tensor to Deep Learning Runtime


C++ 在 AI-Infra(09):系列总结与通关自测

C++ for AI-Infra: Series Recap and Final Self-Test


C++ 在 AI-Infra(08):构建、调试与测试工具链

Build, Debug and Test Toolchain


C++ 在 AI-Infra(07):与 Python 之间——pybind11、Python C API 与 ABI

pybind11, the Python C API and ABI


C++ 在 AI-Infra(06):并发、内存模型、TLS 与守卫

Concurrency, Memory Model, TLS and Guards


C++ 在 AI-Infra(05):宏、静态注册与代码生成

Macros, Static Registration and Code Generation


C++ 在 AI-Infra(04):多态与类型擦除——运行时如何选择实现

Polymorphism and Type Erasure


C++ 在 AI-Infra(03):模板与泛型编程

Templates and Generic Programming


C++ 在 AI-Infra(02):值、引用与所有权——对象模型与 RAII

Value Semantics, Ownership and RAII


C++ 在 AI-Infra(01 下):工程布局——命名空间、库的分层与 CMake

Project Layout: Namespaces, Library Layering and CMake


C++ 在 AI-Infra(01 上):编译模型——从一个 .cpp 到可加载的 .so

The Compilation Model: From a .cpp to a Loadable .so


C++ 在 AI-Infra:从对象模型到算子扩展(总纲)

C++ for AI-Infra, from the Object Model to Operator Extensions


Python 在 AI-Infra(08):系列总结与通关自测

Python for AI-Infra: Series Recap and Final Self-Test


Python 在 AI-Infra:从语言机制到生产交付(总纲)

Python for AI-Infra, from Language Mechanisms to Production Delivery (Overview)


AI-Infra 工程师学习地图:从后端工程师到基础设施贡献者

A Learning Roadmap for AI Infrastructure Engineers


AI 全栈学习地图:造模型、跑模型、用模型的三张图

One System, Three Roles — an Overview of the Three AI Learning Roadmaps


面试手撕代码(19):Infra 岗手撕——并发与系统

Infra Interviews: Thread-Safe LRU, Bounded Queues, Thread Pools, Memory Pools, Blocked GEMM, Ring Allreduce, Paged KV and Token Buckets


LLM

Prompt 与上下文工程:模型这一步该看到什么(总纲)

Prompt and Context Engineering: What the Model Should See at This Step


模型作为组件(07):系列总结与通关自测

The Model as a Component: Series Recap and Final Self-Test


模型作为组件(06):客户端工程——重试、超时、幂等、限流与流式解析

Client Engineering for LLM APIs: Retries, Timeouts, Idempotency, Rate Limits and Stream Parsing


模型作为组件(05):选型——榜单的失真、自己的评测集与供应商的弃用周期

Model Selection Beyond Leaderboards: Your Own Evals, Open vs Closed, Routing and Deprecation Cycles


模型作为组件(04):成本与延迟的账——一次调用花多少钱、慢在哪一段

The Token Cost and Latency Ledger for LLM Applications


模型作为组件(03):API 契约(二)——推理模型:thinking、effort 与跨轮的推理状态

The LLM API Contract, Part 2: Reasoning Models — Thinking, Effort and Cross-turn State


模型作为组件(02):API 契约(一)——消息、工具调用、结构化输出与流式,四家 API 的共同骨架

The LLM API Contract, Part 1: Messages, Tool Calls, Structured Output and Streaming


模型作为组件(01):失效模式——非确定性、幻觉、上下文与越界,把「模型会出错」拆成七条可检测的性质

Failure Modes of LLMs as Components: Nondeterminism, Hallucination, Context and Overreach


模型作为组件:契约、失效模式与选型(总纲)

The Model as a Component: Contract, Failure Modes and Selection


LoRA 专题(04):系列总结与通关自测

LoRA Series: Recap and Final Self-Test


LoRA 专题(03):工程:adapter 文件、合并、多 LoRA 服务与参考模型

LoRA 03: In Production — Adapter Files, Merging, Multi-LoRA Serving and the Free Reference Model


LoRA 专题(02):选参:r、target_modules、alpha、lr 与 QLoRA / DoRA / PiSSA 各让模型变成什么

LoRA 02: Choosing r, target_modules, alpha and lr — and What QLoRA, DoRA and PiSSA Each Change


LoRA 专题(01):低秩假设:为什么两个瘦矩阵够用,以及它省了哪几本账

LoRA 01: The Low-Rank Hypothesis — Why Two Thin Matrices Suffice, and Exactly What They Save


LoRA 专题:SFT 的默认微调方式,从低秩假设到多租户服务(总纲)

LoRA, the Default Way to Fine-Tune: From the Low-Rank Hypothesis to Multi-Tenant Serving — Series Overview


读 Hugging Face 源码(03):tokenizers 与 datasets——从 messages 到 input_ids,从 Arrow 文件到 collate_fn

Inside tokenizers and datasets: the Rust Pipeline, Chat Templates, Arrow Tables, map and Fingerprints


读 Hugging Face 源码(02):generate——一次采样的完整调用链

Inside transformers, Part 2: generate, GenerationConfig, LogitsProcessors, StoppingCriteria and the Decode Loop


读 Hugging Face 源码(01):transformers 模型侧——from_pretrained 怎么把三个文件变成 nn.Module,forward 怎么走到 loss

Inside transformers, Part 1: from_pretrained, the Decoder Stack, Attention Dispatch, KV Cache and the Loss


多模态(10):系列总结与通关自测

Multimodal Models: Series Recap and Final Self-Test


多模态:从视觉编码器到扩散模型(总纲)

Multimodal Models: From Vision Encoders to Diffusion


高效推理与压缩(07):系列总结与通关自测

Efficient Inference and Model Compression: Series Recap and Final Self-Test


高效推理与压缩(06):剪枝、深度缩放与小模型配方

Pruning, Depth Scaling and How Small Models Are Made


高效推理与压缩(05):KV cache 压缩:量化、驱逐与稀疏 attention

KV Cache Compression: Quantization, Eviction and Sparse Attention


高效推理与压缩(04):量化感知训练、低比特与量化模型的评测

Quantization-Aware Training, Extreme Low-Bit and How to Evaluate a Quantized Model


高效推理与压缩(03):训练后量化:误差模型、GPTQ、AWQ 与旋转

Post-Training Quantization: Error Models, GPTQ, AWQ and Rotation


高效推理与压缩(02):投机解码:草稿、接受率与树

Speculative Decoding: Drafters, Acceptance Rates and Draft Trees


高效推理与压缩(01):解码策略、采样与约束生成

Decoding Strategies: Sampling, Truncation, Penalties and Constrained Generation


高效推理与压缩(算法侧):解码、投机、量化与 KV(总纲)

Efficient Inference and Model Compression: The Algorithm Side


算法工程师的实验方法论:用有限的算力得出可信的结论

Experimental Methodology for AI Algorithm Engineers: Credible Conclusions on a Finite Compute Budget


后训练(09):系列总结与通关自测

Post-Training: Series Recap and Final Self-Test


后训练(08):评测:benchmark、LLM-as-judge、Arena 与污染

Evaluating LLMs: Benchmarks, LLM-as-Judge, Arenas and Contamination


后训练(07):蒸馏:logits 级、序列级与 on-policy

Knowledge Distillation for LLMs: Logit-Level, Sequence-Level and On-Policy


后训练(06):Agent 与工具调用的 RL:多轮环境、轨迹数据与延后的奖励

Agentic RL: Multi-Turn Environments, Trajectory Data and Delayed Rewards


后训练(05):推理模型与可验证奖励:R1 的配方、PRM 与 test-time compute

Reasoning Models and Verifiable Rewards: The R1 Recipe, Process Reward Models and Test-Time Compute


后训练(04):离线 RL:从 RLHF 目标推出 DPO 及其变体

Offline Preference Optimization: Deriving DPO from the RLHF Objective, and Its Family


后训练(03):在线 RL:PPO、GRPO 与 RLHF 三件套

Online RL for LLMs: PPO, GRPO and the Policy–Reward–Reference Trio


后训练(02):偏好数据与奖励模型:Bradley-Terry、pairwise loss 与 reward hacking

Preference Data and Reward Models: Bradley-Terry, Pairwise Loss and Reward Hacking


后训练(01):SFT:指令数据、chat template、loss mask 与参数高效微调

Supervised Fine-Tuning: Instruction Data, Chat Templates, Loss Masking and Parameter-Efficient Fine-Tuning


后训练:从 SFT 到可验证奖励(总纲)

Post-Training: From Supervised Fine-Tuning to Verifiable Rewards


预训练(06):系列总结与通关自测

Pretraining: Series Recap and Final Self-Test


预训练(05):训练配方与稳定性:学习率、batch、调度与 loss spike

Pretraining Recipes and Training Stability: Learning Rate, Batch Size, Schedules and Loss Spikes


预训练(04):预训练数据工程:从 Common Crawl 到 15T token,去重、过滤与配比的账

Pretraining Data Engineering: From Common Crawl to 15T Tokens, the Arithmetic of Deduplication, Filtering and Mixing


预训练(03):Scaling law:从 Chinchilla 到"过训练",算力怎么分给参数与数据

Scaling Laws: From Chinchilla to Over-Training, Splitting Compute between Parameters and Data


预训练(02):分词与词表:BPE、词表大小与 token 效率

Tokenizers and Vocabulary: BPE, Vocabulary Size and Token Efficiency


预训练(01):一次预训练是怎么跑起来的:从两个网页文件到一个会续写英文的模型

Pretraining End to End: From Two Common Crawl Files to a Model That Writes English, on a Laptop


预训练:从 tokenizer 到训练配方(总纲)

Pretraining: Tokenizers, Scaling Laws, Data Pipelines and Training Recipes


Transformer 与 LLM:系列总结与通关自测

Transformers and LLMs: Series Recap and Final Self-Test


Transformer 与 LLM(13):多模态:vision encoder 的算量与 image token 的 KV 代价

Multimodal LLMs: The Cost of Vision Encoders and Image Tokens


Transformer 与 LLM(12):量化、投机解码与 LoRA

Quantization, Speculative Decoding and LoRA: Three Ways to Reshape the Computation


Transformer 与 LLM(11):浮点格式、数值稳定性与混合精度

Floating-Point Formats, Numerical Stability and Mixed Precision


Transformer 与 LLM(10):前向的算量与访存量

FLOPs, Bytes and Roofline: Prefill versus Decode


Transformer 与 LLM(09):MTP——改训练目标而不改主干的多 token 预测

Multi-Token Prediction: Denser Supervision from the Same Data, and a Free Speculative Draft


Transformer 与 LLM(08):MoE 的路由、激活参数量与通信形态

Mixture of Experts: Routing, Active Parameters and Communication Patterns


Transformer 与 LLM(07):位置编码与长上下文

Positional Encoding and Long Context: RoPE Wavelengths, Extrapolation and Cost


Transformer 与 LLM(06):Attention 变体与 KV cache

Attention Variants and the KV Cache: Deriving MHA, GQA, MQA and MLA


Transformer 与 LLM(05):Transformer 解剖与参数量

Transformer Anatomy and Parameter Count: From config.json to 8.03B


Transformer 与 LLM(04):手搓 GPT(下)——nanoGPT train.py 与训一个会续写的模型

Building GPT by Hand, Part 2: train.py Line by Line, Then Train One on Shakespeare


Transformer 与 LLM(03):手搓 GPT(上)——nanoGPT model.py 逐行解析

Building GPT by Hand, Part 1: Every Line of nanoGPT's model.py


Transformer 与 LLM(02):一个 token 的旅程——训练侧与推理侧

The Dynamic View: What Happens to a Token During Training and During Inference


Transformer 与 LLM(01):Transformer 长什么样——从一句话到下一个 token

The Static View: Every Box in a Decoder-only Transformer, Computed by Hand


Transformer 与 LLM:结构、实现与算量(总纲)

Transformers and LLMs: Architecture, Implementation and Arithmetic


深度学习基础(07):系列总结与通关自测

Deep Learning Foundations: Series Recap and Final Self-Test


深度学习基础(06):RNN——从 LSTM 到 attention 的诞生

RNN: Backpropagation Through Time, the LSTM Gate as a Residual Path, and How Attention Was Born from the seq2seq Bottleneck


深度学习基础(05):CNN——从 LeNet 到 ResNet,再到 ViT

CNN: Convolution as a Constrained Linear Layer, the ResNet Legacy and How ViT Turns Images into Tokens


深度学习基础(04):正则化与泛化——为什么参数比样本多却不过拟合

Regularization and Generalization: Double Descent, Implicit Bias, Dropout, Weight Decay and When Overfitting Comes Back


深度学习基础(03):优化器——从 SGD 到 AdamW 与学习率调度

Optimizers: From SGD to AdamW, Warmup, Clipping and the Batch-Learning-Rate Scaling Rule


深度学习基础(02):训练为什么不稳定——初始化、归一化与残差

Why Deep Training Is Unstable: Variance Propagation, Initialization, Normalization and Residual Connections


深度学习基础(01):反向传播——手推一个两层网络

Backpropagation by Hand: Shapes, the 2x Rule and Why Activations Must Be Saved


深度学习基础:从反向传播到残差(总纲)

Deep Learning Foundations: From Backpropagation to Residual Connections


LLM 时代的经典机器学习(11):系列总结与通关自测

Classical Machine Learning in the LLM Era: Series Recap and Final Self-Test


LLM 时代的经典机器学习(10):评估——从混淆矩阵到 judge 的一致性

Evaluation: Confusion Matrices, Thresholds, Calibration, Paired Tests, and Every LLM Pitfall They Predict


LLM 时代的经典机器学习(09):去重——MinHash 与 LSH 的概率

Deduplication: Jaccard, MinHash as an Unbiased Estimator, and the LSH S-Curve


LLM 时代的经典机器学习(08):降维——PCA、SVD、t-SNE 与 embedding 的各向异性

Dimensionality Reduction: PCA, SVD, t-SNE / UMAP, and the Anisotropy of Embedding Spaces


LLM 时代的经典机器学习(07):聚类——K-Means、DBSCAN 与「这批语料里有什么」

Clustering: K-Means, DBSCAN, Hierarchical Clustering, and What a Corpus Looks Like in Embedding Space


LLM 时代的经典机器学习(06):集成——随机森林、梯度提升与数据质量分类器的算力账

Ensembles: Random Forests, Gradient Boosting, and Why Data Filters Use Small Models


LLM 时代的经典机器学习(05):SVM 与核方法——最大间隔、核技巧与 attention 的远亲

Support Vector Machines and Kernels: Maximum Margin, the Kernel Trick, and Why Attention Is a Kernel Smoother


LLM 时代的经典机器学习(04):三个基础分类器——朴素贝叶斯、KNN 与决策树

Three Basic Classifiers: Naive Bayes Computes Probabilities, KNN Finds Neighbours, Decision Trees Ask Questions


LLM 时代的经典机器学习(03):逻辑回归与奖励模型——每个分类头的原型

Logistic Regression and Reward Models: The Skeleton of Every Classification Head


LLM 时代的经典机器学习(02):线性回归——最小二乘、梯度下降与 Ridge / Lasso

Linear Regression: Least Squares, Gradient Descent, and Why Ridge Is Weight Decay


LLM 时代的经典机器学习(01):什么是学习——划分、泛化、过拟合与偏差-方差

What Is Learning: Splits, Generalization, Overfitting and the Bias-Variance Decomposition


LLM 时代还要学经典机器学习吗:只讲它在哪里重现(总纲)

Classical Machine Learning in the LLM Era: What Survives and Where It Reappears


算法工程师的工具箱(07):系列总结与通关自测

Tooling for AI Algorithm Engineers: Series Recap and Final Self-Test


算法工程师的工具箱(06):GPU 直觉与实验管理——两个上限、四块显存、能复现

GPU Intuition and Experiment Management: Two Ceilings, Four Memory Buckets, and Reproducibility


算法工程师的工具箱(05):Hugging Face 生态——六个库与一次 LoRA SFT 的组装

The Hugging Face Ecosystem: Six Libraries, a Six-Line LoRA SFT, and Why Reading the Source Is the Fastest Way to Learn


算法工程师的工具箱(04):PyTorch 使用层(下)——混合精度、显存的账与多卡启用

PyTorch in Use, Part 2: Mixed Precision, the Memory Ledger and Turning On Multi-GPU


算法工程师的工具箱(03):PyTorch 使用层(上)——五个对象与二十行训练循环

PyTorch in Use, Part 1: Five Objects and a Twenty-Line Training Loop


算法工程师的工具箱(02):数据科学三剑客——NumPy 的形状直觉、Pandas 的错误分析、Matplotlib 的曲线

The Scientific Python Stack: Shapes and Broadcasting in NumPy, Error Analysis in Pandas, Reading Curves in Matplotlib


算法工程师的工具箱(01):Python 使用层——读懂训练代码的语法、流式过一遍语料、把实验写成脚本

Python in Use for Algorithm Engineers: The Syntax Behind Training Code, Streaming a Corpus, and Scripting an Experiment


算法工程师的工具箱:从一个想法到一次能跑的实验(总纲)

Tooling for AI Algorithm Engineers: NumPy, PyTorch, Hugging Face and the GPU in Your Head


算法工程师的数学(09):系列总结与通关自测

Mathematics for AI Algorithm Engineers: Series Recap and Final Self-Test


算法工程师的数学(08):统计推断与拟合——评测的置信区间与 scaling law

Statistical Inference and Curve Fitting: Confidence Intervals for Benchmarks and How Scaling Laws Are Fit


算法工程师的数学(07):导数、梯度与链式法则——softmax 的梯度与策略梯度

Derivatives, Gradients and the Chain Rule: The Softmax Gradient and the Policy Gradient


算法工程师的数学(06):熵、交叉熵与 KL——从困惑度到 DPO

Entropy, Cross-Entropy and KL Divergence: From Perplexity to the DPO Loss


算法工程师的数学(05):从最大似然到交叉熵——第一个要会推的 loss

From Maximum Likelihood to Cross-Entropy: The First Loss You Should Be Able to Derive


算法工程师的数学(04):概率入门——语言模型是一个条件分布

Probability Basics: A Language Model Is a Conditional Distribution


算法工程师的数学(03):正交与旋转、特征值与 SVD——从 RoPE 到 LoRA

Orthogonal Matrices, Rotations, Eigenvalues and SVD: Why RoPE Encodes Relative Position and Why LoRA Works


算法工程师的数学(02):内积、范数与余弦相似度

Inner Product, Norms and Cosine Similarity: One Language for Attention, Retrieval, Regularization and Quantization Error


算法工程师的数学(01):向量、矩阵与形状——一个 token 过一层要算多少

Vectors, Matrices and Shapes: How Much Compute Does One Token Through One Layer Take


算法工程师的数学:读公式不卡壳的最小集(总纲)

Mathematics for AI Algorithm Engineers: The Minimal Set to Read Papers and Derive Losses


AI 应用工程师学习地图:在非确定性组件之上做可靠产品

A Learning Roadmap for AI Application Engineers


AI 算法工程师学习地图:从数学基础到大模型训练

A Learning Roadmap for AI Algorithm Engineers in the LLM Era


AI 全栈学习地图:造模型、跑模型、用模型的三张图

One System, Three Roles — an Overview of the Three AI Learning Roadmaps


Math

算法工程师的数学(09):系列总结与通关自测

Mathematics for AI Algorithm Engineers: Series Recap and Final Self-Test


算法工程师的数学(08):统计推断与拟合——评测的置信区间与 scaling law

Statistical Inference and Curve Fitting: Confidence Intervals for Benchmarks and How Scaling Laws Are Fit


算法工程师的数学(07):导数、梯度与链式法则——softmax 的梯度与策略梯度

Derivatives, Gradients and the Chain Rule: The Softmax Gradient and the Policy Gradient


算法工程师的数学(06):熵、交叉熵与 KL——从困惑度到 DPO

Entropy, Cross-Entropy and KL Divergence: From Perplexity to the DPO Loss


算法工程师的数学(05):从最大似然到交叉熵——第一个要会推的 loss

From Maximum Likelihood to Cross-Entropy: The First Loss You Should Be Able to Derive


算法工程师的数学(04):概率入门——语言模型是一个条件分布

Probability Basics: A Language Model Is a Conditional Distribution


算法工程师的数学(03):正交与旋转、特征值与 SVD——从 RoPE 到 LoRA

Orthogonal Matrices, Rotations, Eigenvalues and SVD: Why RoPE Encodes Relative Position and Why LoRA Works


算法工程师的数学(02):内积、范数与余弦相似度

Inner Product, Norms and Cosine Similarity: One Language for Attention, Retrieval, Regularization and Quantization Error


算法工程师的数学(01):向量、矩阵与形状——一个 token 过一层要算多少

Vectors, Matrices and Shapes: How Much Compute Does One Token Through One Layer Take


算法工程师的数学:读公式不卡壳的最小集(总纲)

Mathematics for AI Algorithm Engineers: The Minimal Set to Read Papers and Derive Losses


Post-Training

LoRA 专题(04):系列总结与通关自测

LoRA Series: Recap and Final Self-Test


LoRA 专题(03):工程:adapter 文件、合并、多 LoRA 服务与参考模型

LoRA 03: In Production — Adapter Files, Merging, Multi-LoRA Serving and the Free Reference Model


LoRA 专题(02):选参:r、target_modules、alpha、lr 与 QLoRA / DoRA / PiSSA 各让模型变成什么

LoRA 02: Choosing r, target_modules, alpha and lr — and What QLoRA, DoRA and PiSSA Each Change


LoRA 专题(01):低秩假设:为什么两个瘦矩阵够用,以及它省了哪几本账

LoRA 01: The Low-Rank Hypothesis — Why Two Thin Matrices Suffice, and Exactly What They Save


LoRA 专题:SFT 的默认微调方式,从低秩假设到多租户服务(总纲)

LoRA, the Default Way to Fine-Tune: From the Low-Rank Hypothesis to Multi-Tenant Serving — Series Overview


后训练(09):系列总结与通关自测

Post-Training: Series Recap and Final Self-Test


后训练(08):评测:benchmark、LLM-as-judge、Arena 与污染

Evaluating LLMs: Benchmarks, LLM-as-Judge, Arenas and Contamination


后训练(07):蒸馏:logits 级、序列级与 on-policy

Knowledge Distillation for LLMs: Logit-Level, Sequence-Level and On-Policy


后训练(06):Agent 与工具调用的 RL:多轮环境、轨迹数据与延后的奖励

Agentic RL: Multi-Turn Environments, Trajectory Data and Delayed Rewards


后训练(05):推理模型与可验证奖励:R1 的配方、PRM 与 test-time compute

Reasoning Models and Verifiable Rewards: The R1 Recipe, Process Reward Models and Test-Time Compute


后训练(04):离线 RL:从 RLHF 目标推出 DPO 及其变体

Offline Preference Optimization: Deriving DPO from the RLHF Objective, and Its Family


后训练(03):在线 RL:PPO、GRPO 与 RLHF 三件套

Online RL for LLMs: PPO, GRPO and the Policy–Reward–Reference Trio


后训练(02):偏好数据与奖励模型:Bradley-Terry、pairwise loss 与 reward hacking

Preference Data and Reward Models: Bradley-Terry, Pairwise Loss and Reward Hacking


后训练(01):SFT:指令数据、chat template、loss mask 与参数高效微调

Supervised Fine-Tuning: Instruction Data, Chat Templates, Loss Masking and Parameter-Efficient Fine-Tuning


后训练:从 SFT 到可验证奖励(总纲)

Post-Training: From Supervised Fine-Tuning to Verifiable Rewards


Inference

扩散模型推理基础设施(10):系列总结与通关自测

Diffusion Model Inference Infrastructure: Series Recap and Final Self-Test


扩散模型推理基础设施(09):配置、评测与排障——从一张卡的推导到一条伪影的排查

Configuration, Evaluation and Troubleshooting for Diffusion Inference (with Series Summary)


扩散模型推理基础设施(07):serving 形态——请求、批、三段分离、附件、异步任务与成本

Serving Diffusion Models: Request Shapes, Batching, Disaggregation, Add-ons, Async Jobs and Cost


扩散模型推理基础设施(05):多卡并行——序列并行、CFG 并行与 PipeFusion,为什么不是张量并行

Multi-GPU Diffusion Inference: Sequence Parallelism, CFG Parallelism and PipeFusion


扩散模型推理基础设施(03):跨步冗余——TeaCache、First-Block Cache 一族的缓存与跳步

Temporal Redundancy Across Denoising Steps: TeaCache, First-Block Cache and Friends


扩散模型推理基础设施(02):单卡执行——attention 后端、编译、FP8 / INT4 与 offload

Single-GPU Execution: Attention Backends, Compilation, FP8 / INT4 and Offloading


扩散模型推理基础设施(01):负载画像——一次生成在 GPU 上发生什么

Workload Anatomy: FLOPs, Bytes and Seconds of One Diffusion Generation


扩散模型推理基础设施:图像与视频生成的 serving(总纲)

Diffusion Model Inference Infrastructure: Serving Image and Video Generation


高效推理与压缩(07):系列总结与通关自测

Efficient Inference and Model Compression: Series Recap and Final Self-Test


高效推理与压缩(06):剪枝、深度缩放与小模型配方

Pruning, Depth Scaling and How Small Models Are Made


高效推理与压缩(05):KV cache 压缩:量化、驱逐与稀疏 attention

KV Cache Compression: Quantization, Eviction and Sparse Attention


高效推理与压缩(04):量化感知训练、低比特与量化模型的评测

Quantization-Aware Training, Extreme Low-Bit and How to Evaluate a Quantized Model


高效推理与压缩(03):训练后量化:误差模型、GPTQ、AWQ 与旋转

Post-Training Quantization: Error Models, GPTQ, AWQ and Rotation


高效推理与压缩(02):投机解码:草稿、接受率与树

Speculative Decoding: Drafters, Acceptance Rates and Draft Trees


高效推理与压缩(01):解码策略、采样与约束生成

Decoding Strategies: Sampling, Truncation, Penalties and Constrained Generation


高效推理与压缩(算法侧):解码、投机、量化与 KV(总纲)

Efficient Inference and Model Compression: The Algorithm Side


Multimodal

多模态(10):系列总结与通关自测

Multimodal Models: Series Recap and Final Self-Test


多模态(09):自回归图像生成与统一模型

Autoregressive Image Generation and Unified Understanding-Generation Models


多模态(08):Latent diffusion、DiT 与文生图配方

Latent Diffusion, DiT and How Text-to-Image Models Are Built


多模态(07):扩散模型(下):score matching、flow matching 与 classifier-free guidance

Diffusion II: Score Matching, Flow Matching, Noise Schedules, and Classifier-Free Guidance


多模态(06):扩散模型(上):DDPM——加噪、去噪与「预测噪声」

Diffusion I: DDPM — Forward Noising, Reverse Denoising, the ELBO, and DDIM


多模态(05):语音(下):语音理解、语音生成与全双工

Speech II: Speech Understanding, Speech Generation (TTS), Omni Models and Full-Duplex Dialogue


多模态(04):语音(上):从波形到 token——mel 谱、Whisper 与神经 codec

Speech I: From Waveform to Tokens — Mel Spectrograms, Whisper, and Neural Codecs with RVQ


多模态(03):VLM 的训练:数据、阶段与评测

Training a VLM: Data, Stages, Evaluation and Hallucination


多模态(02):VLM 的结构:connector、注入方式与动态分辨率

VLM Architecture: Connectors, Injection Methods and Dynamic Resolution


多模态(01):视觉编码器:CLIP、SigLIP 与自监督 ViT

Vision Encoders: CLIP, SigLIP and Self-Supervised ViTs


多模态:从视觉编码器到扩散模型(总纲)

Multimodal Models: From Vision Encoders to Diffusion


Diffusion

扩散模型推理基础设施(10):系列总结与通关自测

Diffusion Model Inference Infrastructure: Series Recap and Final Self-Test


扩散模型推理基础设施(09):配置、评测与排障——从一张卡的推导到一条伪影的排查

Configuration, Evaluation and Troubleshooting for Diffusion Inference (with Series Summary)


扩散模型推理基础设施(08):三个引擎的对照导读——同一张图的请求在 SGLang Diffusion、vLLM-Omni 与 xDiT 里各走过什么

Three Engines Compared: One Request Through SGLang Diffusion, vLLM-Omni and xDiT


扩散模型推理基础设施(07):serving 形态——请求、批、三段分离、附件、异步任务与成本

Serving Diffusion Models: Request Shapes, Batching, Disaggregation, Add-ons, Async Jobs and Cost


扩散模型推理基础设施(06):少步与自回归——把步数变成系统参数

Few-Step and Autoregressive Generation: When the Step Count Becomes a System Parameter


扩散模型推理基础设施(05):多卡并行——序列并行、CFG 并行与 PipeFusion,为什么不是张量并行

Multi-GPU Diffusion Inference: Sequence Parallelism, CFG Parallelism and PipeFusion


扩散模型推理基础设施(04):视频——长序列 attention 的账与稀疏化

Video Diffusion: The Long-Sequence Attention Bill and How Sparsity Pays It


扩散模型推理基础设施(03):跨步冗余——TeaCache、First-Block Cache 一族的缓存与跳步

Temporal Redundancy Across Denoising Steps: TeaCache, First-Block Cache and Friends


扩散模型推理基础设施(02):单卡执行——attention 后端、编译、FP8 / INT4 与 offload

Single-GPU Execution: Attention Backends, Compilation, FP8 / INT4 and Offloading


扩散模型推理基础设施(01):负载画像——一次生成在 GPU 上发生什么

Workload Anatomy: FLOPs, Bytes and Seconds of One Diffusion Generation


扩散模型推理基础设施:图像与视频生成的 serving(总纲)

Diffusion Model Inference Infrastructure: Serving Image and Video Generation


多模态(10):系列总结与通关自测

Multimodal Models: Series Recap and Final Self-Test


多模态(08):Latent diffusion、DiT 与文生图配方

Latent Diffusion, DiT and How Text-to-Image Models Are Built


多模态(07):扩散模型(下):score matching、flow matching 与 classifier-free guidance

Diffusion II: Score Matching, Flow Matching, Noise Schedules, and Classifier-Free Guidance


多模态(06):扩散模型(上):DDPM——加噪、去噪与「预测噪声」

Diffusion I: DDPM — Forward Noising, Reverse Denoising, the ELBO, and DDIM


多模态:从视觉编码器到扩散模型(总纲)

Multimodal Models: From Vision Encoders to Diffusion


CUDA

AI 平台工程(02):容器里的 GPU——驱动、CUDA、device plugin 与镜像

GPUs in Containers: Driver, CUDA, Device Plugin, DRA and Images


RL 后训练基础设施(03):共置——训练器与推理引擎在同一组 GPU 上共存

Colocation: Handing GPU Memory Back and Forth Between Trainer and Rollout Engine


GPU Kernel 工程(11):系列总结与通关自测

GPU Kernel Engineering: Series Recap and Final Self-Test


GPU Kernel 工程(10):剖析、测试与贡献——把 kernel 做成产品

Profiling, Testing and Contributing: Turning a Kernel into a Product


GPU Kernel 工程(09):量化与融合 kernel——推理系统的其余部分

Quantized and Fused Kernels: The Rest of the Inference Stack


GPU Kernel 工程(08):Attention Kernel——FlashAttention 与 PagedAttention

Attention Kernels: FlashAttention and PagedAttention from Derivation to Code


GPU Kernel 工程(07):Triton——块级编程与编译器的边界

Triton: Block-Level Programming and Where the Compiler Stops


GPU Kernel 工程(06):Tensor Core、CUTLASS 与 CuTe

Tensor Cores, CUTLASS and CuTe: Programming the Matrix Units


GPU Kernel 工程(05):GEMM——从 naive 到分块

GEMM from Naive to Tiled: Reaching the Compute Ceiling on CUDA Cores


GPU Kernel 工程(04):共享内存与 reduction——softmax、LayerNorm 与 online softmax

Shared Memory and Reductions: Softmax, LayerNorm and Online Softmax


GPU Kernel 工程(03):访存合并与 elementwise kernel

Memory Coalescing and Elementwise Kernels: Hitting the Bandwidth Ceiling


GPU Kernel 工程(02):CUDA 编程模型与第一个 kernel

The CUDA Programming Model and Your First Kernel, Measured


GPU Kernel 工程(01):GPU 为什么这样设计——硬件结构与 Roofline

Why GPUs Look the Way They Do: Architecture and the Roofline Model


GPU Kernel 工程:从 CUDA 执行模型到 FlashAttention(总纲)

GPU Kernel Engineering, from the CUDA Execution Model to FlashAttention


Triton

GPU Kernel 工程(11):系列总结与通关自测

GPU Kernel Engineering: Series Recap and Final Self-Test


GPU Kernel 工程(10):剖析、测试与贡献——把 kernel 做成产品

Profiling, Testing and Contributing: Turning a Kernel into a Product


GPU Kernel 工程(09):量化与融合 kernel——推理系统的其余部分

Quantized and Fused Kernels: The Rest of the Inference Stack


GPU Kernel 工程(08):Attention Kernel——FlashAttention 与 PagedAttention

Attention Kernels: FlashAttention and PagedAttention from Derivation to Code


GPU Kernel 工程(07):Triton——块级编程与编译器的边界

Triton: Block-Level Programming and Where the Compiler Stops


GPU Kernel 工程(06):Tensor Core、CUTLASS 与 CuTe

Tensor Cores, CUTLASS and CuTe: Programming the Matrix Units


GPU Kernel 工程(05):GEMM——从 naive 到分块

GEMM from Naive to Tiled: Reaching the Compute Ceiling on CUDA Cores


GPU Kernel 工程(04):共享内存与 reduction——softmax、LayerNorm 与 online softmax

Shared Memory and Reductions: Softmax, LayerNorm and Online Softmax


GPU Kernel 工程(03):访存合并与 elementwise kernel

Memory Coalescing and Elementwise Kernels: Hitting the Bandwidth Ceiling


GPU Kernel 工程(02):CUDA 编程模型与第一个 kernel

The CUDA Programming Model and Your First Kernel, Measured


GPU Kernel 工程(01):GPU 为什么这样设计——硬件结构与 Roofline

Why GPUs Look the Way They Do: Architecture and the Roofline Model


GPU Kernel 工程:从 CUDA 执行模型到 FlashAttention(总纲)

GPU Kernel Engineering, from the CUDA Execution Model to FlashAttention


GPU

AI 平台工程(09):系列总结与通关自测

AI Platform Engineering: Series Recap and Final Self-Test


AI 平台工程(08):可观测、成本与 FinOps

Observability, Cost and FinOps: from DCGM to the Token Bill


AI 平台工程(07):模型网关与多租户——路由、配额与灰度

The Model Gateway: KV-Aware Routing, Multi-Tenancy, Quotas and Canaries


AI 平台工程(06):Serving 平台——从 InferenceService 到 llm-d

Serving Platforms: KServe, Triton, Ray Serve, LeaderWorkerSet and llm-d


AI 平台工程(05):网络与存储——RDMA 进容器、并行文件系统与 checkpoint I/O

Networking and Storage: RDMA in Containers, Parallel File Systems and Checkpoint I/O


AI 平台工程(04):GPU 共享与切分——MIG、时间片、MPS 与 HAMi

Sharing and Partitioning GPUs: MIG, Time-Slicing, MPS and HAMi


AI 平台工程(03):AI 任务调度——gang scheduling、队列与拓扑感知

Scheduling AI Jobs: Gang Scheduling, Queues, Quotas and Topology Awareness


AI 平台工程(02):容器里的 GPU——驱动、CUDA、device plugin 与镜像

GPUs in Containers: Driver, CUDA, Device Plugin, DRA and Images


AI 平台工程(01):引擎的需求清单与平台的整体架构

What Engines Demand from the Platform, and the Platform's Two Layers


AI 平台工程:资源层与交付层(总纲)

AI Platform Engineering: the Resource Layer and the Delivery Layer


通信与互联(09):系列总结与通关自测

Communication and Interconnect: Series Recap and Final Self-Test


通信与互联(08):MoE 的通信——all-to-all、DeepEP 与 GPU 发起的通信

Communication for MoE: All-to-All, DeepEP and GPU-Initiated Networking


通信与互联(07):推理侧的通信——custom all-reduce 与 KV 传输

Communication on the Inference Side: Custom All-Reduce and KV Cache Transfer


通信与互联(06):nccl-tests、调优与排障——从带宽曲线到 hang

nccl-tests, Tuning and Debugging Hangs: From Bandwidth Curves to Flight Recorder


通信与互联(05):PyTorch 的通信栈——ProcessGroupNCCL、stream 语义与计算通信重叠

PyTorch's Communication Stack: ProcessGroupNCCL, Stream Semantics, and Compute-Communication Overlap


通信与互联(04):NCCL 架构——拓扑探测、channel、算法与协议

NCCL Architecture: Topology Detection, Channels, Algorithms and Protocols


通信与互联(03):RDMA 与 GPUDirect——绕过 CPU 和主机内存的数据通路

RDMA and GPUDirect: Bypassing the CPU and Host Memory


通信与互联(02):硬件互联——PCIe、NVLink、NVSwitch 与网络拓扑

Hardware Interconnect: PCIe, NVLink, NVSwitch and Network Topology


通信与互联(01):集合通信原语与代价模型——α-β 模型与 ring all-reduce

Collective Communication Primitives and the Alpha-Beta Cost Model: Deriving Ring All-Reduce


通信与互联:从 NCCL 到 RDMA(总纲)

Communication and Interconnect for AI-Infra, from NCCL to RDMA


GPU Kernel 工程(11):系列总结与通关自测

GPU Kernel Engineering: Series Recap and Final Self-Test


GPU Kernel 工程(10):剖析、测试与贡献——把 kernel 做成产品

Profiling, Testing and Contributing: Turning a Kernel into a Product


GPU Kernel 工程(09):量化与融合 kernel——推理系统的其余部分

Quantized and Fused Kernels: The Rest of the Inference Stack


GPU Kernel 工程(08):Attention Kernel——FlashAttention 与 PagedAttention

Attention Kernels: FlashAttention and PagedAttention from Derivation to Code


GPU Kernel 工程(07):Triton——块级编程与编译器的边界

Triton: Block-Level Programming and Where the Compiler Stops


GPU Kernel 工程(06):Tensor Core、CUTLASS 与 CuTe

Tensor Cores, CUTLASS and CuTe: Programming the Matrix Units


GPU Kernel 工程(05):GEMM——从 naive 到分块

GEMM from Naive to Tiled: Reaching the Compute Ceiling on CUDA Cores


GPU Kernel 工程(04):共享内存与 reduction——softmax、LayerNorm 与 online softmax

Shared Memory and Reductions: Softmax, LayerNorm and Online Softmax


GPU Kernel 工程(03):访存合并与 elementwise kernel

Memory Coalescing and Elementwise Kernels: Hitting the Bandwidth Ceiling


GPU Kernel 工程(02):CUDA 编程模型与第一个 kernel

The CUDA Programming Model and Your First Kernel, Measured


GPU Kernel 工程(01):GPU 为什么这样设计——硬件结构与 Roofline

Why GPUs Look the Way They Do: Architecture and the Roofline Model


GPU Kernel 工程:从 CUDA 执行模型到 FlashAttention(总纲)

GPU Kernel Engineering, from the CUDA Execution Model to FlashAttention


vLLM

AI-Infra 开源贡献指南(05):系列总结与通关自测

Contributing to AI-Infra Open Source: Series Recap and Final Self-Test


AI-Infra 开源贡献指南(04):两个真实 PR 的完整走读——PyTorch 与 vLLM

Two Real Pull Requests, End to End: One in PyTorch, One in vLLM


AI-Infra 开源贡献指南(03):做出一个能被合入的改动

Landing a Mergeable Change: Diff, Tests, Benchmarks, PR, CI and Review


AI-Infra 开源贡献指南(02):找到切入点——从 issue、RFC 到性能回归

Finding Your Entry Point: Issues, RFCs, Roadmaps, CI Failures and Regressions


AI-Infra 开源贡献指南(01):读懂一个百万行的代码库

Reading a Million-Line Codebase: Maps, Entry Points, Symbols, Tests and History


AI-Infra 开源贡献指南(总纲)

A Guide to Contributing to AI-Infra Open Source Projects


AI 平台工程(06):Serving 平台——从 InferenceService 到 llm-d

Serving Platforms: KServe, Triton, Ray Serve, LeaderWorkerSet and llm-d


扩散模型推理基础设施(10):系列总结与通关自测

Diffusion Model Inference Infrastructure: Series Recap and Final Self-Test


扩散模型推理基础设施(08):三个引擎的对照导读——同一张图的请求在 SGLang Diffusion、vLLM-Omni 与 xDiT 里各走过什么

Three Engines Compared: One Request Through SGLang Diffusion, vLLM-Omni and xDiT


扩散模型推理基础设施(07):serving 形态——请求、批、三段分离、附件、异步任务与成本

Serving Diffusion Models: Request Shapes, Batching, Disaggregation, Add-ons, Async Jobs and Cost


扩散模型推理基础设施:图像与视频生成的 serving(总纲)

Diffusion Model Inference Infrastructure: Serving Image and Video Generation


RL 后训练基础设施(09):系列总结与通关自测

RL Post-Training Infrastructure: Series Recap and Final Self-Test


RL 后训练基础设施(08):配置、可观测与排障——从一张卡的比例到一条 hang 的排查

Configuration, Observability and Troubleshooting for RL Post-Training Systems


RL 后训练基础设施(07):verl 源码导读——从一个 GRPO 配置追到每个 worker

Reading verl: From One GRPO Config to Every Worker


RL 后训练基础设施(06):Agentic rollout——多轮、工具、沙箱与环境服务

Agentic Rollout: Multi-Turn Trajectories, Tools, Sandboxes and Environment Services


RL 后训练基础设施(04):权重同步——从训练分片到推理分片

Weight Synchronization: From Training Shards to Inference Shards


RL 后训练基础设施(03):共置——训练器与推理引擎在同一组 GPU 上共存

Colocation: Handing GPU Memory Back and Forth Between Trainer and Rollout Engine


RL 后训练基础设施(02):系统形态——共置、分离与异步

RL System Topologies: Colocated, Disaggregated and Asynchronous


RL 后训练基础设施(01):负载画像——一步 RL 里发生什么

Anatomy of an RL Step: Rollout, Reward and Train


RL 后训练基础设施:rollout 与训练如何共享一组 GPU(总纲)

RL Post-Training Infrastructure: How Rollout and Training Share the Same GPUs


大模型推理系统揭秘(15):vLLM 与 SGLang:同一个请求穿过两套 Serving 系统

vLLM vs SGLang: One Request Through Two LLM Serving Systems


LoRA 专题(03):工程:adapter 文件、合并、多 LoRA 服务与参考模型

LoRA 03: In Production — Adapter Files, Merging, Multi-LoRA Serving and the Free Reference Model


NCCL

扩散模型推理基础设施(05):多卡并行——序列并行、CFG 并行与 PipeFusion,为什么不是张量并行

Multi-GPU Diffusion Inference: Sequence Parallelism, CFG Parallelism and PipeFusion


RL 后训练基础设施(04):权重同步——从训练分片到推理分片

Weight Synchronization: From Training Shards to Inference Shards


通信与互联(09):系列总结与通关自测

Communication and Interconnect: Series Recap and Final Self-Test


通信与互联(08):MoE 的通信——all-to-all、DeepEP 与 GPU 发起的通信

Communication for MoE: All-to-All, DeepEP and GPU-Initiated Networking


通信与互联(07):推理侧的通信——custom all-reduce 与 KV 传输

Communication on the Inference Side: Custom All-Reduce and KV Cache Transfer


通信与互联(06):nccl-tests、调优与排障——从带宽曲线到 hang

nccl-tests, Tuning and Debugging Hangs: From Bandwidth Curves to Flight Recorder


通信与互联(05):PyTorch 的通信栈——ProcessGroupNCCL、stream 语义与计算通信重叠

PyTorch's Communication Stack: ProcessGroupNCCL, Stream Semantics, and Compute-Communication Overlap


通信与互联(04):NCCL 架构——拓扑探测、channel、算法与协议

NCCL Architecture: Topology Detection, Channels, Algorithms and Protocols


通信与互联(03):RDMA 与 GPUDirect——绕过 CPU 和主机内存的数据通路

RDMA and GPUDirect: Bypassing the CPU and Host Memory


通信与互联(02):硬件互联——PCIe、NVLink、NVSwitch 与网络拓扑

Hardware Interconnect: PCIe, NVLink, NVSwitch and Network Topology


通信与互联(01):集合通信原语与代价模型——α-β 模型与 ring all-reduce

Collective Communication Primitives and the Alpha-Beta Cost Model: Deriving Ring All-Reduce


通信与互联:从 NCCL 到 RDMA(总纲)

Communication and Interconnect for AI-Infra, from NCCL to RDMA


RDMA

AI 平台工程(05):网络与存储——RDMA 进容器、并行文件系统与 checkpoint I/O

Networking and Storage: RDMA in Containers, Parallel File Systems and Checkpoint I/O


通信与互联(09):系列总结与通关自测

Communication and Interconnect: Series Recap and Final Self-Test


通信与互联(08):MoE 的通信——all-to-all、DeepEP 与 GPU 发起的通信

Communication for MoE: All-to-All, DeepEP and GPU-Initiated Networking


通信与互联(07):推理侧的通信——custom all-reduce 与 KV 传输

Communication on the Inference Side: Custom All-Reduce and KV Cache Transfer


通信与互联(06):nccl-tests、调优与排障——从带宽曲线到 hang

nccl-tests, Tuning and Debugging Hangs: From Bandwidth Curves to Flight Recorder


通信与互联(05):PyTorch 的通信栈——ProcessGroupNCCL、stream 语义与计算通信重叠

PyTorch's Communication Stack: ProcessGroupNCCL, Stream Semantics, and Compute-Communication Overlap


通信与互联(04):NCCL 架构——拓扑探测、channel、算法与协议

NCCL Architecture: Topology Detection, Channels, Algorithms and Protocols


通信与互联(03):RDMA 与 GPUDirect——绕过 CPU 和主机内存的数据通路

RDMA and GPUDirect: Bypassing the CPU and Host Memory


通信与互联(02):硬件互联——PCIe、NVLink、NVSwitch 与网络拓扑

Hardware Interconnect: PCIe, NVLink, NVSwitch and Network Topology


通信与互联(01):集合通信原语与代价模型——α-β 模型与 ring all-reduce

Collective Communication Primitives and the Alpha-Beta Cost Model: Deriving Ring All-Reduce


通信与互联:从 NCCL 到 RDMA(总纲)

Communication and Interconnect for AI-Infra, from NCCL to RDMA


Megatron

RL 后训练基础设施(09):系列总结与通关自测

RL Post-Training Infrastructure: Series Recap and Final Self-Test


RL 后训练基础设施(04):权重同步——从训练分片到推理分片

Weight Synchronization: From Training Shards to Inference Shards


RL 后训练基础设施:rollout 与训练如何共享一组 GPU(总纲)

RL Post-Training Infrastructure: How Rollout and Training Share the Same GPUs


大规模训练工程(09):系列总结与通关自测

Large-Scale Training Engineering: Series Recap and Final Self-Test


大规模训练工程(08):长时训练的可观测与运维——从指标到 hang 排查

Observability and Operations for Long-Running Training


大规模训练工程(07):训练稳定性与数据管线——loss spike、梯度范数、数据混合与流式加载

Training Stability and the Data Pipeline


大规模训练工程(06):容错与弹性——故障率数学、straggler、SDC 与弹性训练

Fault Tolerance and Elasticity: Failure Math, Stragglers, SDC and Elastic Training


大规模训练工程(05):分布式 checkpoint——格式、异步保存与重分片恢复

Distributed Checkpoint: Format, Async Save and Resharding


大规模训练工程(04):千卡配置实战——并行搭配、micro-batch、激活重计算与 MFU 调优

Configuring a Thousand-GPU Job: Parallelism, Micro-batch, Recompute and MFU


大规模训练工程(03):三个框架——Megatron-LM、DeepSpeed 与 torchtitan 的架构对比与源码导读

Megatron-LM, DeepSpeed and torchtitan: Architecture and Source Guide


大规模训练工程(02):并行策略全景——每种并行切的是哪种状态

A Map of Parallelism: Which State Does Each Strategy Shard


大规模训练工程(01):训练任务的状态解剖——显存账与 MFU

Anatomy of Training State: Memory Accounting and MFU


大规模训练工程:从并行策略到容错恢复(总纲)

Large-Scale Training Engineering, from Parallelism to Fault Tolerance


Distributed Training

RL 后训练基础设施(09):系列总结与通关自测

RL Post-Training Infrastructure: Series Recap and Final Self-Test


RL 后训练基础设施(08):配置、可观测与排障——从一张卡的比例到一条 hang 的排查

Configuration, Observability and Troubleshooting for RL Post-Training Systems


RL 后训练基础设施(05):异步与 off-policy——把同步的墙拆掉之后要补什么

Asynchrony and Off-Policy: What You Owe After Tearing Down the Synchronization Wall


RL 后训练基础设施(02):系统形态——共置、分离与异步

RL System Topologies: Colocated, Disaggregated and Asynchronous


RL 后训练基础设施(01):负载画像——一步 RL 里发生什么

Anatomy of an RL Step: Rollout, Reward and Train


RL 后训练基础设施:rollout 与训练如何共享一组 GPU(总纲)

RL Post-Training Infrastructure: How Rollout and Training Share the Same GPUs


大规模训练工程(09):系列总结与通关自测

Large-Scale Training Engineering: Series Recap and Final Self-Test


大规模训练工程(08):长时训练的可观测与运维——从指标到 hang 排查

Observability and Operations for Long-Running Training


大规模训练工程(07):训练稳定性与数据管线——loss spike、梯度范数、数据混合与流式加载

Training Stability and the Data Pipeline


大规模训练工程(06):容错与弹性——故障率数学、straggler、SDC 与弹性训练

Fault Tolerance and Elasticity: Failure Math, Stragglers, SDC and Elastic Training


大规模训练工程(05):分布式 checkpoint——格式、异步保存与重分片恢复

Distributed Checkpoint: Format, Async Save and Resharding


大规模训练工程(04):千卡配置实战——并行搭配、micro-batch、激活重计算与 MFU 调优

Configuring a Thousand-GPU Job: Parallelism, Micro-batch, Recompute and MFU


大规模训练工程(03):三个框架——Megatron-LM、DeepSpeed 与 torchtitan 的架构对比与源码导读

Megatron-LM, DeepSpeed and torchtitan: Architecture and Source Guide


大规模训练工程(02):并行策略全景——每种并行切的是哪种状态

A Map of Parallelism: Which State Does Each Strategy Shard


大规模训练工程(01):训练任务的状态解剖——显存账与 MFU

Anatomy of Training State: Memory Accounting and MFU


大规模训练工程:从并行策略到容错恢复(总纲)

Large-Scale Training Engineering, from Parallelism to Fault Tolerance


大模型推理

大模型推理系统揭秘(16):系列总结与通关自测

Deep Dive into vLLM: Series Recap and Final Self-Test


大模型推理系统揭秘(15):vLLM 与 SGLang:同一个请求穿过两套 Serving 系统

vLLM vs SGLang: One Request Through Two LLM Serving Systems


大模型推理系统揭秘(14):回到源码:一次请求在 vLLM 内部的真实旅程


大模型推理系统揭秘(13):Serving Infra 的下一站:从模型执行器到分布式智能操作系统


大模型推理系统揭秘(12):硬件解耦:如何不让芯片差异污染 Serving 核心?


大模型推理系统揭秘(11):请求形态的扩展:multi-LoRA 与多模态


大模型推理系统揭秘(10):模型适配:如何跟上变化极快的模型世界?


大模型推理系统揭秘(09):PD 分离:从资源混部走向计算解耦


大模型推理系统揭秘(08):Multi-GPU:一张卡不够时如何扩展?


大模型推理系统揭秘(07):解码的扩展:采样、投机解码与结构化输出


大模型推理系统揭秘(06):GPU 执行:如何让每个 Token 算得更快?


大模型推理系统揭秘(05):KV Cache:LLM Serving 的第一号内存问题


大模型推理系统揭秘(04):Scheduler:GPU 这一轮到底给谁用?


大模型推理系统揭秘(03):鸟瞰 vLLM:一个请求如何穿过整个推理系统?


大模型推理系统揭秘(02):如何衡量一个 LLM Serving 系统?


大模型推理系统揭秘(01):为什么 LLM Serving 比传统 DL 推理难?


大模型推理系统揭秘:从 vLLM 看 LLM Serving Infra 核心技术(总纲)


RL

RL 后训练基础设施(09):系列总结与通关自测

RL Post-Training Infrastructure: Series Recap and Final Self-Test


RL 后训练基础设施(08):配置、可观测与排障——从一张卡的比例到一条 hang 的排查

Configuration, Observability and Troubleshooting for RL Post-Training Systems


RL 后训练基础设施(07):verl 源码导读——从一个 GRPO 配置追到每个 worker

Reading verl: From One GRPO Config to Every Worker


RL 后训练基础设施(06):Agentic rollout——多轮、工具、沙箱与环境服务

Agentic Rollout: Multi-Turn Trajectories, Tools, Sandboxes and Environment Services


RL 后训练基础设施(05):异步与 off-policy——把同步的墙拆掉之后要补什么

Asynchrony and Off-Policy: What You Owe After Tearing Down the Synchronization Wall


RL 后训练基础设施(04):权重同步——从训练分片到推理分片

Weight Synchronization: From Training Shards to Inference Shards


RL 后训练基础设施(03):共置——训练器与推理引擎在同一组 GPU 上共存

Colocation: Handing GPU Memory Back and Forth Between Trainer and Rollout Engine


RL 后训练基础设施(02):系统形态——共置、分离与异步

RL System Topologies: Colocated, Disaggregated and Asynchronous


RL 后训练基础设施(01):负载画像——一步 RL 里发生什么

Anatomy of an RL Step: Rollout, Reward and Train


RL 后训练基础设施:rollout 与训练如何共享一组 GPU(总纲)

RL Post-Training Infrastructure: How Rollout and Training Share the Same GPUs


verl

RL 后训练基础设施(09):系列总结与通关自测

RL Post-Training Infrastructure: Series Recap and Final Self-Test


RL 后训练基础设施(08):配置、可观测与排障——从一张卡的比例到一条 hang 的排查

Configuration, Observability and Troubleshooting for RL Post-Training Systems


RL 后训练基础设施(07):verl 源码导读——从一个 GRPO 配置追到每个 worker

Reading verl: From One GRPO Config to Every Worker


RL 后训练基础设施(06):Agentic rollout——多轮、工具、沙箱与环境服务

Agentic Rollout: Multi-Turn Trajectories, Tools, Sandboxes and Environment Services


RL 后训练基础设施(05):异步与 off-policy——把同步的墙拆掉之后要补什么

Asynchrony and Off-Policy: What You Owe After Tearing Down the Synchronization Wall


RL 后训练基础设施(04):权重同步——从训练分片到推理分片

Weight Synchronization: From Training Shards to Inference Shards


RL 后训练基础设施(03):共置——训练器与推理引擎在同一组 GPU 上共存

Colocation: Handing GPU Memory Back and Forth Between Trainer and Rollout Engine


RL 后训练基础设施(02):系统形态——共置、分离与异步

RL System Topologies: Colocated, Disaggregated and Asynchronous


RL 后训练基础设施(01):负载画像——一步 RL 里发生什么

Anatomy of an RL Step: Rollout, Reward and Train


RL 后训练基础设施:rollout 与训练如何共享一组 GPU(总纲)

RL Post-Training Infrastructure: How Rollout and Training Share the Same GPUs


Kubernetes

AI 平台工程(09):系列总结与通关自测

AI Platform Engineering: Series Recap and Final Self-Test


AI 平台工程(08):可观测、成本与 FinOps

Observability, Cost and FinOps: from DCGM to the Token Bill


AI 平台工程(07):模型网关与多租户——路由、配额与灰度

The Model Gateway: KV-Aware Routing, Multi-Tenancy, Quotas and Canaries


AI 平台工程(06):Serving 平台——从 InferenceService 到 llm-d

Serving Platforms: KServe, Triton, Ray Serve, LeaderWorkerSet and llm-d


AI 平台工程(05):网络与存储——RDMA 进容器、并行文件系统与 checkpoint I/O

Networking and Storage: RDMA in Containers, Parallel File Systems and Checkpoint I/O


AI 平台工程(04):GPU 共享与切分——MIG、时间片、MPS 与 HAMi

Sharing and Partitioning GPUs: MIG, Time-Slicing, MPS and HAMi


AI 平台工程(03):AI 任务调度——gang scheduling、队列与拓扑感知

Scheduling AI Jobs: Gang Scheduling, Queues, Quotas and Topology Awareness


AI 平台工程(02):容器里的 GPU——驱动、CUDA、device plugin 与镜像

GPUs in Containers: Driver, CUDA, Device Plugin, DRA and Images


AI 平台工程(01):引擎的需求清单与平台的整体架构

What Engines Demand from the Platform, and the Platform's Two Layers


AI 平台工程:资源层与交付层(总纲)

AI Platform Engineering: the Resource Layer and the Delivery Layer


RL 后训练基础设施(06):Agentic rollout——多轮、工具、沙箱与环境服务

Agentic Rollout: Multi-Turn Trajectories, Tools, Sandboxes and Environment Services


AI-Application

Prompt 与上下文工程:模型这一步该看到什么(总纲)

Prompt and Context Engineering: What the Model Should See at This Step


模型作为组件(07):系列总结与通关自测

The Model as a Component: Series Recap and Final Self-Test


模型作为组件(06):客户端工程——重试、超时、幂等、限流与流式解析

Client Engineering for LLM APIs: Retries, Timeouts, Idempotency, Rate Limits and Stream Parsing


模型作为组件(05):选型——榜单的失真、自己的评测集与供应商的弃用周期

Model Selection Beyond Leaderboards: Your Own Evals, Open vs Closed, Routing and Deprecation Cycles


模型作为组件(04):成本与延迟的账——一次调用花多少钱、慢在哪一段

The Token Cost and Latency Ledger for LLM Applications


模型作为组件(03):API 契约(二)——推理模型:thinking、effort 与跨轮的推理状态

The LLM API Contract, Part 2: Reasoning Models — Thinking, Effort and Cross-turn State


模型作为组件(02):API 契约(一)——消息、工具调用、结构化输出与流式,四家 API 的共同骨架

The LLM API Contract, Part 1: Messages, Tool Calls, Structured Output and Streaming


模型作为组件(01):失效模式——非确定性、幻觉、上下文与越界,把「模型会出错」拆成七条可检测的性质

Failure Modes of LLMs as Components: Nondeterminism, Hallucination, Context and Overreach


模型作为组件:契约、失效模式与选型(总纲)

The Model as a Component: Contract, Failure Modes and Selection


×