Toggle navigation
Arganzheng's Blog
Tech
Life
Slides
Archive
Tags
About
Search
Archive
Go to article
×
2026
September
2026/09/14
»
[Tech]
AI-Infra 开源贡献指南(01):读懂一个百万行的代码库
Reading a Million-Line Codebase: Maps, Entry Points, Symbols, Tests and History
2026/09/13
»
[Tech]
AI-Infra 开源贡献指南(总纲)
A Guide to Contributing to AI-Infra Open Source Projects
2026/09/12
»
[Tech]
AI 平台工程(08):可观测、成本与 FinOps
Observability, Cost and FinOps: from DCGM to the Token Bill
2026/09/11
»
[Tech]
AI 平台工程(07):模型网关与多租户——路由、配额与灰度
The Model Gateway: KV-Aware Routing, Multi-Tenancy, Quotas and Canaries
2026/09/10
»
[Tech]
AI 平台工程(06):Serving 平台——从 InferenceService 到 llm-d
Serving Platforms: KServe, Triton, Ray Serve, LeaderWorkerSet and llm-d
2026/09/10
»
[Life]
个人博客终于迎来了久违的更新
写了近百篇文章,顺手把博客也重做了一遍
2026/09/09
»
[Tech]
AI 平台工程(05):网络与存储——RDMA 进容器、并行文件系统与 checkpoint I/O
Networking and Storage: RDMA in Containers, Parallel File Systems and Checkpoint I/O
2026/09/09
»
[Tech]
keynote 布局演示:给一份幻灯片配上讲稿
上面是可以翻页的幻灯片,下面是文字稿、参考资料和评论区
2026/09/08
»
[Tech]
AI 平台工程(04):GPU 共享与切分——MIG、时间片、MPS 与 HAMi
Sharing and Partitioning GPUs: MIG, Time-Slicing, MPS and HAMi
2026/09/07
»
[Tech]
AI 平台工程(03):AI 任务调度——gang scheduling、队列与拓扑感知
Scheduling AI Jobs: Gang Scheduling, Queues, Quotas and Topology Awareness
2026/09/06
»
[Tech]
AI 平台工程(02):容器里的 GPU——驱动、CUDA、device plugin 与镜像
GPUs in Containers: Driver, CUDA, Device Plugin, DRA and Images
2026/09/05
»
[Tech]
AI 平台工程(01):引擎的需求清单与平台的整体架构
What Engines Demand from the Platform, and the Platform's Two Layers
2026/09/04
»
[Tech]
AI 平台工程:资源层与交付层(总纲)
AI Platform Engineering: the Resource Layer and the Delivery Layer
2026/09/03
»
[Tech]
RL 后训练基础设施(08):配置、可观测与排障——从一张卡的比例到一条 hang 的排查
Configuration, Observability and Troubleshooting for RL Post-Training Systems
2026/09/02
»
[Tech]
RL 后训练基础设施(07):verl 源码导读——从一个 GRPO 配置追到每个 worker
Reading verl: From One GRPO Config to Every Worker
2026/09/01
»
[Tech]
RL 后训练基础设施(06):Agentic rollout——多轮、工具、沙箱与环境服务
Agentic Rollout: Multi-Turn Trajectories, Tools, Sandboxes and Environment Services
August
2026/08/31
»
[Tech]
RL 后训练基础设施(05):异步与 off-policy——把同步的墙拆掉之后要补什么
Asynchrony and Off-Policy: What You Owe After Tearing Down the Synchronization Wall
2026/08/30
»
[Tech]
RL 后训练基础设施(04):权重同步——从训练分片到推理分片
Weight Synchronization: From Training Shards to Inference Shards
2026/08/29
»
[Tech]
RL 后训练基础设施(03):共置——训练器与推理引擎在同一组 GPU 上共存
Colocation: Handing GPU Memory Back and Forth Between Trainer and Rollout Engine
2026/08/28
»
[Tech]
RL 后训练基础设施(02):系统形态——共置、分离与异步
RL System Topologies: Colocated, Disaggregated and Asynchronous
2026/08/27
»
[Tech]
RL 后训练基础设施(01):负载画像——一步 RL 里发生什么
Anatomy of an RL Step: Rollout, Reward and Train
2026/08/26
»
[Tech]
RL 后训练基础设施:rollout 与训练如何共享一组 GPU(总纲)
RL Post-Training Infrastructure: How Rollout and Training Share the Same GPUs
2026/08/25
»
[Tech]
大模型推理系统揭秘(14):回到源码:一次请求在 vLLM 内部的真实旅程
2026/08/24
»
[Tech]
大模型推理系统揭秘(13):Serving Infra 的下一站:从模型执行器到分布式智能操作系统
2026/08/23
»
[Tech]
大模型推理系统揭秘(12):PD 分离:从资源混部走向计算解耦
2026/08/22
»
[Tech]
大模型推理系统揭秘(11):硬件解耦:如何不让芯片差异污染 Serving 核心?
2026/08/21
»
[Tech]
大模型推理系统揭秘(10):请求形态的扩展:multi-LoRA 与多模态
2026/08/20
»
[Tech]
大模型推理系统揭秘(09):模型适配:如何跟上变化极快的模型世界?
2026/08/19
»
[Tech]
大模型推理系统揭秘(08):Multi-GPU:一张卡不够时如何扩展?
2026/08/18
»
[Tech]
大模型推理系统揭秘(07):解码的扩展:采样、投机解码与结构化输出
2026/08/17
»
[Tech]
大模型推理系统揭秘(06):GPU 执行:如何让每个 Token 算得更快?
2026/08/16
»
[Tech]
大模型推理系统揭秘(05):KV Cache:LLM Serving 的第一号内存问题
2026/08/15
»
[Tech]
大模型推理系统揭秘(04):Scheduler:GPU 这一轮到底给谁用?
2026/08/14
»
[Tech]
大模型推理系统揭秘(03):鸟瞰 vLLM:一个请求如何穿过整个推理系统?
2026/08/13
»
[Tech]
大模型推理系统揭秘(02):如何衡量一个 LLM Serving 系统?
2026/08/12
»
[Tech]
大模型推理系统揭秘(01):为什么 LLM Serving 比传统 DL 推理难?
2026/08/11
»
[Tech]
大模型推理系统揭秘:从 vLLM 看 LLM Serving Infra 核心技术(总纲)
2026/08/03
»
[Tech]
致读者的一封信
这个博客写什么、有哪些好玩的功能、怎么用——以及一份 FAQ
2026/08/01
»
[Slides]
用 Markdown 写在线 PPT
reveal.js + Jekyll 演示 / 使用说明
July
2026/07/29
»
[Tech]
大规模训练工程(08):长时训练的可观测与运维——从指标到 hang 排查
Observability and Operations for Long-Running Training
2026/07/27
»
[Tech]
大规模训练工程(07):训练稳定性与数据管线——loss spike、梯度范数、数据混合与流式加载
Training Stability and the Data Pipeline
2026/07/25
»
[Tech]
大规模训练工程(06):容错与弹性——故障率数学、straggler、SDC 与弹性训练
Fault Tolerance and Elasticity: Failure Math, Stragglers, SDC and Elastic Training
2026/07/23
»
[Tech]
大规模训练工程(05):分布式 checkpoint——格式、异步保存与重分片恢复
Distributed Checkpoint: Format, Async Save and Resharding
2026/07/21
»
[Tech]
大规模训练工程(04):千卡配置实战——并行搭配、micro-batch、激活重计算与 MFU 调优
Configuring a Thousand-GPU Job: Parallelism, Micro-batch, Recompute and MFU
2026/07/19
»
[Tech]
大规模训练工程(03):三个框架——Megatron-LM、DeepSpeed 与 torchtitan 的架构对比与源码导读
Megatron-LM, DeepSpeed and torchtitan: Architecture and Source Guide
2026/07/17
»
[Tech]
大规模训练工程(02):并行策略全景——每种并行切的是哪种状态
A Map of Parallelism: Which State Does Each Strategy Shard
2026/07/15
»
[Tech]
大规模训练工程(01):训练任务的状态解剖——显存账与 MFU
Anatomy of Training State: Memory Accounting and MFU
2026/07/13
»
[Tech]
大规模训练工程:从并行策略到容错恢复(总纲)
Large-Scale Training Engineering, from Parallelism to Fault Tolerance
June
2026/06/27
»
[Tech]
通信与互联(08):MoE 的通信——all-to-all、DeepEP 与 GPU 发起的通信
Communication for MoE: All-to-All, DeepEP and GPU-Initiated Networking
2026/06/24
»
[Tech]
通信与互联(07):推理侧的通信——custom all-reduce 与 KV 传输
Communication on the Inference Side: Custom All-Reduce and KV Cache Transfer
2026/06/20
»
[Tech]
通信与互联(06):nccl-tests、调优与排障——从带宽曲线到 hang
nccl-tests, Tuning and Debugging Hangs: From Bandwidth Curves to Flight Recorder
2026/06/17
»
[Tech]
通信与互联(05):PyTorch 的通信栈——ProcessGroupNCCL、stream 语义与计算通信重叠
PyTorch's Communication Stack: ProcessGroupNCCL, Stream Semantics, and Compute-Communication Overlap
2026/06/13
»
[Tech]
通信与互联(04):NCCL 架构——拓扑探测、channel、算法与协议
NCCL Architecture: Topology Detection, Channels, Algorithms and Protocols
2026/06/10
»
[Tech]
通信与互联(03):RDMA 与 GPUDirect——绕过 CPU 和主机内存的数据通路
RDMA and GPUDirect: Bypassing the CPU and Host Memory
2026/06/06
»
[Tech]
通信与互联(02):硬件互联——PCIe、NVLink、NVSwitch 与网络拓扑
Hardware Interconnect: PCIe, NVLink, NVSwitch and Network Topology
2026/06/03
»
[Tech]
通信与互联(01):集合通信原语与代价模型——α-β 模型与 ring all-reduce
Collective Communication Primitives and the Alpha-Beta Cost Model: Deriving Ring All-Reduce
2026/06/01
»
[Tech]
通信与互联:从 NCCL 到 RDMA(总纲)
Communication and Interconnect for AI-Infra, from NCCL to RDMA
May
2026/05/20
»
[Tech]
GPU Kernel 工程(10):剖析、测试与贡献——把 kernel 做成产品
Profiling, Testing and Contributing: Turning a Kernel into a Product
2026/05/19
»
[Tech]
GPU Kernel 工程(09):量化与融合 kernel——推理系统的其余部分
Quantized and Fused Kernels: The Rest of the Inference Stack
2026/05/18
»
[Tech]
GPU Kernel 工程(08):Attention Kernel——FlashAttention 与 PagedAttention
Attention Kernels: FlashAttention and PagedAttention from Derivation to Code
2026/05/17
»
[Tech]
GPU Kernel 工程(07):Triton——块级编程与编译器的边界
Triton: Block-Level Programming and Where the Compiler Stops
2026/05/16
»
[Tech]
GPU Kernel 工程(06):Tensor Core、CUTLASS 与 CuTe
Tensor Cores, CUTLASS and CuTe: Programming the Matrix Units
2026/05/15
»
[Tech]
GPU Kernel 工程(05):GEMM——从 naive 到分块
GEMM from Naive to Tiled: Reaching the Compute Ceiling on CUDA Cores
2026/05/14
»
[Tech]
GPU Kernel 工程(04):共享内存与 reduction——softmax、LayerNorm 与 online softmax
Shared Memory and Reductions: Softmax, LayerNorm and Online Softmax
2026/05/13
»
[Tech]
GPU Kernel 工程(03):访存合并与 elementwise kernel
Memory Coalescing and Elementwise Kernels: Hitting the Bandwidth Ceiling
2026/05/12
»
[Tech]
GPU Kernel 工程(02):CUDA 编程模型与第一个 kernel
The CUDA Programming Model and Your First Kernel, Measured
2026/05/11
»
[Tech]
GPU Kernel 工程(01):GPU 为什么这样设计——硬件结构与 Roofline
Why GPUs Look the Way They Do: Architecture and the Roofline Model
2026/05/10
»
[Tech]
GPU Kernel 工程:从 CUDA 执行模型到 FlashAttention(总纲)
GPU Kernel Engineering, from the CUDA Execution Model to FlashAttention
2026/05/09
»
[Tech]
多模态(07):自回归图像生成与统一模型
Autoregressive Image Generation and Unified Understanding-Generation Models
2026/05/08
»
[Tech]
多模态(06):Latent diffusion、DiT 与文生图配方
Latent Diffusion, DiT and How Text-to-Image Models Are Built
2026/05/07
»
[Tech]
多模态(05):扩散模型:DDPM、score matching 与 flow matching
Diffusion Models: DDPM, Score Matching and Flow Matching Are One Thing
2026/05/06
»
[Tech]
多模态(04):语音与全模态:音频编码器、codec 与全双工
Speech and Omni Models: Audio Encoders, Neural Codecs and Full-Duplex Dialogue
2026/05/05
»
[Tech]
多模态(03):VLM 的训练:数据、阶段与评测
Training a VLM: Data, Stages, Evaluation and Hallucination
2026/05/04
»
[Tech]
多模态(02):VLM 的结构:connector、注入方式与动态分辨率
VLM Architecture: Connectors, Injection Methods and Dynamic Resolution
2026/05/03
»
[Tech]
多模态(01):视觉编码器:CLIP、SigLIP 与自监督 ViT
Vision Encoders: CLIP, SigLIP and Self-Supervised ViTs
2026/05/02
»
[Tech]
多模态:从视觉编码器到扩散模型(总纲)
Multimodal Models: From Vision Encoders to Diffusion
2026/05/01
»
[Tech]
高效推理与压缩(06):剪枝、深度缩放与小模型配方
Pruning, Depth Scaling and How Small Models Are Made
April
2026/04/30
»
[Tech]
高效推理与压缩(05):KV cache 压缩:量化、驱逐与稀疏 attention
KV Cache Compression: Quantization, Eviction and Sparse Attention
2026/04/29
»
[Tech]
高效推理与压缩(04):量化感知训练、低比特与量化模型的评测
Quantization-Aware Training, Extreme Low-Bit and How to Evaluate a Quantized Model
2026/04/28
»
[Tech]
高效推理与压缩(03):训练后量化:误差模型、GPTQ、AWQ 与旋转
Post-Training Quantization: Error Models, GPTQ, AWQ and Rotation
2026/04/27
»
[Tech]
高效推理与压缩(02):投机解码:草稿、接受率与树
Speculative Decoding: Drafters, Acceptance Rates and Draft Trees
2026/04/26
»
[Tech]
高效推理与压缩(01):解码策略、采样与约束生成
Decoding Strategies: Sampling, Truncation, Penalties and Constrained Generation
2026/04/25
»
[Tech]
高效推理与压缩(算法侧):解码、投机、量化与 KV(总纲)
Efficient Inference and Model Compression: The Algorithm Side
2026/04/24
»
[Tech]
算法工程师的实验方法论:用有限的算力得出可信的结论
Experimental Methodology for AI Algorithm Engineers: Credible Conclusions on a Finite Compute Budget
2026/04/23
»
[Tech]
后训练(08):评测:benchmark、LLM-as-judge、Arena 与污染
Evaluating LLMs: Benchmarks, LLM-as-Judge, Arenas and Contamination
2026/04/22
»
[Tech]
后训练(07):蒸馏:logits 级、序列级与 on-policy
Knowledge Distillation for LLMs: Logit-Level, Sequence-Level and On-Policy
2026/04/21
»
[Tech]
后训练(06):Agent 与工具调用的 RL:多轮环境、轨迹数据与延后的奖励
Agentic RL: Multi-Turn Environments, Trajectory Data and Delayed Rewards
2026/04/20
»
[Tech]
后训练(05):推理模型与可验证奖励:R1 的配方、PRM 与 test-time compute
Reasoning Models and Verifiable Rewards: The R1 Recipe, Process Reward Models and Test-Time Compute
2026/04/19
»
[Tech]
后训练(04):离线 RL:从 RLHF 目标推出 DPO 及其变体
Offline Preference Optimization: Deriving DPO from the RLHF Objective, and Its Family
2026/04/18
»
[Tech]
后训练(03):在线 RL:PPO、GRPO 与 RLHF 三件套
Online RL for LLMs: PPO, GRPO and the Policy–Reward–Reference Trio
2026/04/17
»
[Tech]
后训练(02):偏好数据与奖励模型:Bradley-Terry、pairwise loss 与 reward hacking
Preference Data and Reward Models: Bradley-Terry, Pairwise Loss and Reward Hacking
2026/04/16
»
[Tech]
后训练(01):SFT:指令数据、chat template、loss mask 与参数高效微调
Supervised Fine-Tuning: Instruction Data, Chat Templates, Loss Masking and Parameter-Efficient Fine-Tuning
2026/04/15
»
[Tech]
后训练:从 SFT 到可验证奖励(总纲)
Post-Training: From Supervised Fine-Tuning to Verifiable Rewards
2026/04/13
»
[Tech]
预训练(04):训练配方与稳定性:学习率、batch、调度与 loss spike
Pretraining Recipes and Training Stability: Learning Rate, Batch Size, Schedules and Loss Spikes
2026/04/12
»
[Tech]
预训练(03):预训练数据工程:从 Common Crawl 到 15T token,去重、过滤与配比的账
Pretraining Data Engineering: From Common Crawl to 15T Tokens, the Arithmetic of Deduplication, Filtering and Mixing
2026/04/11
»
[Tech]
预训练(02):Scaling law:从 Chinchilla 到"过训练",算力怎么分给参数与数据
Scaling Laws: From Chinchilla to Over-Training, Splitting Compute between Parameters and Data
2026/04/10
»
[Tech]
预训练(01):分词与词表:BPE、词表大小与 token 效率
Tokenizers and Vocabulary: BPE, Vocabulary Size and Token Efficiency
2026/04/09
»
[Tech]
预训练:从 tokenizer 到训练配方(总纲)
Pretraining: Tokenizers, Scaling Laws, Data Pipelines and Training Recipes
2026/04/09
»
[Tech]
Transformer 与 LLM(08):多模态:vision encoder 的算量与 image token 的 KV 代价
Multimodal LLMs: The Cost of Vision Encoders and Image Tokens
2026/04/08
»
[Tech]
Transformer 与 LLM(07):量化、投机解码与 LoRA
Quantization, Speculative Decoding and LoRA: Three Ways to Reshape the Computation
2026/04/07
»
[Tech]
Transformer 与 LLM(06):浮点格式、数值稳定性与混合精度
Floating-Point Formats, Numerical Stability and Mixed Precision
2026/04/06
»
[Tech]
Transformer 与 LLM(05):MoE 的路由、激活参数量与通信形态
Mixture of Experts: Routing, Active Parameters and Communication Patterns
2026/04/05
»
[Tech]
Transformer 与 LLM(04):位置编码与长上下文
Positional Encoding and Long Context: RoPE Wavelengths, Extrapolation and Cost
2026/04/04
»
[Tech]
Transformer 与 LLM(03):Attention 变体与 KV cache
Attention Variants and the KV Cache: Deriving MHA, GQA, MQA and MLA
2026/04/03
»
[Tech]
Transformer 与 LLM(02):前向的算量与访存量
FLOPs, Bytes and Roofline: Prefill versus Decode
2026/04/02
»
[Tech]
Transformer 与 LLM(01):Transformer 解剖与参数量
Transformer Anatomy and Parameter Count: From config.json to 8.03B
2026/04/01
»
[Tech]
Transformer 与 LLM:结构、算量与数值(总纲)
Transformers and LLMs for Infrastructure Engineers: Architecture, Arithmetic and Numerics
March
2026/03/29
»
[Tech]
深度学习基础(06):RNN——从 LSTM 到 attention 的诞生
RNN: Backpropagation Through Time, the LSTM Gate as a Residual Path, and How Attention Was Born from the seq2seq Bottleneck
2026/03/28
»
[Tech]
深度学习基础(05):CNN——从 LeNet 到 ResNet,再到 ViT
CNN: Convolution as a Constrained Linear Layer, the ResNet Legacy and How ViT Turns Images into Tokens
2026/03/27
»
[Tech]
深度学习基础(04):正则化与泛化——为什么参数比样本多却不过拟合
Regularization and Generalization: Double Descent, Implicit Bias, Dropout, Weight Decay and When Overfitting Comes Back
2026/03/26
»
[Tech]
深度学习基础(03):优化器——从 SGD 到 AdamW 与学习率调度
Optimizers: From SGD to AdamW, Warmup, Clipping and the Batch-Learning-Rate Scaling Rule
2026/03/25
»
[Tech]
深度学习基础(02):训练为什么不稳定——初始化、归一化与残差
Why Deep Training Is Unstable: Variance Propagation, Initialization, Normalization and Residual Connections
2026/03/24
»
[Tech]
深度学习基础(01):反向传播——手推一个两层网络
Backpropagation by Hand: Shapes, the 2x Rule and Why Activations Must Be Saved
2026/03/23
»
[Tech]
深度学习基础:从反向传播到残差(总纲)
Deep Learning Foundations: From Backpropagation to Residual Connections
2026/03/05
»
[Tech]
LLM 时代的经典机器学习(06):评估——从混淆矩阵到 judge 的一致性
Evaluation: Confusion Matrix, Thresholds, AUC, Calibration, Paired Tests and Multiple Comparisons
2026/03/04
»
[Tech]
LLM 时代的经典机器学习(05):去重——MinHash 与 LSH 的概率
Deduplication: Jaccard, MinHash as an Unbiased Estimator, and the LSH S-Curve
2026/03/03
»
[Tech]
LLM 时代的经典机器学习(04):无监督——K-Means、PCA 与 embedding 聚类
Unsupervised Learning: K-Means, DBSCAN, PCA and What a Corpus Looks Like in Embedding Space
2026/03/02
»
[Tech]
LLM 时代的经典机器学习(03):分类器一家——从朴素贝叶斯到梯度提升
A Family of Classifiers: From Naive Bayes to Gradient Boosting, and Why Data Filters Use Small Models
2026/03/01
»
[Tech]
LLM 时代的经典机器学习(02):线性回归与逻辑回归——奖励模型的骨架
Linear and Logistic Regression: The Skeleton of Every Classification Head and Every Reward Model
February
2026/02/28
»
[Tech]
LLM 时代的经典机器学习(01):什么是学习——划分、泛化、过拟合与偏差-方差
What Is Learning: Splits, Generalization, Overfitting and the Bias-Variance Decomposition
2026/02/27
»
[Tech]
LLM 时代还要学经典机器学习吗:只讲它在哪里重现(总纲)
Classical Machine Learning in the LLM Era: What Survives and Where It Reappears
2026/02/26
»
[Tech]
PyTorch 深度实践(10):PyTorch 的工程体系——一次改动如何安全地到达用户
The Engineering System of PyTorch: How a Change Travels Safely from Commit to Production
2026/02/25
»
[Tech]
PyTorch 深度实践(09):分布式 PyTorch
Distributed Training in PyTorch: Collectives, DDP, FSDP, TP, PP, CP and EP
2026/02/24
»
[Tech]
PyTorch 深度实践(08):性能优化与调试
Performance Optimization and Debugging in PyTorch
2026/02/23
»
[Tech]
PyTorch 深度实践(07):编译执行与图优化
Compilation and Graph Optimization in PyTorch
2026/02/22
»
[Tech]
PyTorch 深度实践(06):C++ 扩展与自定义算子
C++ Extensions and Custom Operators in PyTorch
2026/02/21
»
[Tech]
PyTorch 深度实践(05):Dispatcher 与算子系统
The Dispatcher and Operator System in PyTorch
2026/02/20
»
[Tech]
PyTorch 深度实践(04):nn.Module 与训练系统
nn.Module and Training Systems in PyTorch
2026/02/19
»
[Tech]
PyTorch 深度实践(03):自动求导与动态计算图
Autograd and Dynamic Computation Graphs in PyTorch
2026/02/18
»
[Tech]
PyTorch 深度实践(02):Tensor 与内存布局
Tensor Abstraction and Memory Layout in PyTorch
2026/02/17
»
[Tech]
PyTorch 深度实践(01):PyTorch 整体介绍
PyTorch Overall Introduction
2026/02/16
»
[Tech]
PyTorch 深度实践:从 Tensor 到深度学习运行时(总纲)
Deep Dive into PyTorch, from Tensor to Deep Learning Runtime
2026/02/15
»
[Tech]
C++ 在 AI-Infra(08):构建、调试与测试工具链
Build, Debug and Test Toolchain
2026/02/14
»
[Tech]
C++ 在 AI-Infra(07):与 Python 之间——pybind11、Python C API 与 ABI
pybind11, the Python C API and ABI
2026/02/12
»
[Tech]
C++ 在 AI-Infra(06):并发、内存模型、TLS 与守卫
Concurrency, Memory Model, TLS and Guards
2026/02/10
»
[Tech]
C++ 在 AI-Infra(05):宏、静态注册与代码生成
Macros, Static Registration and Code Generation
2026/02/08
»
[Tech]
C++ 在 AI-Infra(04):多态与类型擦除——运行时如何选择实现
Polymorphism and Type Erasure
2026/02/07
»
[Tech]
C++ 在 AI-Infra(03):模板与泛型编程
Templates and Generic Programming
2026/02/06
»
[Tech]
C++ 在 AI-Infra(02):值、引用与所有权——对象模型与 RAII
Value Semantics, Ownership and RAII
2026/02/04
»
[Tech]
C++ 在 AI-Infra(01):从源码到二进制——编译模型与项目布局
Compilation Model and Project Layout
2026/02/02
»
[Tech]
C++ 在 AI-Infra:从对象模型到算子扩展(总纲)
C++ for AI-Infra, from the Object Model to Operator Extensions
January
2026/01/31
»
[Slides]
Java程序员的Python课
用 Java 思维快速理解 Python 的关键差异
2026/01/29
»
[Tech]
Python 在 AI-Infra(07):项目工程化与生产交付
Python Project Engineering and Production Delivery
2026/01/28
»
[Tech]
Python 在 AI-Infra(06):单元测试、问题定位与调试实践
Python Unit Testing, Troubleshooting, and Debugging
2026/01/27
»
[Tech]
Python 在 AI-Infra(05):内存管理与优化
Python Memory Management and Optimization
2026/01/26
»
[Tech]
Python 在 AI-Infra(04):Python的动态机制及工程实践
Python Dynamic Mechanisms and Practice
2026/01/25
»
[Tech]
Python 在 AI-Infra(03):并发、异步与任务协作
Python Concurrency, Asynchrony, and Task Collaboration in AI Systems
2026/01/24
»
[Tech]
Python 在 AI-Infra(02):类型系统与数据契约设计
Python Type System and Data Contract Design
2026/01/23
»
[Tech]
Python 在 AI-Infra(01):语言机制与运行时原理
Python Language Mechanisms and Runtime Internals
2026/01/22
»
[Tech]
Python 在 AI-Infra:从语言机制到生产交付(总纲)
Python for AI-Infra, from Language Mechanisms to Production Delivery (Overview)
2026/01/21
»
[Tech]
算法工程师的工具箱(05):GPU 直觉与实验管理——两个上限、四块显存、能复现
GPU Intuition and Experiment Management: Two Ceilings, Four Memory Buckets, and Reproducibility
2026/01/20
»
[Tech]
算法工程师的工具箱(04):Hugging Face 生态——六个库与一次 LoRA SFT 的组装
The Hugging Face Ecosystem: Six Libraries, a Six-Line LoRA SFT, and Why Reading the Source Is the Fastest Way to Learn
2026/01/19
»
[Tech]
算法工程师的工具箱(03):PyTorch 使用层(下)——混合精度、显存的账与多卡启用
PyTorch in Use, Part 2: Mixed Precision, the Memory Ledger and Turning On Multi-GPU
2026/01/18
»
[Tech]
算法工程师的工具箱(02):PyTorch 使用层(上)——五个对象与二十行训练循环
PyTorch in Use, Part 1: Five Objects and a Twenty-Line Training Loop
2026/01/17
»
[Tech]
算法工程师的工具箱(01):科学计算栈——NumPy 的形状直觉、Pandas 的错误分析、Matplotlib 的曲线
The Scientific Python Stack: Shapes and Broadcasting in NumPy, Error Analysis in Pandas, Reading Curves in Matplotlib
2026/01/16
»
[Tech]
算法工程师的工具箱:从一个想法到一次能跑的实验(总纲)
Tooling for AI Algorithm Engineers: NumPy, PyTorch, Hugging Face and the GPU in Your Head
2026/01/15
»
[Tech]
算法工程师的数学(08):统计推断与拟合——评测的置信区间与 scaling law
Statistical Inference and Curve Fitting: Confidence Intervals for Benchmarks and How Scaling Laws Are Fit
2026/01/14
»
[Tech]
算法工程师的数学(07):导数、梯度与链式法则——softmax 的梯度与策略梯度
Derivatives, Gradients and the Chain Rule: The Softmax Gradient and the Policy Gradient
2026/01/13
»
[Tech]
算法工程师的数学(06):熵、交叉熵与 KL——从困惑度到 DPO
Entropy, Cross-Entropy and KL Divergence: From Perplexity to the DPO Loss
2026/01/12
»
[Tech]
算法工程师的数学(05):从最大似然到交叉熵——第一个要会推的 loss
From Maximum Likelihood to Cross-Entropy: The First Loss You Should Be Able to Derive
2026/01/11
»
[Tech]
算法工程师的数学(04):概率入门——语言模型是一个条件分布
Probability Basics: A Language Model Is a Conditional Distribution
2026/01/10
»
[Tech]
算法工程师的数学(03):正交与旋转、特征值与 SVD——从 RoPE 到 LoRA
Orthogonal Matrices, Rotations, Eigenvalues and SVD: Why RoPE Encodes Relative Position and Why LoRA Works
2026/01/09
»
[Tech]
算法工程师的数学(02):内积、范数与余弦相似度
Inner Product, Norms and Cosine Similarity: One Language for Attention, Retrieval, Regularization and Quantization Error
2026/01/08
»
[Tech]
算法工程师的数学(01):向量、矩阵与形状——一个 token 过一层要算多少
Vectors, Matrices and Shapes: How Much Compute Does One Token Through One Layer Take
2026/01/07
»
[Tech]
算法工程师的数学:读公式不卡壳的最小集(总纲)
Mathematics for AI Algorithm Engineers: The Minimal Set to Read Papers and Derive Losses
2026/01/04
»
[Tech]
AI 应用工程师学习地图:在非确定性组件之上做可靠产品
A Learning Roadmap for AI Application Engineers
2026/01/03
»
[Tech]
AI-Infra 工程师学习地图:从后端工程师到基础设施贡献者
A Learning Roadmap for AI Infrastructure Engineers
2026/01/02
»
[Tech]
AI 算法工程师学习地图:从数学基础到大模型训练
A Learning Roadmap for AI Algorithm Engineers in the LLM Era
2026/01/01
»
[Tech]
AI 全栈学习地图:跑模型、造模型、用模型的三张图
One System, Three Roles — an Overview of the Three AI Learning Roadmaps
2024
July
2024/07/22
»
[Tech]
Python中如何定义POJO
2020
August
2020/08/01
»
[Tech]
测试驱动开发
2019
July
2019/07/28
»
[Tech]
用户行为串联方案
June
2019/06/01
»
[Tech]
参数服务器
May
2019/05/10
»
[Tech]
机器学习中的特征工程
April
2019/04/01
»
[Tech]
一个java大堆引发的『血案』
March
2019/03/12
»
[Tech]
使用CompletableFuture异步编程
January
2019/01/18
»
[Tech]
gitlab如何checkout某个group中的所有项目
2018
December
2018/12/26
»
[Tech]
创建Hadoop FileSystem报Provider org.apache.hadoop.fs.azure.NativeAzureFileSystem not a subtype异常
2018/12/19
»
[Tech]
git分支与maven版本之间的联动
2018/12/18
»
[Tech]
Git分支管理策略
November
2018/11/24
»
[Tech]
Spark数据倾斜及其解决方案
2018/11/20
»
[Tech]
Spark Executor内存管理
2018/11/12
»
[Tech]
Spark RDD
October
2018/10/30
»
[Tech]
设计模式分享
2018/10/24
»
[Tech]
log4j2如何动态的创建logger和appender
2018/10/20
»
[Tech]
Spark如何查看某个applicationId的executor日志
2018/10/19
»
[Tech]
Spark任务读取HDFS文件报Filesystem closed异常
2018/10/13
»
[Tech]
配置Nginx支持CORS的一个『坑』
2018/10/10
»
[Tech]
nginx proxy_pass 的一个『坑』
2018/10/08
»
[Tech]
maven如何deploy到多个repositories
July
2018/07/18
»
[Tech]
关于编码规范的一些建议
June
2018/06/03
»
[Tech]
AI基础架构:从大数据到深度学习
May
2018/05/21
»
[Tech]
微服务架构学习
April
2018/04/30
»
[Tech]
关于微服务架构
2018/04/08
»
[Tech]
Java各种锁介绍
February
2018/02/08
»
[Tech]
kubernetes初体验
January
2018/01/28
»
[Tech]
获取redis集群信息
2018/01/23
»
[Tech]
Redis集群学习
2018/01/23
»
[Tech]
ElasticSearch的节点类型
2018/01/10
»
[Tech]
ElasticSearch如何支持深度分页
2018/01/04
»
[Tech]
使用puppeteer和chrome-headless做暗网抓取
2017
December
2017/12/28
»
[Tech]
ElasticSearch如何支持嵌套属性检索
2017/12/24
»
[Tech]
ElasticSearch的Query Context和Filter Context
November
2017/11/23
»
[Life]
Thanksgiving in 2017
thanks, for everything you did
October
2017/10/28
»
[Tech]
如何查看和设置文件句柄数
August
2017/08/20
»
[Tech]
neo4j如何支持多个label索引查询
July
2017/07/03
»
[Tech]
一个诡异的Antlr4语法问题
June
2017/06/23
»
[Tech]
基于Aerospike实现一个分布式图数据库
May
2017/05/29
»
[Tech]
Aerospike UDF学习笔记
2017/05/20
»
[Life]
快乐课程
你真的会呼吸吗?
2017/05/16
»
[Tech]
markdown中图片如何指定大小
2017/05/15
»
[Life]
阿甘的网络日志
Hello world, hello my new blog
2017/05/05
»
[Tech]
redis slave的key过期机制
2017/05/04
»
[Tech]
Bloom filter在分布式环境中的应用
April
2017/04/23
»
[Tech]
neo4j如何实现存在就更新,否则插入?
2017/04/21
»
[Tech]
neo4j高效数据维护
2017/04/13
»
[Tech]
Titan的pluggable storage backend
2017/04/12
»
[Tech]
DynamoDB学习笔记
2017/04/10
»
[Tech]
neo4j如何批量导入JSON数据
2017/04/09
»
[Tech]
Aerospike学习笔记
March
2017/03/31
»
[Tech]
数据模型和存储系统
2017/03/24
»
[Tech]
Titan使用过程CPU超高问题排查
2017/03/22
»
[Tech]
Titan如何提供REST服务
2017/03/22
»
[Tech]
ArangoDB的索引学习
2017/03/21
»
[Tech]
neo4j学习笔记
2017/03/10
»
[Tech]
Kafka offset lag监控
2017/03/09
»
[Tech]
kafka broker间歇出现CLOSE_WAIT问题
February
2017/02/23
»
[Tech]
图存储引擎学习笔记
January
2017/01/25
»
[Tech]
抓取学习笔记
2017/01/16
»
[Tech]
使用supervisor进行进程监管
2016
December
2016/12/12
»
[Tech]
ElasticSearch存储相关
November
2016/11/26
»
[Tech]
卓有成效的程序员——Mac篇
2016/11/16
»
[Tech]
过载保护
October
2016/10/26
»
[Tech]
Git学习笔记
2016/10/17
»
[Tech]
搜索引擎中的相关性和排序截断
2016/10/17
»
[Tech]
protobuf中的反射
2016/10/17
»
[Tech]
巧用protobuf的自定义options
2016/10/11
»
[Tech]
Protobuf Buffer的缺陷
July
2016/07/16
»
[Tech]
记一个诡异的C++问题
May
2016/05/28
»
[Tech]
互联网广告系统学习笔记
2016/05/05
»
[Tech]
Kerberos学习笔记
April
2016/04/18
»
[Tech]
使用logstash收集nginx访问日志
2016/04/11
»
[Tech]
Kafka实战
2016/04/08
»
[Tech]
安装RabbitMQ
2016/04/08
»
[Tech]
大数据平台学习笔记
March
2016/03/30
»
[Tech]
RAID学习
2016/03/29
»
[Tech]
自动化部署平台设计
2016/03/21
»
[Tech]
keepalived实战
2016/03/12
»
[Tech]
MySQL高可用性方案
2016/03/06
»
[Tech]
logback学习笔记
February
2016/02/19
»
[Tech]
Spring Java-based配置
2016/02/03
»
[Tech]
如何实现一个配置中心
2016/02/01
»
[Tech]
如何单元测试二方库
2015
December
2015/12/07
»
[Tech]
JS跨域问题及解决方案
November
2015/11/16
»
[Tech]
Java8时间处理
October
2015/10/25
»
[Tech]
input too large for RSA cipher
2015/10/23
»
[Tech]
MIME和编码学习笔记
September
2015/09/24
»
[Tech]
分布式文件系统选型和预研
2015/09/16
»
[Tech]
MySQL主从同步学习
2015/09/08
»
[Tech]
Quartz的misfire机制
August
2015/08/27
»
[Tech]
如何让一个Quartz实例不执行任务
2015/08/21
»
[Tech]
怎样获取form-data方式POST的数据
2015/08/11
»
[Tech]
高可用分布式缓存系统
2015/08/11
»
[Tech]
记一次Redis错误排查经历
July
2015/07/30
»
[Tech]
Spring的Bean生命周期和扩展点
2015/07/24
»
[Tech]
Tomcat调优
2015/07/22
»
[Tech]
记一次MySQL主从同步错误处理
2015/07/03
»
[Tech]
Metric监控系统
June
2015/06/11
»
[Tech]
InfluxDB安装和使用
2015/06/10
»
[Tech]
安装OpenTSDB
2015/06/08
»
[Tech]
java服务端监控平台设计
2015/06/06
»
[Tech]
Java Attach API
2015/06/04
»
[Tech]
JMX学习
May
2015/05/28
»
[Tech]
tomcat监控
2015/05/26
»
[Tech]
MySQL主从同步失败
2015/05/15
»
[Tech]
如何限制某个IP对MySQL的访问
March
2015/03/27
»
[Life]
走出象牙塔
给即将踏入社会的师弟师妹们的一些建议
2015/03/09
»
[Tech]
如何自定义Spring XML Bean配置
February
2015/02/07
»
[Tech]
任务调度框架设计和实现
2015/02/06
»
[Tech]
java standalone模板
January
2015/01/27
»
[Tech]
java动态代理和动态类加载
2015/01/24
»
[Tech]
nginx日志格式
2015/01/23
»
[Tech]
Java中如何正确的加载配置文件
2015/01/18
»
[Tech]
Java DNS查询内部实现
2015/01/18
»
[Tech]
如何让java程序优先使用自定义的DNS nameserver
2015/01/16
»
[Tech]
Redis的事务
2015/01/04
»
[Tech]
nginx URL rewrite自动增加请求参数问题
2014
December
2014/12/23
»
[Tech]
应用如何记录集中日志
2014/12/17
»
[Tech]
kibana学习笔记
2014/12/17
»
[Tech]
ElasticSearch如何实现按天翻滚索引
2014/12/16
»
[Tech]
Spring MVC的异常处理机制
2014/12/16
»
[Tech]
如何防止表单重复提交
2014/12/16
»
[Tech]
配置tomcat的access_log
2014/12/13
»
[Tech]
分布式RPC框架如何进行服务寻址和分发
2014/12/11
»
[Tech]
Quartz突然停止执行问题
2014/12/11
»
[Tech]
Java文件读取支持timeout
2014/12/03
»
[Tech]
nginx日志自动按天分隔
2014/12/03
»
[Tech]
如何解决系统盘爆满问题
November
2014/11/17
»
[Tech]
JPA的事务管理器配置
2014/11/16
»
[Tech]
移动终端设备唯一标识
2014/11/06
»
[Tech]
Quartz工作机制
October
2014/10/29
»
[Tech]
日志监控系统
2014/10/22
»
[Tech]
MySQL用户授权
2014/10/20
»
[Tech]
如何构建maven私有仓库
2014/10/14
»
[Tech]
实时消息系统设计与实现
2014/10/10
»
[Tech]
nginx重定向问题
September
2014/09/28
»
[Tech]
Reading搜索
2014/09/26
»
[Tech]
ElasticSearch学习
2014/09/25
»
[Tech]
ElasticSearch的mappings
2014/09/24
»
[Tech]
ElasticSearch字段排序
2014/09/24
»
[Tech]
ElasticSearch的数据类型
2014/09/24
»
[Tech]
ElasticSearch的Analyzer
August
2014/08/16
»
[Tech]
如何提高服务器并发处理能力
2014/08/15
»
[Tech]
select、poll和epoll简介
2014/08/13
»
[Tech]
静态资源服务器迁移MFS方案
2014/08/10
»
[Tech]
服务器编程模型
2014/08/10
»
[Tech]
Java NIO.2
2014/08/06
»
[Tech]
nginx URL rewrite与下载文件名称问题
2014/08/05
»
[Tech]
Java NIO
2014/08/04
»
[Tech]
如何实时同步大量小文件
July
2014/07/30
»
[Tech]
服务端监控方案
2014/07/27
»
[Tech]
如何监控线上应用的运行状态
2014/07/24
»
[Tech]
配置MySQL Slave
2014/07/24
»
[Tech]
使用Redis做简单的消息队列
2014/07/11
»
[Tech]
Spring MVC国际化和本地化
June
2014/06/19
»
[Tech]
log4j详细介绍
2014/06/18
»
[Tech]
log4j日志路径问题
2014/06/17
»
[Tech]
使用拦截器做简单的性能监控
2014/06/05
»
[Tech]
Andriod平台推送系统设计
May
2014/05/23
»
[Tech]
Spring各种依赖注入注解的区别
2014/05/19
»
[Tech]
一个简单分页查询组件实现
April
2014/04/30
»
[Tech]
优雅的Builder模式
2014/04/28
»
[Tech]
动态页面缓存方案
March
2014/03/21
»
[Tech]
小议Android手机隐私数据的存储安全
2014/03/12
»
[Tech]
使用zookeeper实现分布式锁
2014/03/11
»
[Tech]
ZooKeeper简介
2014/03/06
»
[Tech]
开放平台鉴权以及OAuth2.0介绍
2014/03/05
»
[Tech]
闭包
2014/03/03
»
[Tech]
JVM类加载器与ClassNotFoundException和NoClassDefFoundError
2014/03/02
»
[Tech]
Java虚拟机学习笔记
February
2014/02/26
»
[Tech]
海量服务之——灰度发布
2014/02/25
»
[Tech]
HTTPS原理
2014/02/20
»
[Tech]
巧用TheadLocal
2014/02/14
»
[Tech]
SLA和QoS在RPC/OpenAPI容器中的作用
2014/02/10
»
[Tech]
通讯协议序列化思考
2014/02/10
»
[Tech]
分布式系统常用思想和技术总结
2014/02/02
»
[Tech]
Config Server和SLA在RPC中的作用
January
2014/01/26
»
[Tech]
一个有意思的Scala函数式编程例子
2014/01/13
»
[Tech]
网络RPC编码协议学习
2014/01/07
»
[Tech]
Thrift的序列化版本控制
2014/01/02
»
[Tech]
tips for sublime text
2013
December
2013/12/31
»
[Tech]
BTrace实战
2013/12/27
»
[Tech]
循环引用序列化问题
2013/12/22
»
[Tech]
使用EC2和SSH翻墙
2013/12/20
»
[Tech]
mina学习笔记
2013/12/19
»
[Tech]
多次编解码导致的的奇怪乱码问题
2013/12/10
»
[Tech]
一个简单的性能优化和防止DOS攻击的示例
2013/12/06
»
[Tech]
用Photoshop磨皮
2013/12/04
»
[Tech]
如何实现用户认证授权系统
November
2013/11/28
»
[Tech]
spring AOP internal
2013/11/26
»
[Tech]
使用curl和wget模拟REST请求
2013/11/26
»
[Tech]
return async result in java
2013/11/26
»
[Tech]
Java并发学习笔记
2013/11/25
»
[Tech]
高并发下额度限制问题
2013/11/21
»
[Tech]
RESTful CRUD with spring-mvc-and-bootstrap
2013/11/18
»
[Tech]
如何让tomcat不解压你的war包
2013/11/18
»
[Tech]
Rest Response and Exception
2013/11/04
»
[Tech]
如何解决time_wait状态占用端口问题
October
2013/10/08
»
[Tech]
开放平台主动推送预言
September
2013/09/21
»
[Tech]
content negotiation using spring mvc
2013/09/20
»
[Tech]
面试点
August
2013/08/20
»
[Tech]
scala with maven
2013/08/10
»
[Tech]
Scala的Option、Some和None
2013/08/02
»
[Tech]
C# Url Encoding的一些问题
July
2013/07/26
»
[Tech]
移除SVN锁
2013/07/23
»
[Tech]
卓有成效的程序员——windows篇
2013/07/20
»
[Tech]
JVM编码
2013/07/15
»
[Tech]
maven学习笔记
2013/07/10
»
[Tech]
如何在系统启动时完成资源加载
2013/07/09
»
[Tech]
XSS注入防御
2013/07/09
»
[Tech]
CSRF防御
2013/07/06
»
[Tech]
Spring使用@value annotation注入property变量和环境变量
2013/07/02
»
[Tech]
MySQL字符串比较大小写问题
2013/07/01
»
[Tech]
Java Heap OOM问题
June
2013/06/28
»
[Tech]
配置文件串串SHOW
February
2013/02/24
»
[Tech]
负载均衡
January
2013/01/30
»
[Tech]
如何使用tomcat高效调试
2013/01/24
»
[Tech]
如何在远程Linux机器上执行Shell命令
2013/01/24
»
[Tech]
域名解析过程及其相关配置
2013/01/21
»
[Tech]
使用Spring-Security进行登录控制的session问题
2013/01/18
»
[Tech]
Restful Spring MVC
2013/01/11
»
[Tech]
Spring的Bean Scopes
2013/01/11
»
[Tech]
Spring的Bean Scopes实现机制源码剖析
2013/01/07
»
[Tech]
JUnit与Spring的整合——JUnit中的TestCase如何拥有spring的事务管理机制
2013/01/06
»
[Tech]
Spring与web MVC的整合——Spring的应用上下文管理
2013/01/06
»
[Tech]
创建可执行的jar包
2013/01/02
»
[Tech]
JUnit与Spring的整合——JUnit的TestCase如何自动注入Spring容器托管的对象
2012
December
2012/12/30
»
[Tech]
Quartz与Spring的整合-使用Spring的FactoryBean实现动态Properties
2012/12/29
»
[Tech]
Quartz与Spring的整合-Quartz中的job如何自动注入spring容器托管的对象
November
2012/11/15
»
[Tech]
Linux下如何备份旧文件
2012/11/02
»
[Tech]
maven的resources插件
October
2012/10/12
»
[Tech]
Spring事务配置
September
2012/09/11
»
[Tech]
如何不刷新页面上传文件
2012/09/11
»
[Tech]
如何往HttpServletRequest中塞请求参数
August
2012/08/16
»
[Tech]
URL encoding学习笔记
March
2012/03/29
»
[Tech]
海量图片存储思考
2012/03/05
»
[Tech]
Using Fabric to Type Less
2012/03/05
»
[Tech]
尽量用英文写博客
2012/03/04
»
[Tech]
构建可伸缩的大型网站
February
2012/02/28
»
[Tech]
工欲善其事,必先利其器——从零打造你的vim
2012/02/21
»
[Tech]
使用github搭建个人博客
2012/02/15
»
[Tech]
使用rsync进行文件同步
2012/02/09
»
[Tech]
shell模块的另一种组织方式
2012/02/02
»
[Tech]
Linux里复制终端Session(像SecureCRT一样)
2011
September
2011/09/06
»
[Tech]
shell如何模块化和复用——shell深入学习
2011/09/03
»
[Tech]
Linux命令学习之——paste命令
August
2011/08/14
»
[Tech]
python2.x的一个需要注意的地方
2011/08/12
»
[Tech]
ifttt模式语言——sed和awk深入学习
2011/08/10
»
[Tech]
sort和uniq tips
2011/08/10
»
[Tech]
AWK学习与实战
July
2011/07/13
»
[Tech]
Maven快速入门
May
2011/05/26
»
[Tech]
Linux命令学习之——cut命令
2011/05/11
»
[Tech]
shell如何实现ssh免密码登陆
2011/05/08
»
[Tech]
shell语言之我见
March
2011/03/31
»
[Tech]
关于文件描述符和句柄
2011/03/21
»
[Tech]
TCP与UDP的区别
January
2011/01/21
»
[Tech]
sed实战
2011/01/20
»
[Tech]
关于面向对象与面向过程的一些思考
2010
December
2010/12/26
»
[Tech]
语言进化论
2010/12/20
»
[Tech]
关于接口设计的一些思考
November
2010/11/27
»
[Tech]
-exec和xargs的区别
June
2010/06/05
»
[Tech]
Java网络IO编程
May
2010/05/26
»
[Tech]
从面向过程到面向对象——在C中如何实现面向对象编程
2010/05/06
»
[Tech]
进程VS线程
2009
July
2009/07/06
»
[Tech]
从暴风影音事件反思DNS频率攻击漏洞
June
2009/06/18
»
[Tech]
使用Servlet和JSP模拟最小化的SpringMVC框架
May
2009/05/27
»
[Tech]
Groovy元编程——使用invokeMethod和闭包构建DSL和Builder
2009/05/23
»
[Tech]
BDB中的共享区域
2009/05/23
»
[Tech]
BDB中的共享区域──MPOOL
2009/05/07
»
[Tech]
BDB1.6中的初始化过程──MPOOL初始化
2009/05/06
»
[Tech]
BDB事务共享区域
2009/05/06
»
[Tech]
MPOOL共享内存
2009/05/06
»
[Tech]
BDB日志共享区域
2009/05/06
»
[Tech]
BDB锁共享区域
2008
December
2008/12/09
»
[Tech]
如何确保C库可以正确被C++客户端程序调用
ABOUT ME
天造之才,皆有其用。
振翅高飞,无须在梦中。
微信公众号
知
HOT TAGS
BDB
7
database
15
transaction
4
Spring MVC
5
Java
29
shell
6
linux
7
productivity
3
jekyll
4
博客
3
spring
15
maven
6
tomcat
3
JVM
3
mysql
8
缓存
3
log4j
3
redis
7
消息队列
4
nginx
7
elasticsearch
12
生活
4
性能优化
3
分布式
4
kafka
3
git
3
高并发
3
图数据库
8
neo4j
5
aerospike
3
Kubernetes
11
架构
3
AI
157
spark
6
敏捷
3
Python
15
AI-Infra
96
LLM
64
Agent
4
Roadmap
4
Math
9
PyTorch
22
C++
9
Machine Learning
8
Deep Learning
7
Transformer
14
Pretraining
5
Post-Training
9
RLHF
7
Evaluation
3
Inference
7
Quantization
3
Multimodal
8
Diffusion
3
CUDA
13
Triton
11
GPU
29
NCCL
10
RDMA
10
Megatron
11
DeepSpeed
6
Distributed Training
14
torchtitan
7
Observability
3
大模型推理
15
RL
9
verl
9
vLLM
11
RECOMMEND
美团技术团队
国内少有的持续高质量的一线工程实践,尤其是分布式与搜索推荐
Jay Alammar
用图讲透 Transformer / LLM 的第一人,The Illustrated 系列
Eugene Yan
机器学习系统与 LLM 应用落地的长文,工程视角
deeplearning.ai
吴恩达的课程与 The Batch 周报,跟进 AI 进展的低成本方式
×