解剖极简 SFT 训练器:每个决策,教科书和框架怎么说

TRL 的 SFTTrainer 一行就能训,为什么还有人手写 SFT 循环?我把 forge 仓库 Stage 1 的每个决策(chat template、prompt masking、collate 与 loss、bf16、LR schedule、数据、验证)逐一摆到台面上,对照两套教材(rlhf-book、hands-on-modern-rl)、四个生产框架(alignment-handbook、trl、open-instruct、OpenRLHF)和六路信息源:哪些决策有共识,哪些有真实的配方之争,哪些两本教材各自糊弄过去了。

八月 28, 2026 · 13 分钟 · 6418 字 · njx

Context Engineering: The Shape of Evidence

A note on RAG and context engineering: retrieval quality, evidence shape, memory boundaries, and why context is a product surface.

五月 22, 2026 · 3 分钟 · 637 字 · njx

Evals As Instruments: What Demos Hide

A note on evaluation as an instrument: failure cases, metrics, benchmark design, product loops, and the discipline of measuring agents.

五月 22, 2026 · 4 分钟 · 729 字 · njx

Agent Design: Loops, Tools, and Memory

A field note on designing agents as observable loops, with tools, memory, failure recovery, and product boundaries.

五月 22, 2026 · 4 分钟 · 747 字 · njx