<?xml version="1.0" encoding="utf-8" standalone="yes"?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom" xmlns:content="http://purl.org/rss/1.0/modules/content/">
  <channel>
    <title>Dapo on Njx&#39;Log</title>
    <link>https://njx-njx.github.io/tags/dapo/</link>
    <description>Recent content in Dapo on Njx&#39;Log</description>
    <image>
      <title>Njx&#39;Log</title>
      <url>https://njx-njx.github.io/blog-cover.jpg</url>
      <link>https://njx-njx.github.io/blog-cover.jpg</link>
    </image>
    <generator>Hugo -- 0.161.1</generator>
    <language>zh</language>
    <lastBuildDate>Sun, 30 Aug 2026 12:05:00 +0800</lastBuildDate>
    <atom:link href="https://njx-njx.github.io/tags/dapo/index.xml" rel="self" type="application/rss+xml" />
    <item>
      <title>截断的回答该怎么判：DAPO Overlong 在修什么</title>
      <link>https://njx-njx.github.io/posts/dapo-overlong-reward-shaping/</link>
      <pubDate>Sun, 30 Aug 2026 12:05:00 +0800</pubDate>
      <guid>https://njx-njx.github.io/posts/dapo-overlong-reward-shaping/</guid>
      <description>RL 训练必须有生成长度上限，被截断的回答默认按答错处理——这会把“无法判定”的样本错判成“错”样本，把奖励噪声直接注入组内 advantage。DAPO 的两个方案：Overlong Filtering 承认不可判定、屏蔽其 loss；Soft Overlong Punishment 在悬崖前修一段线性缓冲坡，让长度控制变成连续信号。本文拆解机制、公式、配置数值，并把它和 Dr.GRPO、token-mean 拼成完整的长度问题图景。</description>
    </item>
    <item>
      <title>一票还是十票：token-mean 与 sample-mean 在争什么</title>
      <link>https://njx-njx.github.io/posts/token-mean-vs-sample-mean/</link>
      <pubDate>Sun, 30 Aug 2026 12:05:00 +0800</pubDate>
      <guid>https://njx-njx.github.io/posts/token-mean-vs-sample-mean/</guid>
      <description>同一个 batch 里，500 token 的回答和 5000 token 的回答，谁对梯度的贡献大？GRPO 说一样大，DAPO 说长的大十倍，两家都能给出理由。这是损失聚合方式之争——sample-mean 每条回答一票，token-mean 每个 token 一票。DAPO 论文 §3.3 说长 CoT 场景必须换 token 级，但换完会把长度倾斜从 response 级赶到 question 级。本文拆解三个公式、两家的论证和一个少有人讲的 tradeoff。</description>
    </item>
    <item>
      <title>一道题采几条回答：GRPO 组大小 G 买来了什么</title>
      <link>https://njx-njx.github.io/posts/grpo-group-size-tradeoffs/</link>
      <pubDate>Sat, 29 Aug 2026 12:00:00 +0800</pubDate>
      <guid>https://njx-njx.github.io/posts/grpo-group-size-tradeoffs/</guid>
      <description>DeepSeekMath 每题采 64 条回答，教学实现默认 8 条，最新论文 2 条也声称够用——同一个旋钮差出 32 倍。G 是 GRPO 唯一的 baseline 来源和最大的算力开关，它到底买来了什么？本文拆解 G 的三重身份：控制变量的样本数、有效梯度组的覆盖率、std 归一化的分母质量，以及动态采样为什么能用过滤代替加大 G。</description>
    </item>
    <item>
      <title>熵是怎么塌掉的：DAPO 的 Clip-Higher 在修什么</title>
      <link>https://njx-njx.github.io/posts/dapo-clip-higher-entropy-collapse/</link>
      <pubDate>Wed, 26 Aug 2026 14:30:00 +0800</pubDate>
      <guid>https://njx-njx.github.io/posts/dapo-clip-higher-entropy-collapse/</guid>
      <description>naive GRPO 复现 R1 时熵快速坍缩、组内样本趋同，DAPO 的 Clip-Higher 只把裁剪上界从 0.2 抬到 0.28。这一个数背后，是 PPO-Clip 在概率轴上的结构性不对称。本文拆解失效链条、机制、证据与边界。</description>
    </item>
  </channel>
</rss>
