<?xml version="1.0" encoding="utf-8" standalone="yes"?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom" xmlns:content="http://purl.org/rss/1.0/modules/content/">
  <channel>
    <title>Precision on Njx&#39;Log</title>
    <link>https://njx-njx.github.io/tags/precision/</link>
    <description>Recent content in Precision on Njx&#39;Log</description>
    <image>
      <title>Njx&#39;Log</title>
      <url>https://njx-njx.github.io/blog-cover.jpg</url>
      <link>https://njx-njx.github.io/blog-cover.jpg</link>
    </image>
    <generator>Hugo -- 0.161.1</generator>
    <language>zh</language>
    <lastBuildDate>Thu, 27 Aug 2026 13:00:00 +0800</lastBuildDate>
    <atom:link href="https://njx-njx.github.io/tags/precision/index.xml" rel="self" type="application/rss+xml" />
    <item>
      <title>同一个模型，两个概率：训练推理不一致的代价</title>
      <link>https://njx-njx.github.io/posts/tis-train-inference-mismatch/</link>
      <pubDate>Thu, 27 Aug 2026 13:00:00 +0800</pubDate>
      <guid>https://njx-njx.github.io/posts/tis-train-inference-mismatch/</guid>
      <description>rollout 用推理引擎、训练用训练引擎，同一个模型在两边的数值对不上：实际采样的分布与你假设的策略存在偏差，名义上的 on-policy 悄悄变成 off-policy。本文拆解差异来源、有偏梯度与部署差距两个后果、算法补丁谱系，以及 Sea AI Lab 归到 BF16 精度上的根源结论。</description>
    </item>
  </channel>
</rss>
