Jiaxing Ni
19, post-training researcher @ ByteDance. CS at Beijing Institute of Technology, Xu Teli Elite Class.
I work where harness, eval, and post-training meet — the engineering side of modern LLM training. Good systems feel quiet on the surface. Underneath, there are rollout pipelines, reward traces, eval slices, and a few scripts watching the night shift.
Now
Post-train researcher @ ByteDance. Current research threads:
- A full harness–eval–post-train rollout system
- Interpretable generalization in evals
- High-quality SFT data synthesis
- Agentic RL environment construction
Before
- Moonshot AI (Kimi) — researcher: harness (swarm, Kimi work memory), post-training (rollout system), eval (science / multi-turn / single-turn).
- paperboy — agent engineer: a context OS built on user-context modeling.
- ByteDance Xpert — expert data pipeline and translation bench.
- Maxagent — eval research: deterministic evidence extraction from transcripts, and why an eval generalizes.
- PKU Wangxuan Institute — research intern on fine-grained sports video understanding: 3D pose reconstruction, action localization.
- Earlier fragments: quadruped robot simulation training and vision models (Baidu Software Cup), PaddlePaddle / PaddleDetection modules, and a series of agent architectures — AI notes, research reading assistant, code generation, video, companion.
Coordinates
- Identity: researcher, builder, engineer.
- Base: Beijing / Shanghai.
- Current focus: rollout systems, post-training recipes, evals and benchmarks, reasoning RL and agentic RL.
- Taste: prototype quickly, measure honestly, keep only what survives contact with real edges.
Toolkit
- Languages: Python, C++.
- ML: PyTorch, PaddlePaddle, Transformers, DeepSpeed.
- Infra: Docker, Git, distributed training/inference.
- Agents: harness design, tool loops, context engineering, eval pipelines.
Tools are not ornaments here. A tool earns its place when the next experiment becomes less vague.
About This Blog
A field notebook on the mechanisms behind LLM training: what a knob actually does, what breaks when you turn it, and what the metric said.
Topics orbit harness and rollout systems, model architecture, mid-training and post-training, evals and benchmarks, reasoning RL and agentic RL, RSI, and data synthesis. Longer posts often come with an interactive report edition — see Reports.
Built with Hugo and PaperMod, hosted on GitHub Pages, and open source on GitHub.