Product Evals: A Travel Planner Taste Test

A product note on evaluating an AI travel planner: itinerary quality, OOD scenes, long-context consistency, recommendation taste, and user loops.

五月 22, 2026 · 4 分钟 · 664 字 · njx

Context Engineering: The Shape of Evidence

A note on RAG and context engineering: retrieval quality, evidence shape, memory boundaries, and why context is a product surface.

五月 22, 2026 · 3 分钟 · 637 字 · njx

Agentic RL: The Long Shadow of Feedback

A field note on agentic RL: reward design, behavior shaping, credit assignment, online and offline evaluation, and the feedback loops behind agents.

五月 22, 2026 · 4 分钟 · 749 字 · njx

Evals As Instruments: What Demos Hide

A note on evaluation as an instrument: failure cases, metrics, benchmark design, product loops, and the discipline of measuring agents.

五月 22, 2026 · 4 分钟 · 729 字 · njx

Agent Design: Loops, Tools, and Memory

A field note on designing agents as observable loops, with tools, memory, failure recovery, and product boundaries.

五月 22, 2026 · 4 分钟 · 747 字 · njx