WikiSkill: Operational Memory for Agents

WikiSkill’s strongest idea is the separation of raw traces, persistent operational memory, and executable skills, with validation-gated rollback only for the executable layer.

August 31, 2026 · Sai Boorlagadda

Beyond Static Leaderboards: Predictive Validity for the Evaluation of LLM Agents

The paper correctly prioritizes out-of-distribution rank transfer, but its proposed score conflates predictive validity with risk-adjusted utility.

June 21, 2026 · Sai Boorlagadda

From Static Templates to Dynamic Runtime Graphs

The paper offers a useful template/realized-graph/trace distinction and reporting protocol, but lacks a reproducible survey methodology.

June 21, 2026 · Sai Boorlagadda