【摘要】创作型 Agent 的评测不能只盯固定 benchmark。更可靠的做法,是用 Scenario Model 每轮生成样本外任务,再让 baseline 和 candidate 在同一批新题上做配对比较 阅读全文