arbor
Autonomously improve a real artifact (code, training recipe, agent harness, data pipeline, prompt) against an objective and an evaluator, using Hypothesis Tree Refinement (HTR) from the Arbor paper. Use this whenever someone wants to iteratively optimize something over many experiments without overfitting — e.g. "get my model's eval score up", "improve this agent/harness", "tune this pipeline", "beat the baseline on this benchmark", "run a search over approaches and keep the best", "do an MLE-bench / Kaggle-style optimization", or any long-horizon "make this artifact better and don't just memorize the dev set" task. Trigger it even when the user doesn't say "Arbor" or "hypothesis tree" but describes repeated experiment-and-evaluate loops, branching exploration of competing ideas, or worries about a dev/test gap. Runs Claude itself as the coordinator with subagent executors in isolated git worktrees; for the standalone `arbor` CLI tool see references/arbor-upstream.md.
适合你,如果需要反复实验迭代提升评估指标而不在测试集上过拟合
npx oh-my-skill add k-dense-ai/scientific-agent-skills/arborcurl -fsSL https://oh-my-skill.com/install.sh | bash -s -- k-dense-ai/scientific-agent-skills/arbornpx oh-my-skill verify k-dense-ai/scientific-agent-skills/arbor怎么用
商店整理自技能原文 · 版本 ab2f84a · 表述以原文为准装上后,Claude 会变成自动实验员:根据你给的任务、评估标准和现有代码或配置,自己设计并执行多轮实验,记录每次尝试和结果,最后给出一个优化后的版本。
当你想要反复改进某个代码或模型,并有一个评估分数作为目标时触发,例如想提高测试分数或调优参数。