<?xml version="1.0" encoding="utf-8" standalone="yes"?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom" xmlns:content="http://purl.org/rss/1.0/modules/content/">
  <channel>
    <title>Benchmark on CYR&#39;ML</title>
    <link>https://chengyongru.github.io/tags/benchmark/</link>
    <description>Recent content in Benchmark on CYR&#39;ML</description>
    <generator>Hugo</generator>
    <language>en-us</language>
    <lastBuildDate>Mon, 15 Jun 2026 10:30:20 +0800</lastBuildDate>
    <atom:link href="https://chengyongru.github.io/tags/benchmark/index.xml" rel="self" type="application/rss+xml" />
    <item>
      <title>关于 Claw-SWE-Bench</title>
      <link>https://chengyongru.github.io/blog/notebook/%E5%85%B3%E4%BA%8Eclaw%20bench/</link>
      <pubDate>Fri, 12 Jun 2026 18:03:38 +0800</pubDate>
      <guid>https://chengyongru.github.io/blog/notebook/%E5%85%B3%E4%BA%8Eclaw%20bench/</guid>
      <description>&lt;p&gt;最近沉迷 agent，所以对相关论文和 bench 比较感兴趣。下面这篇是我最近看到的比较炸裂的一篇。&lt;/p&gt;
&lt;p&gt;&lt;a href=&#34;https://arxiv.org/html/2606.12344v1&#34;&gt;https://arxiv.org/html/2606.12344v1&lt;/a&gt;
&lt;a href=&#34;https://github.com/opensquilla/claw-swe-bench&#34;&gt;https://github.com/opensquilla/claw-swe-bench&lt;/a&gt;&lt;/p&gt;
&lt;p&gt;我的结论很简单：这篇可以看看 bench 设计思路，但是实验结论基本没什么意义。原因也很直接：论文把单次运行、混杂对比和不完整实验网格写成了看起来很确定的结论。作者明知每个组合只跑一次，还在正文里讨论排名、机制和框架优劣，多多少少有点让我脑内自动浮现派大星翻白眼流哈喇子表情包。&lt;/p&gt;</description>
    </item>
  </channel>
</rss>
