Hi, thanks for creating ResearchClawBench! I'm considering doing a submission soon.
As I was reviewing the runs from the public leaderboard, I noticed that the same sub-task may be evaluated as Mode A (Objective Evaluation) in some cases and Mode B (Subjective Evaluation) in other cases. This means the judge may be applying inconsistent scoring rubrics.
For example, for the Physics 000 task, if we compare the public ResearchHarness runs for Claude Opus 4.6 and 4.7, we see:
- For ResearchHarness (Claude-Opus-4.6): the judge determines Result 1 as Mode B, Result 2 as Mode A, and Result 3 as Mode A.
- But for ResearchHarness (Claude-Opus-4.7): the judge determines Result 1 as Mode A, Result 2 as Mode A, and Result 3 as Mode B.
Specifically, Results 1 and 3 are evaluated with different modes across the two runs.
Would you consider fixing the evaluation mode for each subtask beforehand? This could eliminate an additional source of noise in the evaluation.
Hi, thanks for creating ResearchClawBench! I'm considering doing a submission soon.
As I was reviewing the runs from the public leaderboard, I noticed that the same sub-task may be evaluated as Mode A (Objective Evaluation) in some cases and Mode B (Subjective Evaluation) in other cases. This means the judge may be applying inconsistent scoring rubrics.
For example, for the Physics 000 task, if we compare the public ResearchHarness runs for Claude Opus 4.6 and 4.7, we see:
Specifically, Results 1 and 3 are evaluated with different modes across the two runs.
Would you consider fixing the evaluation mode for each subtask beforehand? This could eliminate an additional source of noise in the evaluation.