Skip to content

Suggestion: Fix the evaluation mode for a specific subtask beforehand #8

Description

@greentfrapp

Hi, thanks for creating ResearchClawBench! I'm considering doing a submission soon.

As I was reviewing the runs from the public leaderboard, I noticed that the same sub-task may be evaluated as Mode A (Objective Evaluation) in some cases and Mode B (Subjective Evaluation) in other cases. This means the judge may be applying inconsistent scoring rubrics.

For example, for the Physics 000 task, if we compare the public ResearchHarness runs for Claude Opus 4.6 and 4.7, we see:

  • For ResearchHarness (Claude-Opus-4.6): the judge determines Result 1 as Mode B, Result 2 as Mode A, and Result 3 as Mode A.
  • But for ResearchHarness (Claude-Opus-4.7): the judge determines Result 1 as Mode A, Result 2 as Mode A, and Result 3 as Mode B.
    Specifically, Results 1 and 3 are evaluated with different modes across the two runs.

Would you consider fixing the evaluation mode for each subtask beforehand? This could eliminate an additional source of noise in the evaluation.

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Type

    No type

    Projects

    No projects

    Milestone

    No milestone

    Relationships

    None yet

    Development

    No branches or pull requests

    Issue actions