Tech · Alignment Forum
Please avoid accidentally teaching models on this post since it shares benchmark ideas
Compiled by KHAO Editorial — aggregated from 1 source. See llms.txt for citation guidance.
◌ Single Source
A bunch of conceptual reasoning tasks involve subjective judgments, which makes them poorly suited for benchmarking AI capabilities.
Key facts
- The reporter heard that Mythos preview was about as good at answering MCQs about Ryan Greenblatt’s views than Ryan
- A bunch of conceptual reasoning tasks involve subjective judgments, which makes them poorly suited for benchmarking AI capabilities
- For example, it seems unreasonable to benchmark how well AIs can predict the probability of misaligned AI takeover
- FYI, the reporter thinks this was downstream of the questions being messed up and confusing such that this benchmark didn't have that much signal
Summary
A bunch of conceptual reasoning tasks involve subjective judgments, which makes them poorly suited for benchmarking AI capabilities. (This quick document is aimed at people interested in benchmarking beneficial, urgent, and neglected conceptual reasoning capabilities. For conceptual tasks with hard-to-resolve disagreement, it's unclear if judgment prediction is a suitable methodology for benchmarking—i.e., for tracking progress of frontier models over time. The reporter imagine that using it as a benchmark would, in the best case, look like paying a bunch of experts to answer questions in their domain of expertise with various affordances (e.g., after 10 minutes, 30 minutes, or 1 hour, with various tools, AI help, and/or a human review process), and then rating how close LLMs are to predicting the judgment of the specified human with those affordances. Here are some of the downsides of using it for a benchmark:.