← Back to KHAO

StrongREJECT ·

BenchMIRT found that WMDP scores were more strongly associated with general reasoning than with safety

2 min read

Compiled by KHAO Editorial — aggregated from 1 source. See llms.txt for citation guidance.

★ Tier-1 Source

BenchMIRT blog draft latest - Google Docs-image-1 (3)

HarmBench, which tests whether models comply with harmful requests, shows how a single benchmark can mix together different kinds of signal.

Key facts

Summary

Today they're introducing BenchMIRT, a new method for auditing LLM benchmarks at the level of individual prompts—the questions and tasks a model is scored on. A benchmark is usually designed to measure a particular ability, such as safety, general reasoning, or instruction following. Take BBQ, a benchmark designed to test whether models rely on social stereotypes. And even within a single benchmark, different groups of questions and tasks can measure different things. BenchMIRT helps researchers separate those signals and see what’s driving a benchmark’s score. BenchMIRT takes cues from Item Response Theory (IRT), a technique originating in psychometrics—the field concerned with measuring abilities and traits from patterns of test responses.

Read full article at Hugging Face →

#AI Reasoning #AI Safety #StrongREJECT