CrucibleBench places language models in a persistent MUD, a text world where NPCs remember, trust accumulates, and mistakes leave traces, and scores what they do over 50 turns with hidden social objectives.
Status: Phase 1 proof-of-concept released · 13 models · 650 runs · $99.59 billed. Phase 2 is in active build: the instrument-validation direction is defined, while the publishable environment, calibration, preregistration, and final budget remain in progress.
Read in full here: