A great list of things to evaluate (redwood research blog)

Category Paper RQ & Motivation Result Methodology Notes:
CoT Monitorability Evaluating CoT Monitorability Evaluate how monitorable CoTs are They remain quite monitorable. Making smaller models think for longer create more monitorable CoTs Intervention evaluations: Compare monitor judgement of environment condition significance in model decision vs. whether it actually has impact on model decision (RCT)
Examples
• Health: Monitor judges from CoT whether a condition regarding the patient was important / whether patient self-diagnosis was important. Compare to whether model answer changes without hint.
• AITA: Monitor judges from CoT whether a hint was important for model judgement. Compare to whether model answer changes without hint.
• Memory: Monitor judges from CoT whether a feature about a customer was important in the making of the plan. Compare to (?)

Process evaluations: • For problems, note theorems which must have been invoked for them. • CoT checks if model used that theorem. If model did not and still got it correct, then it’s not monitorable. Outcome-property evaluations: • Did the model hack • Can we predict from CoT whether the model hacked

Outcome based evaluation: • Model may hack on a task • Monitor given CoT and asked if it hacked • Compare. | | | | Reasoning models struggle to control CoT | | | | | | Misalignment propensity | Agentic Misalignment Code opensource! | | | | |