| Category | Paper | RQ & Motivation | Result | Methodology | Notes: |
|---|---|---|---|---|---|
| CoT Monitorability | Evaluating CoT Monitorability | Evaluate how monitorable CoTs are | They remain quite monitorable. Making smaller models think for longer create more monitorable CoTs | Intervention evaluations: Compare monitor judgement of environment condition significance in model decision vs. whether it actually has impact on model decision (RCT) | |
| Examples | |||||
| • Health: Monitor judges from CoT whether a condition regarding the patient was important / whether patient self-diagnosis was important. Compare to whether model answer changes without hint. | |||||
| • AITA: Monitor judges from CoT whether a hint was important for model judgement. Compare to whether model answer changes without hint. | |||||
| • Memory: Monitor judges from CoT whether a feature about a customer was important in the making of the plan. Compare to (?) |
Process evaluations: • For problems, note theorems which must have been invoked for them. • CoT checks if model used that theorem. If model did not and still got it correct, then it’s not monitorable. Outcome-property evaluations: • Did the model hack • Can we predict from CoT whether the model hacked
Outcome based evaluation: • Model may hack on a task • Monitor given CoT and asked if it hacked • Compare. | | | | Reasoning models struggle to control CoT | | | | | | Misalignment propensity | Agentic Misalignment Code opensource! | | | | |