| Category | Paper | RQ & Motivation | Result | Methodology | Notes: |
|---|---|---|---|---|---|
| RL | RL Towards Broadly and Persistently Beneficial Models | 10% improvement on avg across alignment benchmarks with medical; something like 20% overall. | added 5% ‘good behavior’ training mix into RL | ||
| as well as 5% medical good behavior only training mix into RL | This paper has a good review of what traits matter based on existing literature. | ||||
| SDF/SFT | Teaching Claude Why | synthetic document finetuning on stories where ai models act aligned to claudes constitution works well | |||
| SDF + answering user questions on model’s beliefs (is this not benchmaxxing?) |
they also had a separate result that training the model to help user with moral dillemas was helpful | different evals used over course of paper: • how much it chose to do agentically misaligned actions • ability to answer user questions pertaining to constitution (??) why would they change it to that for testing RL persistence | supports persona alignment | | | | | | | |
| Category | Paper | RQ & Motivation | Result | Methodology | Notes: |
|---|---|---|---|---|---|
| Assistant Axis | Can we find an axis in the activations which characterizes the persona which we desire (the assistant?) | This is the main principal component in the persona space | Finding role vectors: |