Gutenberg PBC is founded by former frontier lab and big tech technical staff. We are funded by Coefficient Giving, and angels from Anthropic, OpenAI, and CRV. We are looking for technical collaborators and early hires.
We use SAE’s for hypothesis generation. We were the first to apply this method to agent transcripts. In practice, this finds complementary hypotheses to judges. Our SDK is available as a developer tool.
It’d be great to have whitebox access to frontier models to catch eval awareness and reward hacking. But we’ll never get that. How far can one get by prefilling one model with another model’s completions? Could you make the claim that if “deception” is active in Kimi’s mind, then it’s likely also active in Mythos? This could be a useful tripwire
We’re running experiments like this right now
We are motivated by this to add interp techniques to scalable oversight.
To register hiring interest, fill this form out. For other reachouts, [email protected].