Gutenberg PBC is founded by former frontier lab and big tech technical staff. We are funded by Coefficient Giving, and angels from Anthropic, OpenAI, and CRV. We are looking for technical collaborators and early hires.

Hypothesis generation

We use SAE’s for hypothesis generation. We were the first to apply this method to agent transcripts. In practice, this finds complementary hypotheses to judges. Our SDK is available as a developer tool.

  1. We are looking for further research collaborations with posttrainers. We look to apply this to behavioral evaluations as well.
  2. We have research to improve SAEs and autointerp with agents and NLA’s.

Whitebox proxy

It’d be great to have whitebox access to frontier models to catch eval awareness and reward hacking. But we’ll never get that. How far can one get by prefilling one model with another model’s completions? Could you make the claim that if “deception” is active in Kimi’s mind, then it’s likely also active in Mythos? This could be a useful tripwire

We’re running experiments like this right now

  1. There are 3 major variables: model size, model architecture, and vector. The first step is to benchmark the “faithfulness” of a whitebox proxy to a target model across these dimensions. A priori, we expect faithfulness to be “jagged”
  2. Next is to improve faithfulness
    1. The expensive upper bound is to distill an untrusted blackbox onto a trusted whitebox to make the representations more similar.
      1. From there, we can look for cheaper ways to improve faithfulness, eg if we only care about “deception faithfulness” what’s the minimum distillation needed to meet our requirements
    2. There may be other clever ideas, such as changing representations in-context

We are motivated by this to add interp techniques to scalable oversight.

To register hiring interest, fill this form out. For other reachouts, [email protected].