By Austin; still a draft
The AI safety movement should push itself to be dramatically more transparent to the public.
To date, the AI safety movement has been one of the strongest forces for clarity and wisdom in the world. The movement has been prescient on the subject of concerns from existential risk, seriously grappling with outcomes others dismissed as sci-fi nonsense.
Society is now waking up to the potential threats of advanced AI. I understand that many in the movement are feeling the crunch, and thinking more carefully about about optics and what they publish. Even so, acting transparently is more important than ever.
If you want labs to be transparent, you should be transparent too. AI 2040 proposes “Total Research Transparency” for labs to open up their research, algorithms, LLM weights. Safety should do likewise. Model good behavior, to convince labs that this is an acceptable and correct way to behave.
I think AI safety people are unusually virtuous; you should display that virtuosity. “Nor do they light a lamp and then put it under a bushel basket; it is set on a lampstand, where it gives light to all in the house.” (Matthew 5:15)
Transparency ties yourself to the mast, and forces you to be virtuous. Famous maxim: “Act as though what you do might end up on the front page on the NYT”. And, what better way to enforce that then to publish everything you think and do?
You can’t keep things private anyways, given stylometry and cheap intelligence. Actions cast a shadow in the world, and AI will be able to detect that shadow, reconstruct that action. (More here.)
Transparency was a cornerstone of this movement. It is part of what drew me (and many others) to engage and buy into the beliefs of AI safety. Public writings, auditable spreadsheets. Sticking with transparency would demonstrate that the movement has integrity and self-consistency.
501c3 public charities are obliged to some transparency already, around financing and executive salaries. And, ~all AI safety orgs abide by the letter of the law. But perhaps, you should be exemplary with regards to its spirit.
It is a duty of powerful actors to be transparent, and AI safety is becoming powerful. Society already asks for transparency from our political leaders, our government, our labs and megacorps. AI safety wishes to influence major actors and pass sweeping regulation. “Dress for the job you want”; prove yourself worthy of this role.
Transparency helps with internal coordination. It scales well. AI safety is about to undergo hypergrowth. It’s no longer 200 people in the Bay Area who all go to each other’s parties. Knowing what other people think and are doing, regardless of who is in or out of particular group chats, will be key to scaling up.
Transparency helps with external recruitment and fundraising. To date, a lot of AI safety people have come to the movement via public writings and transparent reasoning. Many more great people will join as well, if you continue this way. (If transparency is good for recruiting, does that mean that only public comms-focused orgs like 80k and Bluedot should be transparent, versus our research and policy orgs? I’d argue no.)
Any accounting of transparency in AI safety must begin with the saga of Open Philanthropy.
Givewell began as a strong force for transparency among nonprofits. Holden and Elie published constant updates, maintained a listing of their own mistakes, made their spreadsheets available for public critique, and engaged in good faith with commenters. Take a look at the Givewell Blog circa 2007 for a taste of this; I find their attitude beautiful and inspirational. This transparency was key to the early EA movement, and I believe this helped to convince Dustin Moskovitz and Cari Tuna to start funding Givewell with major amounts of money. Together, they started Givewell Labs to explore causes beyond global health, which became the behemoth Open Philanthropy.
Sadly, this golden era did not last. In 2016, Holden published a major update on how they’re thinking about openness (mostly: less open, due to its costs). Beyond their stated reasons, I suspect that once Good Ventures (Dustin & Cari) became more committed to funding OpenPhil, OpenPhil just had less pressure to continue making its thinking visible to the broader public.
From there, things only became less transparent. I expect Holden remained a major driver for transparency, but he left OpenPhil in 2024. (Around this time, he started the Cold Takes series, which I greatly appreciate for its transparency and wisdom; it has been formative for my views on xrisk.) Then in 2025, OpenPhil gave up on being “Open Philanthropy” altogether and renamed itself to Coefficient Giving.
After some reflection, I consider this sequence of events as a failure of principled thinking. I hesitate to say this because Holden himself was (and remains) one of my personal heroes; he’s certainly in the top 5 of people who have shaped my views. (Others include Scott Alexander, Paul Graham, and Jesus). But: I think that Holden and others didn’t properly honor the role of transparency, in the growth of Givewell and later OpenPhil. Phrased extremely uncharitably: being transparent is what got CG to where it is; to discontinue it now is a bait-and-switch.