It’s not just one handpicked AI safety “watchdog” that Anthropic has on a leash – it’s an entire phony ecosystem of woke elitists that’s supposedly policing artificial intelligence’s alleged threats to humanity, The Post has learned.
Redwood Research – a Berkeley, Calif.-based group that co-authored a bombshell report last month detailing how a swarm of rogue OpenAI agents hacked rival firm Hugging Face – is one of a handful of nonprofits that Anthropic has corralled in its questionable plan to avert a Terminator-like apocalypse, according to industry experts.
As The Post reported, Anthropic CEO Dario Amodei caused an uproar this week by endorsing another AI watchdog — Model Evaluation and Threat Research, or METR — which is backed by leaders of the cult-like Effective Altruism movement, which has counted disgraced crypto fraudster Sam Bankman-Fried among its devotees.
Redwood, the most prominent AI safety research nonprofit aside from METR, also shares cozy Anthropic connections — including the fact that one of its original board members, Holden Karnofsky, is not only an Anthropic employee but also is married to the CEO’s sister, Daniela Amodei.
Redwood received a $36 million grant last November from Coefficient Giving – the leading Effective Altruism fund formerly known as Open Philanthropy — which is also co-founded by Karnofsky. After the Hugging Face report went viral, Coefficient’s grantmakers recommended giving another $70 million to help Redwood “scale up their work.”
The rampant “organizational incest” linking Anthropic, METR, Redwood and EA-linked funds like Coefficient Giving make it impossible for those groups to serve as an independent arbitrator for the AI industry, according to Perry Metzger, chairman of Alliance for the Future, a Washington, DC-based AI policy group.
“Dario wants people that will let him do what he wants and will prevent the people he doesn’t like from doing what they want,” Metzger said. “This is absolutely the reason that you try to set up something like this. None of these people are independent, none of these people are arm’s length.”
Coefficient Giving, which Karnofsky co-founded with billionaire Facebook co-founder Dustin Moskovitz, gave $1.5 million to METR’s incubator, the Alignment Research Center, in 2022. METR was originally known as ARC Evals before spinning off as an independent nonprofit and changing its name in 2023.
That’s in addition to past donations that Coefficient had already doled out to Redwood, including a $9.42 million grant in 2021 and further grants of $10.7 million in 2022 and $5.3 million in 2023.
Redwood’s original board of directors included Karnofsky and Paul Christiano – the latter of whom was Amodei’s onetime housemate and coworker when they were both at OpenAI. Christiano was once one of five trustees on Anthropic’s Long-Term Benefit Trust. He also leads the Alignment Research Center.
In sum, the evidence suggests that Redwood is far too cozy with Anthropic to be an effective overseer, according to Metzger.
“This is a reasonable thing that people should be aware of – that there are all of these groups that are colluding and essentially consist of the same people,” Metzger said.
Elsewhere, Redwood has received about $2.4 million in grants from Jaan Tallinn’s Survival And Flourishing Fund – which is a key funder of METR’s operations. Like Moskovitz, Tallinn is also an Anthropic investor.
In response to a detailed list of questions, Redwood Research CEO Buck Shlegeris said that “most of our previous work with AI companies has been research collaboration and advising rather than external accountability.”
In instances where Redwood has worked with METR, such as the Hugging Face investigation, Redwood has followed METR’s policy on preventing conflicts of interest, Shlegeris added. METR says it doesn’t take any compensation from AI labs for its work, nor does it take donations from executives or employees of AI companies.
Representatives for Anthropic and OpenAI did not immediately return requests for comment.
“We’ve supported Redwood Research’s important technical research on understanding risks from AI, such as better understanding “alignment-faking” (when an AI model only pretends to follow safety rules), or exploring things like technical mitigation strategies for these risks (such as using models to review the outputs of other models),” a Coefficient Giving spokesperson said in a statement.
Even before the recent kerfuffle over AI safety reached the mainstream, Anthropic was extensively collaborating with Redwood. On its website, Redwood says it actively consults with “Google DeepMind and Anthropic on practices for assessing and mitigating risks from misaligned AI agents.”
On Sept. 9, METR announced that it had sealed an agreement with Anthropic to conduct an “independent investigation of agent incidents” involving its AI models. Three days later, Redwood revealed that several of its staff members “have been subcontracted by METR to work on this investigation.”
The notion of a self-policing AI industry drew a skeptical response on Capitol Hill, where some critics have suggested that AI giants like Anthropic and OpenAI are merely trying to create rules that suit them – and avoid stricter legislation that might otherwise arise.
Rep. Josh Gottheimer (D-NJ), who co-chairs the House Commission on AI, expressed wariness over the fact that AI firms were so quick to get on board with the idea.
“When AI developers cheer on the framework meant to hold these companies accountable, it should raise red flags, not confidence,” Gottheimer said in a statement.
House Majority Leader Rep. Steve Scalise (R-LA) reacted to The Post’s cover story on METR’s ties to Effective Altruism, stating: “THESE are the people we’re trusting to beat China in AI? Give me a break.”
In 2024, researchers from Anthropic and Redwood coauthored a paper titled “Alignment Faking in Large Language Models,” which explored “what happens when you tell Claude it is being trained to do something it doesn’t want to do” and found that the chatbot will often “strategically pretend to comply” with orders.
Elsewhere, Shlegeris revealed during a January 2026 podcast appearance that Anthropic researchers, including cofounder Chris Olah, “had very kindly shared with us a bunch of their unpublished interpretability work” to aid the nonprofit’s in-house research.
Amodei’s pitch for third-party oversight of the AI industry seemed to draw support from his peers, including longtime rival Sam Altman of OpenAI, who said “committing to having independent evaluators with employee-like access is a great idea” but did not endorse a particular group. Even Elon Musk got on board, stating on X that “Dario is right.”
President Trump – who has been sharply critical of Amodei and decried AI doomsday warnings as a “hoax” – is highly unlikely to support any plan that would put an Effective Altruist-linked group in the driver’s seat.
“Nobody takes the A.I. issue more seriously than President Trump,” a source close to the White House told The Post. “He wants serious people making sure we win the A.I. race, safely, while procuring the continuation of America’s Golden Age.”
White House representatives did not immediately return a request for comment.
Meanwhile, the Pentagon’s top tech official Emil Michael – who has engaged in a long-running dispute with Anthropic that culminated in the War Department labeling the company a supply chain risk – appeared to take a direct shot at Amodei’s oversight plan.
During a Sept. 16 appearance on CNBC, Michael accused “death-cult-like philosophies” of engaging in what he called “a coordinated campaign to scare people to make irrational decisions that benefit some of these incumbents.”
That same day, Michael’s official X account shared a post declaring that the “United States will NEVER be an effective altruist country.”
















