
IFP's 23 AI-Automation Proposals Aren't Law Yet
A think tank's policy menu for automated AI R&D carries no binding force today, but a benchmark jump from an AI research agent and a new game-theory paper on trust suggest the report understates how fast the window to legislate is closing.
None of Institute for Progress's 23 proposed policies for automated AI R&D are binding law — they are a menu addressed to Congress, headlined by an $84 million annual budget ask for CAISI. A concurrent benchmark jump, Intology's Locus agent scoring 51.6% on PostTrainBench+ versus a rival's 23.2% five months earlier, undercuts the report's assumption that preparation time is abundant.
The Weights Desk · 5 min read- IFP's 23 recommendations across 7 categories are proposals to Congress and the executive branch, not enacted policy — zero are currently binding.
- The most concrete near-term ask is $84 million/year in appropriations for CAISI, plus transparency, incident-reporting, and whistleblower-protection legislation.
- Intology's Locus agent jumped from a prior best of 23.2% (March 2026) to 44.7%, then 51.6% on the harder PostTrainBench+ variant using over 4,000 H100-hours — a vendor-reported but directionally telling automation signal.
- A new MIT/Columbia paper, 'Racing to Ruin,' finds transparency — IFP's top-ranked recommendation category — can destabilize cooperation at intermediate trust levels before it helps, complicating a transparency-first strategy.
- The appropriations line for CAISI and any verification-technology funding are the signals worth tracking, not the full 23-item list as an undifferentiated bloc.
Nothing in the Institute for Progress's newly publicized 23-point plan for automated AI R&D is binding. It is a policy menu, not a mandate — and the most useful way to read it is by asking which of its seven recommendation categories could plausibly become law soon, which are aspirational noise, and whether the timeline it assumes still holds. A benchmark result and a new economics paper surfaced in the same week suggest the timeline is shorter, and the mechanism IFP leans on hardest is shakier, than the report lets on.
What's Actually Binding: Nothing
IFP published "How Should the US Prepare for Increasingly Automated AI R&D?" in August 2026, organizing 23 specific measures into seven categories — transparency, state capacity, risk management, verification technology, resilience, extending the US lead, and international-cooperation option value. None have been introduced as legislation or executive action. IFP researcher Tim Fist framed the report as a response to OpenAI, Anthropic, and, he said, over 1,300 employees across major AI companies who want the US government to prepare to "pace" automated AI R&D — itself an unresolved and vaguely defined ask.
Binding Soon, If Adopted
The report's most concrete, appropriations-shaped items are what to actually watch: a request that Congress fund the Center for AI Standards and Innovation at a minimum of $84 million per year, plus legislation covering incident reporting, whistleblower protections, and disclosure of model behavior specifications. It also calls for ensuring US electrical capacity keeps pace with data-center buildout and for using the existing US-China bilateral AI dialogue to jointly manage capability-growth risk. Each is a discrete, trackable ask — unlike the report's more diffuse calls to "accelerate diffusion" or "create option value."
The Number That Explains the Urgency
IFP's report leans on OpenAI and Anthropic leadership's own public estimates of full AI R&D automation by 2028. A live data point in the same newsletter issue tests that clock: Intology's Locus agent, self-reported and not independently audited, scored 44.7% on PostTrainBench — versus a prior best of 23.2% just five months earlier, in March 2026 — then 51.6% on the harder PostTrainBench+ variant using over 4,000 H100-hours, edging past the 51.1% human baseline on post-training Qwen3 base models. Vendor-claimed, but directionally consistent with the automation curve IFP is trying to get ahead of.
Why 'Transparency First' Might Not Work
MIT economist Drew Fudenberg and Columbia's Andrew Koh, in a July 31, 2026 arXiv paper titled "Racing to Ruin," formally model AI-style development races and find that under low trust between competitors, every equilibrium races to disaster with probability one. Their sharper finding complicates IFP's top-ranked category: transparency is double-edged. Faster detection of a rival's moves can first destroy an early-stopping equilibrium at intermediate trust levels — because it becomes cheaper to keep racing while waiting for proof the rival stopped — before eventually restoring cooperation once detection gets fast enough to be self-enforcing.
The Prioritized List, If You Read Nothing Else
Ranked by what is both trackable and load-bearing given the trust dynamics above: first, the CAISI appropriations line, because it is a concrete number Congress can vote on. Second, incident-reporting and whistleblower legislation, because Fudenberg and Koh's model implies transparency needs to be paired with trust-building, not deployed alone. Third, IFP's own verification-technology line item, since it is the closest thing in the report to the trust-restoring mechanism the game theory calls for. International-coordination language ranks lowest for near-term probability.
The verdict: IFP's report is a reasonable preparation checklist, but its framing implicitly assumes there is time to work through all 23 items in sequence. The Locus benchmark jump argues that assumption is optimistic, and the Fudenberg-Koh model argues the report's lead recommendation — transparency — is not self-sufficient without the trust and verification infrastructure IFP buries further down its list. Readers tracking this space should watch the CAISI appropriations fight and whether verification-tech funding survives it, not the 23-item list as a whole.
- Is any part of the IFP's 23-point automation plan currently law?
- No. The Institute for Progress published the recommendations as a policy report in August 2026, addressed to Congress and the executive branch. None of the seven categories or 23 individual items have been enacted; they are a preparatory menu, not a mandate.
- What is PostTrainBench+ and why does it matter to this story?
- It's an expanded-compute variant of PostTrainBench, a benchmark where an agent must post-train an open-weight model above baseline. Intology's Locus agent scored 51.6% on it using over 4,000 H100-hours, beating the 51.1% human baseline — evidence the automated-R&D timeline IFP is preparing for is arriving faster than a static 23-item list can track.
- What does the 'Racing to Ruin' paper say about transparency in AI competition?
- Drew Fudenberg (MIT) and Andrew Koh (Columbia) show formally that under low trust between competitors, every equilibrium races toward disaster with certainty. Transparency is double-edged: faster detection of a rival's actions can first destroy an early-stopping equilibrium at intermediate trust levels before eventually restoring it once detection is fast enough to make stopping self-enforcing.
- Import AI 468: 23 RSI ideas; PostTrainBench+; and how trust and transparency interplay with AI racing — Import AI
- How Should the US Prepare for Increasingly Automated AI R&D? — Institute for Progress
- Racing to Ruin — arXiv
- PostTrainBench: Can LLM Agents Automate LLM Post-Training? — arXiv
- Intology announcement of Locus results on PostTrainBench/PostTrainBench+ — X (Intology)
- Tim Fist on IFP's automated-AI-R&D policy recommendations — X (Tim Fist)