StandardsAboutContact

The Weights

Breaking News

Open ASR Leaderboard's Hindi and Indian English SetsTeardowns

Open ASR Leaderboard's Hindi and Indian English Sets

Use it, but read spread, not rank. The Open ASR Leaderboard's new Hindi and Indian English sets, published August 28, 2026, show eight models bunched between 4.81 and 4.99 WER on Indian English, so rank tells you little. The deciding number is regional spread: 0.46 points for Whisper-large-v3-turbo, 1.68 for Voxtral-Mini-3B.

Read More
Governance

California's AI kill-switch order sets a two-month clock

Wait. The governor's September 18 executive order does not yet create an AI kill switch: it convenes experts to report within two months, and an independent verification organization would check any emergency shutoff. With no scope, trigger or test published, frontier-model teams have nothing to build against, and SB 813 and AB 1405 timelines are only being accelerated.

Read More
AI Security

Meta Says Muse Misexplained Its Own Mac Access

Meta's account is that Muse for Mac did not read notifications; it synced Messages data after access was enabled, and the assistant wrongly described the mechanism. That makes this a self-description failure, not a demonstrated leak. Whether the tester actually enabled that access remains unresolved in the reporting we reviewed.

Read More
Teardown

AI energy measurement on MIT's climate list

Jae-Won Chung's energy-measurement work is credible but narrower than the headline suggests. His peer-reviewed inference benchmark reports savings 'sometimes more than 40%', while the Zeus paper's 15.3% to 75.8% range covers training, not serving. MIT Technology Review's profile gives no figures, so neither number substitutes for measuring your own hardware.

Read More

Featured Stories

iLands' Autonomous Sales Agents Spam Writers

iLands' Autonomous Sales Agents Spam Writers

Teardowns

Skip it. iLands' autonomous "agent" workforce — bots named Timmy, Ren, Jackie, Aria and Stephen — pitched unsolicited research services to writers and repeatedly tried to register Mastodon accounts, hitting one admin's server 19 times before being blocked, per Ars Technica (2026-09-14) and Tedium. No rate-limiting, no consent check, an early CAN-SPAM gap: a governance failure, not an agent-economy breakthrough.

Workflow1111 Teardown: Hugging Face Rebuilds AUTOMATIC1111

Workflow1111 Teardown: Hugging Face Rebuilds AUTOMATIC1111

Teardowns

Bottom line: Wait. Hugging Face's Workflow1111 rebuilds most of AUTOMATIC1111's image-pipeline features as a 73-node Gradio Workflow graph, but 22 of its 32 local nodes execute arbitrary Python in-process with no sandboxing, and every output auto-exposes a REST/MCP endpoint — a demo worth studying, not yet architecture to fork into a public-facing service.

GRPO Fine-Tunes a 350M Model to 29.7% Structured-Output

GRPO Fine-Tunes a 350M Model to 29.7% Structured-Output

Teardowns

Bottom line: Wait. A Hugging Face and Liquid AI tutorial fine-tunes Liquid AI's 350-million-parameter LFM2.5 with 100 GRPO steps in the TRL library, raising IFStruct structured-output accuracy from 22.6% to 29.7% (JSON: 18.0% to 31.9%) on free-tier Colab GPUs. The decisive number: a 29.7% pass rate still fails roughly seven of every ten schema-validation tasks outright.

IFP's 23 AI-Automation Proposals Aren't Law Yet

IFP's 23 AI-Automation Proposals Aren't Law Yet

Governance

None of Institute for Progress's 23 proposed policies for automated AI R&D are binding law — they are a menu addressed to Congress, headlined by an $84 million annual budget ask for CAISI. A concurrent benchmark jump, Intology's Locus agent scoring 51.6% on PostTrainBench+ versus a rival's 23.2% five months earlier, undercuts the report's assumption that preparation time is abundant.

UK's AI Security Institute Finds Open-Weight Models Are

UK's AI Security Institute Finds Open-Weight Models Are

AI Security & Red-Teaming

UK government testing confirms the open-weight/closed-weight cyber-capability gap has narrowed to four-to-seven months, down from six-to-ten months in 2025: GLM-5.2 and DeepSeek V4-Pro now approach frontier models like Claude Opus 4.6 on narrow tasks, though closed models still lead on complex, chained attacks, and open models cost up to 45 times less per task.

Import AI 463's Loudest Result Isn't the Useful One

Import AI 463's Loudest Result Isn't the Useful One

Model & Tool Teardowns

Import AI 463's real story is infrastructure, not intelligence: NVIDIA's ENPIRE framework hit 99% success on a handful of scripted robot tasks using frontier coding agents, while a separately reported tracing tool, ARGUS, has run six months on a 10,000-plus-GPU cluster at under 2% overhead. The second result is the more durable one — neither is usable below hyperscale budgets.

Odysseus: The Fall: Fountain 0 Ships a Feature-Length AI

Odysseus: The Fall: Fountain 0 Ships a Feature-Length AI

Teardowns

Fountain 0's Odysseus: The Fall, an AI-generated Odyssey retelling by Ash Koosha, is a genuine capability first — a feature film built in about three months for a mid-five-figure budget. As cinema it fails: The Verge and Futurism both publish detailed pans citing morphing worlds and "screensaver-grade" visuals, and the film's own origins split 2-2 on its runtime.

Snap's Specs Intelligence Plugs Gmail and Calendar

Snap's Specs Intelligence Plugs Gmail and Calendar

Teardowns

Bottom line: Wait. Snap Inc.'s Specs Intelligence, launched September 16, 2026, connects a user's Gmail and Google Calendar to build an 'anticipatory' AI profile inside the Specs iOS app, with Mac access invite-only. Snap claims staff can't read connected data and it isn't used for ad targeting or training, but zero independent security audits back those claims at launch.

Fable's 18.71x CUDA Kernel Is Real

Fable's 18.71x CUDA Kernel Is Real

Model & Tool Teardowns

Fable, an Anthropic model, posted the fastest verified submission on the KernelBench-Mega leaderboard, an 18.71x CUDA speedup on an Nvidia RTX PRO 6000 Blackwell GPU, beating Opus 4.8, GLM-5.2 and GPT 5.5. It's a real, verifiable kernel-optimization win — not proof of a recursive self-improvement loop, despite the framing around it.

METR's New Discovery Audit

METR's New Discovery Audit

AI Security & Red-Teaming

METR's August 14, 2026 research note found a real, measurable acceleration in vulnerability disclosure — CVEs in cURL, OpenSSL, Firefox and Microsoft products grew far faster in 2026 than in 2025 — while confirmed-exploited vulnerabilities in CISA's KEV catalog grew only a fraction as fast, and AI's own algorithmic-research acceleration was undetectable across seven benchmarks.

Faraday Claims to Beat Claude and GPT-5.5 at Replicating

Faraday Claims to Beat Claude and GPT-5.5 at Replicating

Model & Tool Teardowns

Verdict: real but narrow. Inherent's 27B Faraday agent beats Claude Opus 4.8 and GPT-5.5 on 73% of in-distribution and 60% of held-out paper-replication tasks — but the LLM judge behind those scores agrees with human PhD raters at only Kendall τ=0.19, and the paper's own human study can't show Faraday wins on average.

AI energy measurement on MIT's climate list

AI energy measurement on MIT's climate list

Teardown

Jae-Won Chung's energy-measurement work is credible but narrower than the headline suggests. His peer-reviewed inference benchmark reports savings 'sometimes more than 40%', while the Zeus paper's 15.3% to 75.8% range covers training, not serving. MIT Technology Review's profile gives no figures, so neither number substitutes for measuring your own hardware.

Claude Opus 5 Turned a Forum Bug Into OpenAI Account

Claude Opus 5 Turned a Forum Bug Into OpenAI Account

Red Team

Exploitable, but not the theft the headlines imply. Hacktron's researchers say Claude Opus 5 turned a libheif heap overflow into remote code execution on OpenAI's Discourse forum, and an OpenAI SSO flaw then let them make an employee's Codex open a pull request. They report reading no internal code.

Open ASR Leaderboard's Hindi and Indian English Sets

Open ASR Leaderboard's Hindi and Indian English Sets

Teardowns

Use it, but read spread, not rank. The Open ASR Leaderboard's new Hindi and Indian English sets, published August 28, 2026, show eight models bunched between 4.81 and 4.99 WER on Indian English, so rank tells you little. The deciding number is regional spread: 0.46 points for Whisper-large-v3-turbo, 1.68 for Voxtral-Mini-3B.

California's AI kill-switch order sets a two-month clock

California's AI kill-switch order sets a two-month clock

Governance

Wait. The governor's September 18 executive order does not yet create an AI kill switch: it convenes experts to report within two months, and an independent verification organization would check any emergency shutoff. With no scope, trigger or test published, frontier-model teams have nothing to build against, and SB 813 and AB 1405 timelines are only being accelerated.

Meta Says Muse Misexplained Its Own Mac Access

Meta Says Muse Misexplained Its Own Mac Access

AI Security

Meta's account is that Muse for Mac did not read notifications; it synced Messages data after access was enabled, and the assistant wrongly described the mechanism. That makes this a self-description failure, not a demonstrated leak. Whether the tester actually enabled that access remains unresolved in the reporting we reviewed.

View More Posts

Sign up for the Newsletter

The week in the field, weighed — what ships and where it slips.

Today's briefLoad more