StandardsAboutContact

The Weights

Breaking News

People Stop Saying "I Don't Know" the MomentThe Fine Print

People Stop Saying "I Don't Know" the Moment

Access to an AI answer, not its accuracy, drove people to nearly stop saying 'I don't know': the rate fell from about 44 percent to roughly 3 percent across a five-experiment, 3,000-plus-participant study, even though the AI was usually wrong. Confidence rose while correct answers dropped to about a third of baseline, per a PsyArXiv preprint reported by The Decoder.

Read More
Teardown

Multiverse Computing's Ising-Optimization Pruning

Bottom line: Wait. Multiverse Computing's block-removal method, detailed on Hugging Face and in arXiv preprint 2602.00161, lifts Llama-3.3-70B-Instruct from a 54.0-MMLU baseline to 76.9 at 50% depth compression — but the edge shrinks to parity at 8B scale, the low-energy-equals-good-model premise breaks after retraining, and zero independent reproductions of any number exist.

Read More
Teardown

GPT-6 Astra Hits 80% on an IKEA-Error Benchmark

OpenAI's GPT-6 Astra scored 80 percent on Epoch AI's 60-photo Furniture Assembly Benchmark, nearly tripling Claude Opus 4.5's 28 percent from November 2025 and beating Claude Fable 5.1 (70 percent) and Claude Opus 5 (61 percent). The gain is real, but the test is tiny and Astra takes three minutes per photo — too slow for live assembly help.

Read More
Teardown

Nvidia's SoL-Pi Is Real, Open-Source Code

Bottom line: Wait. Nvidia's SoL-Pi (arXiv:2609.20519) is real, MIT-licensed, npm-installable code for the Pi coding-agent harness — not vaporware. It cuts EdgeBench tokens up to 49% at 93.7-94.3% score retention, but on Terminal-Bench 4 it solves only 15 of 63 tasks versus 18 for both Pi and Codex, a real capability regression its own headline numbers don't disclose.

Read More

Featured Stories

KDE's 'AI-Native Desktop' Is a Conference Talk

KDE's 'AI-Native Desktop' Is a Conference Talk

Teardown

Verdict: Wait it. The decisive number: zero of three proposed upstream changes — declarative reconciliation, per-widget capabilities, richer Activity metadata — has a merge request, a prototype, or a published threat model. KDE contributors Eva Brucherseifer and Jan Muehlig pitched an 'AI-native' Plasma at Akademy 2026; it remains a talk, not code.

ChatGPT Voice Now Reads Your Email and Slack

ChatGPT Voice Now Reads Your Email and Slack

Teardown

OpenAI's September 23, 2026 update lets ChatGPT Voice call the Gmail, Calendar, and Slack connectors that text-mode ChatGPT and ChatGPT Work already had, running on GPT-6 Astra, Sol, and Luna. Verdict: this closes a modality gap, not a new capability — and OpenAI has published no confirmation mechanism or safety documentation for the wider voice-triggered action surface.

Trump Tells UN the US Will Rename 'Artificial

Trump Tells UN the US Will Rename 'Artificial

The Fine Print

Trump told the UN General Assembly on September 22, 2026 that the US government will replace 'artificial intelligence' with 'super intelligence' in official documents, calling 'artificial' inaccurate. The White House transcript is the only paperwork behind the change: no executive order exists, and it collides with 'superintelligence,' the term AI researchers reserve for capability no shipped system has reached.

Grok 4.7 Teardown: Cheap, Mid-Pack, and Still Behind

Grok 4.7 Teardown: Cheap, Mid-Pack, and Still Behind

Teardown

Bottom line: Skip it for agentic coding and frontier-reasoning work. Grok 4.7 scores 46 on the Artificial Analysis Intelligence Index versus 53 for both Claude Fable 5.1 and GPT-6 — the one number that decides this call. Its sole edge, $2/$6 per million tokens, only pays off for cost-sensitive, non-agentic tasks where mid-pack reasoning is acceptable.

OpenAI's New Math Advisory Panel Can't Touch

OpenAI's New Math Advisory Panel Can't Touch

The Fine Print

OpenAI's new nine-mathematician Advisory Group on Mathematics and Artificial Intelligence, hosted at the Institute for Advanced Study, will review how the company assesses and discloses AI-generated math claims after a disputed Navier-Stokes proof and a credit dispute with NYU's Tristan Buckmaster. Verdict: a credible governance step explicitly barred from the one lever — pacing — that caused the crisis.

Opus 5.5 and GPT-6 Sol/Luna: The Price Cuts Are Real

Opus 5.5 and GPT-6 Sol/Luna: The Price Cuts Are Real

Teardown

Bottom line: Use it for cost-sensitive routing. Anthropic cut Claude Opus 5.5 to $4/$20 per million tokens; OpenAI cut GPT-6 Sol and Luna by up to 58% — per each company's own pricing page, with no stated permanence. Opus 5.5 reportedly exceeded its 128,000-token output limit under maximum reasoning and returned nothing; verify before migrating production agents.

Cloudflare's Python Workers Reach GA

Cloudflare's Python Workers Reach GA

Teardown

Cloudflare's Python Workers exited preview on September 21, 2026, running CPython compiled to WebAssembly via Pyodide inside its V8-based workerd runtime, with bindings into Workers AI, R2, D1, and Hyperdrive. Verdict: usable for single-call, I/O-bound LLM and RAG glue code, but multiprocessing and threading are non-functional — ruling out parallel inference batching and much of the numerical Python stack.

Meta's Muse AI Assistant Has a Local-Escalation 0-Day

Meta's Muse AI Assistant Has a Local-Escalation 0-Day

Red Team

Security researcher Patrick Wardle disclosed a proof-of-concept, not-a-mused, showing any unprivileged local process on a Mac can rewrite an undocumented Muse setting to hijack the AI assistant's account and, in his demo, a paired iPhone's location and Bluetooth radio. Exploitation still requires an attacker who can already run code as the user, via malware or a pasted ClickFix command.

California's New Data-Center Ratepayer Law Doesn't

California's New Data-Center Ratepayer Law Doesn't

The Fine Print

WAIT: California's SB 886 tells the Public Utilities Commission to finalize a tariff protecting ratepayers from data-center grid costs by January 1, 2028 — over a year after Gov. Gavin Newsom signed the seven-bill package on Sept. 21, 2026. An interim filing window opens January 1, 2027, letting some data centers lock in contracts before that protection exists.

AI Chiefs Back a Slowdown, but the Fine Print Binds No One

AI Chiefs Back a Slowdown, but the Fine Print Binds No One

The Fine Print

Treat the AI industry's 'doomer turn' as noise for compliance: executives endorsing a slowdown creates no legal obligation. Only the reported Anthropic embedded-evaluator pledge is a concrete commitment, and it is voluntary. The one empirical result, a Google DeepMind study of 100 agents, is an unreviewed lab finding, not an in-the-wild exploit.

Huang's "0% Chance" on AI Doom Is a Position

Huang's "0% Chance" on AI Doom Is a Position

The Signal

Jensen Huang's "0% chance" that AI ends the world by 2030 is Noise: CBS News describes no model or evaluation behind it. The Signal is the Nvidia chief executive's argument that existing product-liability and cybersecurity law is enough, because that claim is checkable and determines who answers when deployed AI fails.

Meta Says Muse Misexplained Its Own Mac Access

Meta Says Muse Misexplained Its Own Mac Access

Red Team

Meta's account is that Muse for Mac did not read notifications; it synced Messages data after access was enabled, and the assistant wrongly described the mechanism. That makes this a self-description failure, not a demonstrated leak. Whether the tester actually enabled that access remains unresolved in the reporting we reviewed.

California's AI kill-switch order sets a two-month clock

California's AI kill-switch order sets a two-month clock

The Fine Print

Wait. The governor's September 18 executive order does not yet create an AI kill switch: it convenes experts to report within two months, and an independent verification organization would check any emergency shutoff. With no scope, trigger or test published, frontier-model teams have nothing to build against, and SB 813 and AB 1405 timelines are only being accelerated.

Open ASR Leaderboard's Hindi and Indian English Sets

Open ASR Leaderboard's Hindi and Indian English Sets

Teardown

Use it, but read spread, not rank. The Open ASR Leaderboard's new Hindi and Indian English sets, published August 28, 2026, show eight models bunched between 4.81 and 4.99 WER on Indian English, so rank tells you little. The deciding number is regional spread: 0.46 points for Whisper-large-v3-turbo, 1.68 for Voxtral-Mini-3B.

Claude Opus 5 Turned a Forum Bug Into OpenAI Account

Claude Opus 5 Turned a Forum Bug Into OpenAI Account

Red Team

Exploitable, but not the theft the headlines imply. Hacktron's researchers say Claude Opus 5 turned a libheif heap overflow into remote code execution on OpenAI's Discourse forum, and an OpenAI SSO flaw then let them make an employee's Codex open a pull request. They report reading no internal code.

View More Posts

Sign up for the Newsletter

The week in the field, weighed — what ships and where it slips.

Today's briefLoad more