StandardsAboutContact
The Weights
  1. KDE's 'AI-Native Desktop' Is a Conference Talk

    At Akademy 2026, two longtime KDE contributors pitched a 'sovereign' Plasma that assembles itself around a personal AI model. The proposal has a talk title and three architectural asks; it has no upstream code, no prototype, and no published threat model.

  2. Meta's Glasses Shipped No Facial Recognition at Connect

    None of the hardware Meta unveiled at Connect 2026 carries facial recognition. But Wired caught Meta's own app running a near-built version of the feature months earlier, and a 2024 student project already proved the same attack works without Meta's help.

  3. Meta's Muse Proves Agentic Checkout Works

    Meta's Muse agent checks out on real, live payment rails — Stripe's Link wallet and Shopify's Shop Pay — but the company has published no accuracy or error data for purchases the agent completes on its own, and Amazon has already blocked it over security concerns. For any team building an agent with its own spend authority, that gap matters more than the download count.

  4. ChatGPT Voice Now Reads Your Email and Slack

    OpenAI wired ChatGPT Voice into the same Gmail, Calendar, and Slack connectors text-mode ChatGPT already had, on newly upgraded GPT-6 models. The capability isn't new; what's missing is any fresh account of how voice-triggered actions get confirmed before they fire.

  5. People Stop Saying "I Don't Know" the Moment

    A five-experiment study of more than 3,000 people found that merely showing an AI-generated answer collapsed participants' willingness to admit uncertainty, even when that answer was wrong. The result, reported by The Decoder, reframes AI overreliance as a measurable calibration failure rather than a vague usability concern.

  6. Multiverse Computing's Ising-Optimization Pruning

    Multiverse Computing says treating LLM block removal as an Ising-glass optimization problem beats standard pruning by nearly 23 MMLU points on a 70B model. The number is real inside the company's own paper; nobody outside it has checked it yet.

  7. GPT-6 Astra Hits 80% on an IKEA-Error Benchmark

    Epoch AI's Furniture Assembly Benchmark shows a real, multi-vendor jump in vision-model error-spotting since November 2025. The sample size and processing time mean the obvious 'AI assembly buddy' pitch doesn't hold up yet.

  8. Nvidia's SoL-Pi Is Real, Open-Source Code

    Nvidia's SoL-Pi harness optimizer is a genuine, MIT-licensed, npm-installable release, not a paper-only research result. But on Terminal-Bench 4, the paper's own numbers show it solving fewer tasks than the baseline it was built to beat.

  9. Nemotron 3 Diarization: Nvidia's Free 100M-Parameter

    Nvidia's new open-weight speaker-diarization model tops an independent benchmark at 14.72% error, doubling its predecessor's speaker capacity to eight. The catch: its speaker labels are anonymous, session-scoped guesses, and Nvidia's documentation warns against treating them as verified identity.

  10. AI Mushroom ID: Skip It — Best Model Still Calls Poisonous

    A Quesma benchmark tested 16 vision-language models against 1,040 verified mushroom photos. The top model, Google's Gemini 3.8 Flash, named the right species on the first guess only 65% of the time and still misclassified a poisonous mushroom as edible in 11% of cases.

  11. Gemini 4 Argon ties GPT-6 Astra on Artificial Analysis

    Google's first frontier model in over seven months lands in the top tier of independent testing without taking the lead. It uses more than twice the output tokens of GPT-6 Astra per task, and its cost advantage depends on introductory pricing.

  12. Photo Scrubber: A Local Face-Blur Tool Built by an AI

    Simon Willison released an experimental browser tool that blurs faces and removes metadata, built with GPT-6 Astra on MediaPipe and BlazeFace. The architecture is sound for privacy, but the post gives no detection-recall figure, and for a redaction tool that is the number that matters.