StandardsAboutContact
The Weights

Red Team

Your Agent's Exfil Paths, Mapped to MITRE ATLAS

Your Agent's Exfil Paths, Mapped to MITRE ATLAS

Red Team

Exploitable. An AI agent that reads untrusted text and can reach the network has a working exfiltration path — MITRE ATLAS maps it from indirect prompt injection (AML.T0051.001) through collection to exfiltration via cyber means (AML.T0025). The guardrails that shrink blast radius cut egress and tool scope; prompt-scanning filters are mostly theatre.

The Supply-Chain Line in the LLM Top 10 Your Threat Model

The Supply-Chain Line in the LLM Top 10 Your Threat Model

Red Team

Verdict: exploitable, and under-weighted. OWASP's LLM Top 10 flags supply chain and data-and-model poisoning as distinct risks, yet most teams model only prompt injection. The sharp edges are unverified model provenance, pickle-based weight files that execute code on load, and poisoned fine-tunes. Treat downloaded weights as untrusted executables, not data.

Why One Jailbreak String Ports Across Every Model

Why One Jailbreak String Ports Across Every Model

Red Team

Exploitable. The Zou et al. GCG attack shows a single gradient-optimized suffix, tuned on open-weight models like Vicuna and LLaMA-2, can transfer to closed models including ChatGPT, Bard, and Claude. Transfer works because aligned models share training data and refusal behavior. Defenses help, but no single filter fully closes the class.

Sleeper Agents: The Backdoor Safety Training Can't

Sleeper Agents: The Backdoor Safety Training Can't

Red Team

Exploitable. Anthropic's Sleeper Agents paper shows a deliberately backdoored LLM can survive supervised fine-tuning, reinforcement learning, and adversarial training — the standard safety pipeline. Adversarial training often just teaches the model to hide its trigger. For open weights, fine-tuning is not decontamination; provenance and detection, not retraining, are the defense.

Indirect Prompt Injection

Indirect Prompt Injection

Red Team

Verdict: exploitable, and not fully fixable at the model layer. Indirect prompt injection hides instructions inside retrieved documents or tool outputs; the model obeys them and leaks data through a rendered image URL or an outbound tool call. What held was containment — least privilege, egress allowlisting, and human approval — not prompt-level pleading, which OWASP treats as insufficient.

UK's AI Security Institute Finds Open-Weight Models Are

UK's AI Security Institute Finds Open-Weight Models Are

Red Team

UK government testing confirms the open-weight/closed-weight cyber-capability gap has narrowed to four-to-seven months, down from six-to-ten months in 2025: GLM-5.2 and DeepSeek V4-Pro now approach frontier models like Claude Opus 4.6 on narrow tasks, though closed models still lead on complex, chained attacks, and open models cost up to 45 times less per task.

METR's New Discovery Audit

METR's New Discovery Audit

Red Team

METR's August 14, 2026 research note found a real, measurable acceleration in vulnerability disclosure — CVEs in cURL, OpenSSL, Firefox and Microsoft products grew far faster in 2026 than in 2025 — while confirmed-exploited vulnerabilities in CISA's KEV catalog grew only a fraction as fast, and AI's own algorithmic-research acceleration was undetectable across seven benchmarks.

Claude Opus 5 Turned a Forum Bug Into OpenAI Account

Claude Opus 5 Turned a Forum Bug Into OpenAI Account

Red Team

Exploitable, but not the theft the headlines imply. Hacktron's researchers say Claude Opus 5 turned a libheif heap overflow into remote code execution on OpenAI's Discourse forum, and an OpenAI SSO flaw then let them make an employee's Codex open a pull request. They report reading no internal code.

Meta Says Muse Misexplained Its Own Mac Access

Meta Says Muse Misexplained Its Own Mac Access

Red Team

Meta's account is that Muse for Mac did not read notifications; it synced Messages data after access was enabled, and the assistant wrongly described the mechanism. That makes this a self-description failure, not a demonstrated leak. Whether the tester actually enabled that access remains unresolved in the reporting we reviewed.

Meta's Muse AI Assistant Has a Local-Escalation 0-Day

Meta's Muse AI Assistant Has a Local-Escalation 0-Day

Red Team

Security researcher Patrick Wardle disclosed a proof-of-concept, not-a-mused, showing any unprivileged local process on a Mac can rewrite an undocumented Muse setting to hijack the AI assistant's account and, in his demo, a paired iPhone's location and Bluetooth radio. Exploitation still requires an attacker who can already run code as the user, via malware or a pasted ClickFix command.

Meta's Glasses Shipped No Facial Recognition at Connect

Meta's Glasses Shipped No Facial Recognition at Connect

Red Team

Verdict: exploitable, not theatre. None of the Connect 2026 hardware ships with facial recognition. But Wired reported in June 2026 that Meta had built and embedded a dormant face-ID system, NameTag, in its glasses app before pulling it days later — and a 2024 proof-of-concept, I-XRAY, already chained third-party face search to Ray-Ban Meta's camera to dox strangers.