StandardsAboutContact
The Weights
GPT-6 Astra Hits 80% on an IKEA-Error Benchmark

GPT-6 Astra Hits 80% on an IKEA-Error Benchmark

Epoch AI's Furniture Assembly Benchmark shows a real, multi-vendor jump in vision-model error-spotting since November 2025. The sample size and processing time mean the obvious 'AI assembly buddy' pitch doesn't hold up yet.

OpenAI's GPT-6 Astra scored 80 percent on Epoch AI's 60-photo Furniture Assembly Benchmark, nearly tripling Claude Opus 4.5's 28 percent from November 2025 and beating Claude Fable 5.1 (70 percent) and Claude Opus 5 (61 percent). The gain is real, but the test is tiny and Astra takes three minutes per photo — too slow for live assembly help.

The Weights Desk · 4 min read

OpenAI's GPT-6 Astra correctly flagged assembly errors in IKEA furniture photos 80 percent of the time, according to an independent benchmark Epoch AI published September 23, 2026 — nearly triple the 28 percent ceiling Claude Opus 4.5 hit ten months earlier, in November 2025. It's a genuine capability jump on a task vision-language models have historically flubbed. But the benchmark is 60 photographs across three furniture builds, and Astra needs roughly three minutes to analyze each one, so anyone picturing a phone held up mid-build for instant correction should recalibrate.

What Epoch AI Actually Measured

Epoch AI's Furniture Assembly Benchmark (FAB) photographs IKEA pieces mid-build, some correctly assembled and some doctored with deliberate errors, and asks a model to spot the mistake against the official instruction manual, using zoom tools and a Python interpreter to inspect the image. Grading combines whether the model names the right faulty step with whether its description of the error matches, itself checked by an LLM judge. The design tests a narrow but genuinely hard skill: mapping a 2D photo onto a multi-step physical process and catching where it diverged.

The Score Jump Is Real — the Sample Isn't Big Enough to Lean On

Claude Opus 4.5 scored 28 percent in November 2025; ten months later, GPT-6 Astra scored 80 percent, Claude Fable 5.1 scored 70 percent, and Claude Opus 5 scored 61 percent, per Epoch AI. That's a real, multi-model gain, not a cherry-picked win for one lab. But the entire test set is 60 images spanning three furniture pieces — too small to establish whether the gain reflects general physical reasoning or overfitting to IKEA-style diagrams specifically, a limitation Epoch AI itself flags in the write-up.

Three Minutes a Photo Kills the Obvious Use Case

The headline framing — a model that catches your shelf mistake — implies live troubleshooting mid-assembly. Epoch AI clocked Astra's median response time at about three minutes per photo, two to ten times faster than the other models it tested but nowhere near conversational. That rules out the assistant-looking-over-your-shoulder scenario entirely; the realistic near-term use is asynchronous — submit a photo, get a verdict minutes later, closer to a quality-control checkpoint than a live coach.

Who Should Actually Care

The plausible near-term buyer isn't a consumer assembling a bookshelf — it's furniture retailers, logistics QA teams, or repair and warranty operations that already process photos asynchronously and could use an 80-percent-accurate error-spotter to triage returns or damage claims. Robotics and appliance-repair applications, floated in The Decoder's report, remain speculative extensions of a benchmark built entirely around flat-pack furniture and not validated on either domain.

The Verdict

The jump from 28 percent to 80 percent in ten months is a legitimate signal that frontier vision-language models are getting sharply better at step-by-step physical-process verification, and it isn't just OpenAI — Anthropic's two most recent Claude releases also cleared 60 percent. But treat this as an early, narrow-domain result, not proof of general physical reasoning: 60 images, three products, three minutes a photo. Buy the direction; don't buy the 'AI assembly buddy' framing until the benchmark grows past one furniture brand and the latency drops by an order of magnitude.

Does this mean GPT-6 Astra can guide someone through assembling furniture in real time?
No. Epoch AI clocked a median response time of about three minutes per photo, which the benchmark's own authors say is too slow for live, step-by-step assembly guidance.
How rigorous is the Epoch AI Furniture Assembly Benchmark?
It's a narrow 60-image test spanning three IKEA builds, graded partly by an LLM judge — a legitimate independent signal, but too small to prove the accuracy jump generalizes beyond flat-pack furniture to broader physical or spatial reasoning.
How do the actual scores compare?
GPT-6 Astra: 80 percent. Claude Fable 5.1: 70 percent. Claude Opus 5: 61 percent. Claude Opus 4.5, the November 2025 baseline: 28 percent, per Epoch AI's published results.
  1. Can AI spot mistakes in IKEA assembly? — Epoch AI
  2. OpenAI's GPT-6 Astra can now tell you exactly where you screwed up your IKEA shelf — The Decoder