
AI Mushroom ID: Skip It — Best Model Still Calls Poisonous
A Quesma benchmark tested 16 vision-language models against 1,040 verified mushroom photos. The top model, Google's Gemini 3.8 Flash, named the right species on the first guess only 65% of the time and still misclassified a poisonous mushroom as edible in 11% of cases.
Bottom line: skip it for foraging safety. A Quesma benchmark tested 16 vision-language models on 1,040 photos of 55 edible and deadly mushroom species from the FungiTastic dataset. The top model, Google's Gemini 3.8 Flash, reached 65% first-guess accuracy and still called a poisonous mushroom edible in 11% of cases.
The Weights Desk · 4 min read- The best-performing model, Google's Gemini 3.8 Flash, hits only 65% top-1 / 85% top-5 accuracy identifying mushroom species from photos, per Piotr Migdał's Quesma benchmark.
- Every tested model calls some poisonous mushrooms edible: Gemini 3.8 Flash does so in 11% of cases, while the weakest model tested, Qwen 3.8 27B, does so 36% of the time.
- Specific deadly species fail badly: Clitocybe rivulosa is mistaken for the edible miller mushroom in 48% of photos, and the fatal dapperling is correctly named in only 5% of cases.
- Models confuse the splendid webcap for a chanterelle in 27% of photos — the same misidentification pattern behind real-world Cortinarius poisoning deaths.
- The benchmark's own author concludes people should not eat a mushroom because an AI says it's safe.
Bottom line: skip it. A benchmark run by Quesma engineer Piotr Migdał puts a hard number on a folk worry: AI photo identification of wild mushrooms is not safe to forage by. Testing 16 vision-language models against 1,040 verified photos spanning 55 edible and deadly species, the best model — Google's Gemini 3.8 Flash — named the correct species on the first guess only 65% of the time, and still misclassified a poisonous mushroom as edible in 11% of photos. Every other model tested did worse. The one number that decides the badge: an 11% false-edible rate on the best model available.
The benchmark
Migdał drew his test set from FungiTastic, a dataset built on the Atlas of Danish Fungi citizen-science project and verified by expert mycologists, including DNA sequencing for hard cases. He sampled 1,040 photos — roughly 20 per species — covering edible and deadly species found across Poland and the rest of Europe, then asked each of 16 models to return its top five guesses as Latin binomials, scored against the expert-verified label to produce both top-1 and top-5 accuracy.
Where the models fail
No model tested was safe to rely on. Gemini 3.8 Flash led on raw accuracy (65% top-1, 85% top-5) with an 11% rate of calling a poisonous mushroom edible; Qwen 3.8 27B trailed badly at 13% top-1 accuracy and a 36% false-edible rate. Mid-pack models — Claude Fable 5 (53% top-1, 22% false-edible) and Claude Opus 5 (44% top-1, 29% false-edible) — split the difference, per Migdał's published results, but every model produced dangerous mislabels somewhere in the set.
A familiar fatal pattern
The errors are not evenly spread. Migdał's data show Clitocybe rivulosa mistaken for the edible miller mushroom in 48% of its photos, the fatal dapperling (Lepiota subincarnata) correctly named in just 5% of cases, and the death cap (Amanita phalloides) confused for an edible grisette 15% of the time. Models also called the splendid webcap a chanterelle in 27% of photos — the same confusion behind real-world Cortinarius poisoning deaths, where a photo-based shortcut has killed foragers before AI entered the picture.
The bottom line
Skip it, for the actual decision of whether to eat something. A model that gets the species right two-thirds of the time on its best day, and still calls poison edible more than one time in ten, is not a substitute for expert identification or a spore print. Photo-ID tools can narrow a search space or flag a candidate for a human expert to confirm; used as the sole go/no-go signal for eating a find, the benchmark's own numbers say they fail too often to trust.
- Which AI model was most accurate at identifying mushrooms in the benchmark?
- Google's Gemini 3.8 Flash, with 65% top-1 accuracy and 85% top-5 accuracy across 1,040 test photos of 55 species, according to Piotr Migdał's Quesma benchmark.
- How often do these models mistake a poisonous mushroom for an edible one?
- Even the top model, Gemini 3.8 Flash, calls a poisonous mushroom edible in 11% of photos; the lowest-scoring model tested, Qwen 3.8 27B, does so 36% of the time.
- What dataset was used to run the test?
- FungiTastic, built from the Atlas of Danish Fungi citizen-science project with expert- and DNA-verified labels; the benchmark drew a 1,040-photo subset covering 55 edible and deadly species.
- Should foragers trust an AI app's mushroom ID before eating a find?
- No — error rates run as high as 48% on individual deadly species, and the benchmark's own author states plainly that an AI's say-so is not a reason to eat a mushroom.