It's clear that this feature missed the mark. Some of the images generated are inaccurate or even offensive. [...] So what went wrong? In short, two things. First, our tuning to ensure that Gemini showed a range of people failed to account for cases that should clearly not show a range. And second, over time, the model became way more cautious than we intended and refused to answer certain prompts entirely -- wrongly interpreting some very anodyne prompts as sensitive. These two things led the model to overcompensate in some cases, and be over-conservative in others, leading to images that were embarrassing and wrong.
Curated from blog.google · 23 February 2024 →
Gemini's image generation launched at the start of February 2024 and was withdrawn three weeks later, after it returned racially varied images for prompts where the history is specific, including German soldiers of 1943 and America's founders, and refused a run of ordinary requests. Google paused image generation of people entirely while it reworked the feature. The post is worth reading as an engineering account rather than an apology: a correction applied globally to avoid one failure mode produced a second one, and a safety layer drifted more conservative than anyone had specified. Both are properties of a system tuned after the fact rather than of the model underneath it.