It's clear that this feature missed the mark. Some of the images generated are inaccurate or even offensive. [...] So what went wrong? In short, two things. First, our tuning to ensure that Gemini showed a range of people failed to account for cases that should clearly not show a range. And second, over time, the model became way more cautious than we intended and refused to answer certain prompts entirely -- wrongly interpreting some very anodyne prompts as sensitive. [...]
Gemini's image generation launched at the start of February 2024 and was withdrawn three weeks later, after it returned racially varied images for prompts where the history is specific, including German soldiers of 1943 and America's founders, and refused a run of ordinary requests. Google paused image generation of people entirely while it reworked the feature. The post is worth reading as an engineering account rather than an apology: a correction applied globally to avoid one failure mode produced a second one, and a safety layer drifted more conservative than anyone had specified. Both are properties of a system tuned after the fact rather than of the model underneath it.
Prabhakar Raghavan, Google, in Google