During a livestream on Tuesday, OpenAI CEO Sam Altman unveiled the first major enhancement to ChatGPT’s image-generation capabilities in over a year.
ChatGPT now utilizes OpenAI’s GPT-4o model to natively generate and modify images and photos. While GPT-4o has long powered the AI-driven chatbot, it previously could only generate and edit text—not images.
Altman announced that GPT-4o’s native image generation is now available in ChatGPT and Sora, OpenAI’s AI video-generation tool, for subscribers of the company’s $200-per-month Pro plan. OpenAI also stated that the feature will soon roll out to Plus and free-tier ChatGPT users, as well as developers accessing the company’s API service.
Compared to its predecessor, DALL-E 3, GPT-4o’s image-generation model takes slightly longer to process but produces more accurate and detailed images. Additionally, GPT-4o can edit existing images, including those containing people, by transforming them or “inpainting” foreground and background elements.
To power this new capability, OpenAI informed The Wall Street Journal that GPT-4o was trained on “publicly available data” and proprietary datasets obtained through partnerships with companies such as Shutterstock.
Many generative AI vendors regard their training data as a competitive asset and closely guard details about it. Additionally, disclosing such information carries potential legal risks, particularly concerning intellectual property claims—another reason companies remain tight-lipped about their datasets.
“We are mindful of artists’ rights regarding our outputs, and we have policies in place to prevent the generation of images that directly imitate the work of any living artist,” said OpenAI’s Chief Operating Officer Brad Lightcap in a statement to The Wall Street Journal.
OpenAI provides an opt-out form allowing creators to request the removal of their works from its training datasets. The company also honors requests to prevent its web-crawling bots from collecting training data, including images, from websites.
ChatGPT’s expanded image-generation capabilities follow Google’s recent rollout of experimental native image output for Gemini 2.0 Flash, one of its flagship AI models. The feature quickly went viral on social media—though not entirely for positive reasons. Users discovered that Gemini 2.0 Flash’s image-generation tool lacked sufficient safeguards, enabling the removal of watermarks and the creation of images featuring copyrighted characters.



