ChatGPT Images 2.5 is here: This time, other raw image models really need to be on alert.

CN
链捕手
Follow
1 hour ago

Author: Xiao Jing, Tencent Technology

Editor: Xu Qingyang, Tencent Technology

On September 8, local time in the United States, OpenAI released the latest image generation model ChatGPT Images 2.5.

The new version focuses on enhancing image editing, fidelity to reference images, and the consistency of multiple modifications. When users modify an image consecutively, previously adjusted elements are retained more easily. When changing people, products, or backgrounds, it is also easier to alter only specified parts. The generation speed is up to 50% faster than Images 2.0.

At the same time, ChatGPT has added Sketch, Templates, and image commenting features. Users can directly draw sketches as references, or mark areas on images that need modification. Developers can use the GPT-Image-2.5 Flare and GPT-Image-2.5 Sunburst through the API.

OpenAI states that currently users generate more than 3 billion images weekly through ChatGPT Images and the GPT-Image models in the API. Images 2.5 is already available to ChatGPT, ChatGPT Work, and Codex users, covering web, desktop, and mobile platforms.

ChatGPT Images 2.5 is here: This time, other image generation models should really be nervous

01 Modify one part without messing up the whole image

The real trouble with AI image generation often appears after the first image is generated. For example, if the person is already satisfactory and only wants to change clothes. If the product is already determined, one might just want to change the background. If the poster is already done, one might only want to change a line of text. In the past, continuing to edit could cause other parts of the image to change unexpectedly.

Images 2.5 places a significant emphasis on precise editing. OpenAI provides examples including “full body edit” and “multi-city travel ticket.” These demonstrations use a sequence of images to showcase the modification process, primarily observing whether the model can adjust specified content per instructions while retaining the original subject, composition, and other details.

ChatGPT Images 2.5 is here: This time, other image generation models should really be nervous

The “multi-city travel ticket” case is relatively easy to understand. If an image only needs one piece of information changed, the model needs to identify the specific object for modification while maintaining the overall design structure of the ticket. This capability is more practical for product images, advertising materials, and brand posters than merely generating a nice-looking image.

Another focal point is multi-round editing. OpenAI demonstrates three cases: “cube rotation,” “travel infographic,” and “birthday candle.” They are not finished after one generation but undergo multiple modifications consecutively. OpenAI aims to prove that prior modifications can be preserved while later operations do not continuously damage the completed content.

ChatGPT Images 2.5 is here: This time, other image generation models should really be nervous

This is also the most notable aspect of the Images 2.5 upgrade.

Twitter user @thesoragirls conducted a straightforward test: drawing a candle in the image while designating a petal to remain unchanged, and then observing whether the model would alter the original content while modifying other areas. Her test results showed that the fixed petal was retained while other parts changed as specified.

ChatGPT Images 2.5 is here: This time, other image generation models should really be nervous

However, consecutive editing did not completely resolve detail issues. Japanese animator @genel_ai found in actual testing that while the official emphasis was that image quality could be maintained after multiple edits, artifacts and texture changes might still appear as edits stacked.

ChatGPT Images 2.5 is here: This time, other image generation models should really be nervous

In other words, 2.5 shows some improvement in stability, but quality loss can still occur in complex scenes.

02 Reference images, more stable

The second change in Images 2.5 is its handling of reference images.

Users can provide photos of people, pets, or other subjects and request them to enter a new scene, change visual styles, or alter compositions. OpenAI states that the new version better retains the recognizable features of individuals while improving lighting and texture performance.

The official examples include five cases: “redesigned baby portrait,” “dog in clothing,” “photo booth avatar,” “group party composite photo,” and “made bed.”

ChatGPT Images 2.5 is here: This time, other image generation models should really be nervous

These examples cover various situations. The baby portrait and photo booth avatar primarily focus on whether the character's features can be maintained; the dog case checks whether the subject can retain its original image after changing clothes; the group party composite photo involves multiple characters appearing in one image; “made bed” is closer to regular image editing, requiring specific adjustments to the original scene.

ChatGPT Images 2.5 is here: This time, other image generation models should really be nervous

This ability is crucial for work that requires continuous use of the same person, product, or brand material. API users can generate different versions based on the same reference image, reducing significant changes that may appear each time a new generation is made.

Complex images are also a focus of the official cases this time. OpenAI showcased a 1950s-style family illustration: a family standing in front of a massive cylindrical space habitat containing green landscapes, lakes, and futuristic buildings.

ChatGPT Images 2.5 is here: This time, other image generation models should really be nervous

The OpenAI official introduction page displays a travel page for Yichang: combining travel destination and attraction information with images and text content on one page, creating a complete travel guide. It requires simultaneously handling multiple images, different levels of text, and page layouts, not just completing individual elements.

ChatGPT Images 2.5 is here: This time, other image generation models should really be nervous

There is also a set of nine-grid modernist posters from the medieval period, each with different geometric shapes and text, including “Create,” “Grow together,” and “Choose kindness.”

Another image is an impressionist depiction of the streets of San Francisco, with streets extending down between colorful houses, revealing the bay and the Golden Gate Bridge.

ChatGPT Images 2.5 is here: This time, other image generation models should really be nervous

Additionally, there are cream-gold Lake Como wedding invitations, inverted futuristic cities, eight retro stamps from U.S. national parks, solar flare science slides, blue-gold earth mosaics, ChatGPT sticker posters, and futuristic cities at night in the rain.

ChatGPT Images 2.5 is here: This time, other image generation models should really be nervous

These cases cover different types such as illustrations, posters, invitations, infographics, educational visuals, and product promotional images. OpenAI's intention to display is clear: when the prompt specifies structure, text, visual style, and specific elements simultaneously, Images 2.5 can execute these requirements more completely.

Fashion entrepreneur Yana Welinder believes that Images 2.5 shows significant improvements in fashion design performance. She noted that while using the old models, reference designs could be retained but sometimes resulted in a flatter final effect; she feels that 2.5 allows the original design to present a more complete visual effect.

ChatGPT Images 2.5 is here: This time, other image generation models should really be nervous

However, some have reported conflicting test results. AI and software engineer Mark Kretschmann conducted a noise and artifact test in a forest scene.

ChatGPT Images 2.5 is here: This time, other image generation models should really be nervous

He believes that while 2.5 has many improvements, the test results are inconsistent. He then compared Images 2.5 with Images 2.0, suggesting that in several of his tested cases, 2.0 actually performed better in terms of realism.

Thus, the more accurate judgment currently is that Images 2.5 mainly improves editing control and complex instruction handling, but the actual effects still vary across different types of images.

03 A few strokes, and it understands

In addition to the model itself, OpenAI has added several new operational methods to ChatGPT.

The most direct one is Sketch. Users can draw sketches directly in ChatGPT and then have the model generate the final image based on this sketch. For example, if one wants to design a room, they can first sketch the general layout; if they want to create a piece of clothing, they can outline the shape; or even just draw a rough composition, and hand it off to ChatGPT to complete. Then, they can add visual styles and other requirements later.

When using, simply input “@Sketch” to activate it. This effectively reduces dependence on textual prompts. Some compositions are difficult to describe in words, especially the positions, proportions, and rough outlines of objects. Now users can draw first, then let the model finish.

Twitter user @fquolodasha rated GPT-Image2.5’s hand-drawing abilities highly after testing, stating that this time the results were “pretty awesome,” and remarked on OpenAI’s recent speed of updates and capability enhancements.

ChatGPT Images 2.5 is here: This time, other image generation models should really be nervous

Templates address another issue. ChatGPT has added some common creative formats, such as Poster and Merch. Users first select a template, then fill in the information they want to express, design elements, and visual styles, eliminating the need to start from scratch with a blank canvas every time.

Image sharing also added prompt options. Users can now share the prompt used to generate an image when sharing that image. Others can then use it to replace their own photos and details to continue generating.

An example provided by OpenAI is a portrait in 1980s style: curly hair, colorful jacket, gold chain, neon lights, and a music player. Other users can adopt this concept and insert their own photos.

ChatGPT Images 2.5 is here: This time, other image generation models should really be nervous

04 Two API versions

Images 2.5 has also entered the API. OpenAI launched GPT-Image-2.5 Flare and GPT-Image-2.5 Sunburst.

Flare is the default version, focusing on speed, quality, and editing capabilities. OpenAI claims it generates higher quality than GPT-Image-2 while reducing latency by 50%, making it suitable for social content, product experiences, visual searches, rapid prototyping, and large-scale image generation.

Sunburst, on the other hand, leans towards tasks that require fine control, such as formal advertising materials and high-quality product images. It allows for longer generation times in exchange for higher editing precision.

Early user feedback released by OpenAI also focuses on editing control.

Axultan Alimkulov, product head at Higgsfield AI, specifically noted that Flare can better understand which elements should stay unchanged. For film, UGC, and advertising teams, this means that when modifying an element, the original character, composition, and visual identity can be preserved as much as possible.

Serial entrepreneur @gkxspace observed this update within commercial image workflows. He believes that one of the biggest limitations of AI generated images in the past was that while the first image might be great, it was challenging to stably create a second or third based on it. The changes in Images 2.5 regarding consecutive editing, speed, and fidelity to reference images provide more practical uses in scenarios like e-commerce dressing, brand material extension, and continuous illustrations.

ChatGPT Images 2.5 is here: This time, other image generation models should really be nervous

Currently, Images 2.5 is available to ChatGPT, ChatGPT Work, and Codex users, and the Flare and Sunburst in the API have also been released. OpenAI maintains mechanisms such as prompt and image safety checks, C2PA metadata, and invisible watermarks for identifying images generated by its tools.

From official examples and current user tests, it is clear that the focus of this update is highly concentrated: making post-image generation modifications easier to control, ensuring reference images are more stable in continuous use, and incorporating operations like Sketch, templates, and image annotations into the workflow.

As for realism, details, and image quality after multi-round editing, existing tests have shown different results. The extent to which Images 2.5 can improve these issues still requires more practical use to validate.

免责声明:本文章仅代表作者个人观点,不代表本平台的立场和观点。本文章仅供信息分享,不构成对任何人的任何投资建议。用户与作者之间的任何争议,与本平台无关。如网页中刊载的文章或图片涉及侵权,请提供相关的权利证明和身份证明发送邮件到support@aicoin.com,本平台相关工作人员将会进行核查。

Share To
APP

X

Telegram

Facebook

Reddit

CopyLink