Row 21706

Row ID: 21706 | Dataset Entry | Axioma AXP Content Repository

Content Data

This page contains data entry 21706 from the Axioma AXP content repository. The structured data below represents the complete record for this entry.

This is incorrect. GPT-4o *was* trained to directly output images, and as a result its image consistency is much improved compared to Dall-E 3.

From [Hello GPT-4o](https://openai.com/index/hello-gpt-4o/):

> With GPT-4o, we trained a single new model end-to-end across text, vision, and audio, meaning that all inputs and outputs are processed by the same neural network. Because GPT-4o is our first model combining all of these modalities, we are still just scratching the surface of exploring what the model can do and its limitations.

Followed immediately by explicit examples of its new image output capabilities.

These are not yet active in ChatGPT, though. If you try these examples, it gives the poorly consistent results exactly as Dall-E 3 would give.

This isn't them cherry-picking; this is a legitimately different model. We don't have access to the new capabilities yet.

FieldValue
text This is incorrect. GPT-4o *was* trained to directly output images, and as a result its image consistency is much improved compared to Dall-E 3. From [Hello GPT-4o](https://openai.com/index/hello-gpt-4o/): > With GPT-4o, we trained a single new model end-to-end across text, vision, and audio, meaning that all inputs and outputs are processed by the same neural network. Because GPT-4o is our first model combining all of these modalities, we are still just scratching the surface of exploring what…
label r/openai
dataType comment
communityName r/OpenAI
datetime 2024-05-21
username_encoded Z0FBQUFBQm5Lak1BQjdncGMwRzhoS0dXaWo0bDI0QkdzRmljTnlpVHNvc3Vqdm44RjZqLXgta1k2ZDFxbndaWU1YVjlMc2E0SFBQS0QtNU5DcUxmWlRMRHBZUVdha3hZS0E9PQ==
url_encoded Z0FBQUFBQm5Lak9QelNfVDI4WnplWFRjYnlmQTBBUlA5M2NXR0c5R0VRVm9oZzVIYTRUWUhtaXZMZ1BGSkxkM21KN1NkNkN5U3pBb3BoU2J0aFl2dTFQLTNEbHJvajNqbUhDTUxXdmktVXpEYVM0RFZMSFJPTTVsVXNTVmJ4elVINnpXQlhWM09HTTlzZjVNRk1VZDFlMTF6T3dDbFNlQXZrVVFiYVF2Y1RGcEoyUzk4QnpRV1RSZ3NLbVduR0FVQ05ZMFZwdVJVT01RdEFYYVRjdzc2X1dabl83XzJkOEdHdz09

Raw Record

{
  "text": "This is incorrect. GPT-4o *was* trained to directly output images, and as a result its image consistency is much improved compared to Dall-E 3.\n\nFrom [Hello GPT-4o](https://openai.com/index/hello-gpt-4o/):\n\n> With GPT-4o, we trained a single new model end-to-end across text, vision, and audio, meaning that all inputs and outputs are processed by the same neural network. Because GPT-4o is our first model combining all of these modalities, we are still just scratching the surface of exploring what the model can do and its limitations.\n\nFollowed immediately by explicit examples of its new image output capabilities.\n\nThese are not yet active in ChatGPT, though. If you try these examples, it gives the poorly consistent results exactly as Dall-E 3 would give. \n\nThis isn't them cherry-picking; this is a legitimately different model. We don't have access to the new capabilities yet.",
  "label": "r/openai",
  "dataType": "comment",
  "communityName": "r/OpenAI",
  "datetime": "2024-05-21",
  "username_encoded": "Z0FBQUFBQm5Lak1BQjdncGMwRzhoS0dXaWo0bDI0QkdzRmljTnlpVHNvc3Vqdm44RjZqLXgta1k2ZDFxbndaWU1YVjlMc2E0SFBQS0QtNU5DcUxmWlRMRHBZUVdha3hZS0E9PQ==",
  "url_encoded": "Z0FBQUFBQm5Lak9QelNfVDI4WnplWFRjYnlmQTBBUlA5M2NXR0c5R0VRVm9oZzVIYTRUWUhtaXZMZ1BGSkxkM21KN1NkNkN5U3pBb3BoU2J0aFl2dTFQLTNEbHJvajNqbUhDTUxXdmktVXpEYVM0RFZMSFJPTTVsVXNTVmJ4elVINnpXQlhWM09HTTlzZjVNRk1VZDFlMTF6T3dDbFNlQXZrVVFiYVF2Y1RGcEoyUzk4QnpRV1RSZ3NLbVduR0FVQ05ZMFZwdVJVT01RdEFYYVRjdzc2X1dabl83XzJkOEdHdz09"
}

Entry Information