Row 7785

Row ID: 7785 | Dataset Entry | Axioma AXP Content Repository

Content Data

This page contains data entry 7785 from the Axioma AXP content repository. The structured data below represents the complete record for this entry.

OpenAI claims that *GPT-4o* is much better at image recognition than *GPT-4-turbo*. I wanted to see how it fares with screenshots of Web UI pages. I see two potential goals here that I'm interested in: 1. Accessibility 2. Automated web testing (to replace frameworks like Selenium or Cypress)

I made a YouTube video showing some of the tests I did: [https://www.youtube.com/watch?v=ZFzBDPpeP04](https://www.youtube.com/watch?v=ZFzBDPpeP04)

However, the conclusion is that both *GPT-4o* and *GPT-4-turbo* are equally bad at this task. I couldn't see any kind of improvement with *GPT-4o*. It's a bit disappointing, because when I test it with regular photos, both models are very good. But they really struggle with details on Web UI screenshots.

I'm curious of other people tried similar use cases and what were your results.

FieldValue
text OpenAI claims that *GPT-4o* is much better at image recognition than *GPT-4-turbo*. I wanted to see how it fares with screenshots of Web UI pages. I see two potential goals here that I'm interested in: 1. Accessibility 2. Automated web testing (to replace frameworks like Selenium or Cypress) I made a YouTube video showing some of the tests I did: [https://www.youtube.com/watch?v=ZFzBDPpeP04](https://www.youtube.com/watch?v=ZFzBDPpeP04) However, the conclusion is that both *GPT-4o* and *GP…
label r/openai
dataType post
communityName r/OpenAI
datetime 2024-05-17
username_encoded Z0FBQUFBQm5LakwzWDI1NHN4VlRKTmxubExnUjNzMEJBU0JUeExjaTNYUm9Ibld4Q0ctYUNtSDFRVE5MVkVLUmVlTE5YT3dTTzdiRUJURFVPd05pa2l2b1RUc0pfam1RX3c9PQ==
url_encoded Z0FBQUFBQm5Lak9ISHBpUjN1MW81V1psWV9GdE9TTHE4T3FyRFZzLWo4eFhlNjRkSG1vLXA4VkNkRmxiMXRldXZ2U0p1OUtyTDlBa182U1pKZVRCdlpoSFpCUUtwSWVDWjBxazBDUU5fUFdiNkoxNnppY0t6RkgyS19RRzl2T0ZmV1NwMXhRVXBqWWtNbktrT05MaXJ0dENxeHItdVdZcTdCSUREZVZEREs5M3lNSWQyeDJEMGdUX1dtUlBFanZQbWcyaGRpYmVDb2FHM2JIU1o3SWxXYXJHUlB0YVRKMU5Hdz09

Raw Record

{
  "text": "OpenAI claims that *GPT-4o* is much better at image recognition than *GPT-4-turbo*.  I wanted to see how it fares with screenshots of Web UI pages. I see two potential goals here that I'm interested in:  \n1. Accessibility  \n2. Automated web testing (to replace frameworks like Selenium or Cypress)\n\nI made a YouTube video showing some of the tests I did: [https://www.youtube.com/watch?v=ZFzBDPpeP04](https://www.youtube.com/watch?v=ZFzBDPpeP04)\n\nHowever, the conclusion is that both *GPT-4o* and *GPT-4-turbo* are equally bad at this task. I couldn't see any kind of improvement with *GPT-4o*.  It's a bit disappointing, because when I test it with regular photos, both models are very good. But they really struggle with details on Web UI screenshots.\n\nI'm curious of other people tried similar use cases and what were your results.",
  "label": "r/openai",
  "dataType": "post",
  "communityName": "r/OpenAI",
  "datetime": "2024-05-17",
  "username_encoded": "Z0FBQUFBQm5LakwzWDI1NHN4VlRKTmxubExnUjNzMEJBU0JUeExjaTNYUm9Ibld4Q0ctYUNtSDFRVE5MVkVLUmVlTE5YT3dTTzdiRUJURFVPd05pa2l2b1RUc0pfam1RX3c9PQ==",
  "url_encoded": "Z0FBQUFBQm5Lak9ISHBpUjN1MW81V1psWV9GdE9TTHE4T3FyRFZzLWo4eFhlNjRkSG1vLXA4VkNkRmxiMXRldXZ2U0p1OUtyTDlBa182U1pKZVRCdlpoSFpCUUtwSWVDWjBxazBDUU5fUFdiNkoxNnppY0t6RkgyS19RRzl2T0ZmV1NwMXhRVXBqWWtNbktrT05MaXJ0dENxeHItdVdZcTdCSUREZVZEREs5M3lNSWQyeDJEMGdUX1dtUlBFanZQbWcyaGRpYmVDb2FHM2JIU1o3SWxXYXJHUlB0YVRKMU5Hdz09"
}

Entry Information