Row 7785
Content Data
This page contains data entry 7785 from the Axioma AXP content repository. The structured data below represents the complete record for this entry.
OpenAI claims that *GPT-4o* is much better at image recognition than *GPT-4-turbo*. I wanted to see how it fares with screenshots of Web UI pages. I see two potential goals here that I'm interested in: 1. Accessibility 2. Automated web testing (to replace frameworks like Selenium or Cypress)
I made a YouTube video showing some of the tests I did: [https://www.youtube.com/watch?v=ZFzBDPpeP04](https://www.youtube.com/watch?v=ZFzBDPpeP04)
However, the conclusion is that both *GPT-4o* and *GPT-4-turbo* are equally bad at this task. I couldn't see any kind of improvement with *GPT-4o*. It's a bit disappointing, because when I test it with regular photos, both models are very good. But they really struggle with details on Web UI screenshots.
I'm curious of other people tried similar use cases and what were your results.
| Field | Value |
|---|---|
| text | OpenAI claims that *GPT-4o* is much better at image recognition than *GPT-4-turbo*. I wanted to see how it fares with screenshots of Web UI pages. I see two potential goals here that I'm interested in: 1. Accessibility 2. Automated web testing (to replace frameworks like Selenium or Cypress) I made a YouTube video showing some of the tests I did: [https://www.youtube.com/watch?v=ZFzBDPpeP04](https://www.youtube.com/watch?v=ZFzBDPpeP04) However, the conclusion is that both *GPT-4o* and *GP… |
| label | r/openai |
| dataType | post |
| communityName | r/OpenAI |
| datetime | 2024-05-17 |
| username_encoded | Z0FBQUFBQm5LakwzWDI1NHN4VlRKTmxubExnUjNzMEJBU0JUeExjaTNYUm9Ibld4Q0ctYUNtSDFRVE5MVkVLUmVlTE5YT3dTTzdiRUJURFVPd05pa2l2b1RUc0pfam1RX3c9PQ== |
| url_encoded | Z0FBQUFBQm5Lak9ISHBpUjN1MW81V1psWV9GdE9TTHE4T3FyRFZzLWo4eFhlNjRkSG1vLXA4VkNkRmxiMXRldXZ2U0p1OUtyTDlBa182U1pKZVRCdlpoSFpCUUtwSWVDWjBxazBDUU5fUFdiNkoxNnppY0t6RkgyS19RRzl2T0ZmV1NwMXhRVXBqWWtNbktrT05MaXJ0dENxeHItdVdZcTdCSUREZVZEREs5M3lNSWQyeDJEMGdUX1dtUlBFanZQbWcyaGRpYmVDb2FHM2JIU1o3SWxXYXJHUlB0YVRKMU5Hdz09 |
Raw Record
{
"text": "OpenAI claims that *GPT-4o* is much better at image recognition than *GPT-4-turbo*. I wanted to see how it fares with screenshots of Web UI pages. I see two potential goals here that I'm interested in: \n1. Accessibility \n2. Automated web testing (to replace frameworks like Selenium or Cypress)\n\nI made a YouTube video showing some of the tests I did: [https://www.youtube.com/watch?v=ZFzBDPpeP04](https://www.youtube.com/watch?v=ZFzBDPpeP04)\n\nHowever, the conclusion is that both *GPT-4o* and *GPT-4-turbo* are equally bad at this task. I couldn't see any kind of improvement with *GPT-4o*. It's a bit disappointing, because when I test it with regular photos, both models are very good. But they really struggle with details on Web UI screenshots.\n\nI'm curious of other people tried similar use cases and what were your results.",
"label": "r/openai",
"dataType": "post",
"communityName": "r/OpenAI",
"datetime": "2024-05-17",
"username_encoded": "Z0FBQUFBQm5LakwzWDI1NHN4VlRKTmxubExnUjNzMEJBU0JUeExjaTNYUm9Ibld4Q0ctYUNtSDFRVE5MVkVLUmVlTE5YT3dTTzdiRUJURFVPd05pa2l2b1RUc0pfam1RX3c9PQ==",
"url_encoded": "Z0FBQUFBQm5Lak9ISHBpUjN1MW81V1psWV9GdE9TTHE4T3FyRFZzLWo4eFhlNjRkSG1vLXA4VkNkRmxiMXRldXZ2U0p1OUtyTDlBa182U1pKZVRCdlpoSFpCUUtwSWVDWjBxazBDUU5fUFdiNkoxNnppY0t6RkgyS19RRzl2T0ZmV1NwMXhRVXBqWWtNbktrT05MaXJ0dENxeHItdVdZcTdCSUREZVZEREs5M3lNSWQyeDJEMGdUX1dtUlBFanZQbWcyaGRpYmVDb2FHM2JIU1o3SWxXYXJHUlB0YVRKMU5Hdz09"
}
Entry Information
- Entry ID: 7785
- Repository: Axioma AXP
- Dataset: arrmlet/reddit_dataset_36
- Total Entries: 100,000