Row 80431
Content Data
This page contains data entry 80431 from the Axioma AXP content repository. The structured data below represents the complete record for this entry.
As far as I know it is pretty well known that OHE has a lot of downfalls, such as creating a sparse matrix, having very large amount of dimensions as you increase the training text size and also not being able to capture the semantics of sentences/documents like embeddings can.
I dont have any source for this on the top of my head, but it shouldn't be too hard to find in other trusted sources other than papers. If you really want to find that one specific, I'd recommend either searching in google scholar or creating a boolean query in a database such as IEEE with anything you can remember from the text
| Field | Value |
|---|---|
| text | As far as I know it is pretty well known that OHE has a lot of downfalls, such as creating a sparse matrix, having very large amount of dimensions as you increase the training text size and also not being able to capture the semantics of sentences/documents like embeddings can. I dont have any source for this on the top of my head, but it shouldn't be too hard to find in other trusted sources other than papers. If you really want to find that one specific, I'd recommend either searching in goo… |
| label | r/datascience |
| dataType | comment |
| communityName | r/datascience |
| datetime | 2024-05-24 |
| username_encoded | Z0FBQUFBQm5Lak1sbFRsVkRTaEU0dC1tOTU4VWlQNUFwTWg5eTRnU0cwc2F2eFN1QTZXYTNxNkhVZERadW5aTmZTLUVuMmVhMnFCLXk0VUtaaXBsQUg5YWJLQVdUU1NmbGc9PQ== |
| url_encoded | Z0FBQUFBQm5Lak8ydDRrTTRRbi12X3VkelhMWlluMHZqcGRhVVNnTTFKaUU3TmIxUEc2U2NnN1RQWWZPUHNTUUF0d2Z3UkNWb1I1YTJVNDlTUEk4NXJqdTQ0eld3czNnM1lKdlB0MEFmR1JaRld4eDhqUDh0ekYwOVdzMjhpNE5Pb1hTR0hJYV9ubkc0N09sV2ZvdTUzdDhfTmJ2UmtqWGlQckZydWUyN013dHdOamtBdjdpV25CRXA4SFQ5SDRVVzA1dFE2b010S2ZE |
Raw Record
{
"text": "As far as I know it is pretty well known that OHE has a lot of downfalls, such as creating a sparse matrix, having very large amount of dimensions as you increase the training text size and also not being able to capture the semantics of sentences/documents like embeddings can. \n\nI dont have any source for this on the top of my head, but it shouldn't be too hard to find in other trusted sources other than papers. If you really want to find that one specific, I'd recommend either searching in google scholar or creating a boolean query in a database such as IEEE with anything you can remember from the text",
"label": "r/datascience",
"dataType": "comment",
"communityName": "r/datascience",
"datetime": "2024-05-24",
"username_encoded": "Z0FBQUFBQm5Lak1sbFRsVkRTaEU0dC1tOTU4VWlQNUFwTWg5eTRnU0cwc2F2eFN1QTZXYTNxNkhVZERadW5aTmZTLUVuMmVhMnFCLXk0VUtaaXBsQUg5YWJLQVdUU1NmbGc9PQ==",
"url_encoded": "Z0FBQUFBQm5Lak8ydDRrTTRRbi12X3VkelhMWlluMHZqcGRhVVNnTTFKaUU3TmIxUEc2U2NnN1RQWWZPUHNTUUF0d2Z3UkNWb1I1YTJVNDlTUEk4NXJqdTQ0eld3czNnM1lKdlB0MEFmR1JaRld4eDhqUDh0ekYwOVdzMjhpNE5Pb1hTR0hJYV9ubkc0N09sV2ZvdTUzdDhfTmJ2UmtqWGlQckZydWUyN013dHdOamtBdjdpV25CRXA4SFQ5SDRVVzA1dFE2b010S2ZE"
}
Entry Information
- Entry ID: 80431
- Repository: Axioma AXP
- Dataset: arrmlet/reddit_dataset_36
- Total Entries: 100,000