Row 8058
Content Data
This page contains data entry 8058 from the Axioma AXP content repository. The structured data below represents the complete record for this entry.
Hi Guys, I'm needing to decide which library to use for named entity recognition. I've used spaCy, which works well, but I need a library that allows me to categorize entities and also sub-entities. Has anyone done something similar? I mean, where the same word can be more than one entity. spaCy offers the SpanCat pipeline, which theoretically allows this, but l've had trouble creating the training corpus. I think it's because they expect you to purchase an annotation text framework like Prodigy.
| Field | Value |
|---|---|
| text | Hi Guys, I'm needing to decide which library to use for named entity recognition. I've used spaCy, which works well, but I need a library that allows me to categorize entities and also sub-entities. Has anyone done something similar? I mean, where the same word can be more than one entity. spaCy offers the SpanCat pipeline, which theoretically allows this, but l've had trouble creating the training corpus. I think it's because they expect you to purchase an annotation text framework like Prodigy… |
| label | r/machinelearning |
| dataType | post |
| communityName | r/MachineLearning |
| datetime | 2024-05-18 |
| username_encoded | Z0FBQUFBQm5Lakwzb1Y1MFAxSWNrNHIycUFWbUVKTXRTNDhoSURLdVFSNXJsenpiaHFfR3ctcFdOWVJzOFBhMkw0VnFiNHVZNlItblR5bkp2T3MyR2JwbkROb0prVFU2d2c9PQ== |
| url_encoded | Z0FBQUFBQm5Lak9ITXJaNFlhVHhvV3dUREgxbjdYTnVUajd0NndKU3ZOTTNMWW8wM2Q5aVhVZUtvbDI5TElmSVByX1hKZmJZLTAyUXFuREdpQ3FwM19KTXZvb2p4bjJvSjYyWFhxd3duWm5vYTkzbTBSS1lIWUxtOTlHQUJHVmlic2U3OElRbElZUThXU3lMQVdaZkx6MG8xeFppdHdvWFBBa0Q1a203YlFjVjlHelhvQjdxVG5zS2k0b3pzSHBiZWotYkdQanc5Q0wwOU1xcTRGTTFYeGdCTUphTkNlcmJEZz09 |
Raw Record
{
"text": "Hi Guys, I'm needing to decide which library to use for named entity recognition. I've used spaCy, which works well, but I need a library that allows me to categorize entities and also sub-entities. Has anyone done something similar? I mean, where the same word can be more than one entity. spaCy offers the SpanCat pipeline, which theoretically allows this, but l've had trouble creating the training corpus. I think it's because they expect you to purchase an annotation text framework like Prodigy.",
"label": "r/machinelearning",
"dataType": "post",
"communityName": "r/MachineLearning",
"datetime": "2024-05-18",
"username_encoded": "Z0FBQUFBQm5Lakwzb1Y1MFAxSWNrNHIycUFWbUVKTXRTNDhoSURLdVFSNXJsenpiaHFfR3ctcFdOWVJzOFBhMkw0VnFiNHVZNlItblR5bkp2T3MyR2JwbkROb0prVFU2d2c9PQ==",
"url_encoded": "Z0FBQUFBQm5Lak9ITXJaNFlhVHhvV3dUREgxbjdYTnVUajd0NndKU3ZOTTNMWW8wM2Q5aVhVZUtvbDI5TElmSVByX1hKZmJZLTAyUXFuREdpQ3FwM19KTXZvb2p4bjJvSjYyWFhxd3duWm5vYTkzbTBSS1lIWUxtOTlHQUJHVmlic2U3OElRbElZUThXU3lMQVdaZkx6MG8xeFppdHdvWFBBa0Q1a203YlFjVjlHelhvQjdxVG5zS2k0b3pzSHBiZWotYkdQanc5Q0wwOU1xcTRGTTFYeGdCTUphTkNlcmJEZz09"
}
Entry Information
- Entry ID: 8058
- Repository: Axioma AXP
- Dataset: arrmlet/reddit_dataset_36
- Total Entries: 100,000