Row 9563
Content Data
This page contains data entry 9563 from the Axioma AXP content repository. The structured data below represents the complete record for this entry.
Hi, I have a tabular dataset, of which some are labelled and a large portion is unlabelled. I'm trying to minimize the log-loss on the unlabelled data so overfitting on it would be perfectly fine. What would be the best approach? I tried pseudo-labels (predicting the unlabelled data and adding the most confident samples to the training data) but it made almost no difference on the test loss.
Plus, I know the results (as the overall log-loss value) of a couple of predictions on this unlabelled dataset. Any way to utilize that?
| Field | Value |
|---|---|
| text | Hi, I have a tabular dataset, of which some are labelled and a large portion is unlabelled. I'm trying to minimize the log-loss on the unlabelled data so overfitting on it would be perfectly fine. What would be the best approach? I tried pseudo-labels (predicting the unlabelled data and adding the most confident samples to the training data) but it made almost no difference on the test loss. Plus, I know the results (as the overall log-loss value) of a couple of predictions on this unlabelled d… |
| label | r/machinelearning |
| dataType | comment |
| communityName | r/MachineLearning |
| datetime | 2024-05-20 |
| username_encoded | Z0FBQUFBQm5Lakw1ZlphWUFxREVoVUxNTk5NYlNFTnJERzU1V2hEYVhaNUNhelYzanI3SFI3OGhVeVdwT2hpOGRBZmo4OXBldlljS3pOMmg4cVdkVEpMYnprQm41VFZTZEE9PQ== |
| url_encoded | Z0FBQUFBQm5Lak9JWlFiM1ZoWkxMTVkyNGlqa1EwSlNqckdjUk5KeG5ZWThMeWxiTGtPZEU3dlMwcUhCc0VPSmhlRl9SNXBnSDZWV2NhUWxOTjRkdU5XbjZNY3V5YjhWVDNDX2xEcFY3czZWWHJHQVVPckZrOWhwOTl3SGlNdDN3RHlyWE5ERk9QM3lzWElQT2pGLThBY2tGZ0w0Y1dyUXNzSG5wbVhOMFJYSmc3REE0NXNvX25tczVwajVveWVYdVlDdlpjek5aOTJG |
Raw Record
{
"text": "Hi, I have a tabular dataset, of which some are labelled and a large portion is unlabelled. I'm trying to minimize the log-loss on the unlabelled data so overfitting on it would be perfectly fine. What would be the best approach? I tried pseudo-labels (predicting the unlabelled data and adding the most confident samples to the training data) but it made almost no difference on the test loss.\n\nPlus, I know the results (as the overall log-loss value) of a couple of predictions on this unlabelled dataset. Any way to utilize that?",
"label": "r/machinelearning",
"dataType": "comment",
"communityName": "r/MachineLearning",
"datetime": "2024-05-20",
"username_encoded": "Z0FBQUFBQm5Lakw1ZlphWUFxREVoVUxNTk5NYlNFTnJERzU1V2hEYVhaNUNhelYzanI3SFI3OGhVeVdwT2hpOGRBZmo4OXBldlljS3pOMmg4cVdkVEpMYnprQm41VFZTZEE9PQ==",
"url_encoded": "Z0FBQUFBQm5Lak9JWlFiM1ZoWkxMTVkyNGlqa1EwSlNqckdjUk5KeG5ZWThMeWxiTGtPZEU3dlMwcUhCc0VPSmhlRl9SNXBnSDZWV2NhUWxOTjRkdU5XbjZNY3V5YjhWVDNDX2xEcFY3czZWWHJHQVVPckZrOWhwOTl3SGlNdDN3RHlyWE5ERk9QM3lzWElQT2pGLThBY2tGZ0w0Y1dyUXNzSG5wbVhOMFJYSmc3REE0NXNvX25tczVwajVveWVYdVlDdlpjek5aOTJG"
}
Entry Information
- Entry ID: 9563
- Repository: Axioma AXP
- Dataset: arrmlet/reddit_dataset_36
- Total Entries: 100,000