Row 27818

Row ID: 27818 | Dataset Entry | Axioma AXP Content Repository

Content Data

This page contains data entry 27818 from the Axioma AXP content repository. The structured data below represents the complete record for this entry.

**TL;DR:** LLMs can be used as active learning components because they are good at finding difficult or diverse examples - even outperforming few-shot learning methods.

​

While LLMs such as GPT-4 are commonly used in the training process of smaller BERT-like models (pseudo-labelling or data augmentation), we wondered if they could also be used as active learning agents. In contrast to conventional active learning, we see the opportunity that LLMs are really good at capturing the diversity and difficulty of examples and do not face a cold start problem. Conventional active learning often requires many seed instances, which makes it somewhat unattractive for many tasks where BERT models already achieve good performance in few-shot scenarios.

​

We experiment with different LLMs as active learning components and indeed show that they can significantly improve performance in few-shot scenarios:

[Few-shot results \(32 instances\) with different LLMs on GLUE tasks.](https://preview.redd.it/eyhmz4u0wr1d1.jpg?width=2306&format=pjpg&auto=webp&s=1d7a35b6e3c14d1c094283e90a9155fec3337bf1)

We also show that active learning with GPT-4 can outperform the few-shot learning method SetFit:

[Comparison on AGNews \(32 instances\)](https://preview.redd.it/ji0a0nj8wr1d1.jpg?width=1651&format=pjpg&auto=webp&s=19a75253d7e0c966563af21a0de86b63b1d84387)

**Paper:** [http://arxiv.org/abs/2405.10808](http://arxiv.org/abs/2405.10808)

FieldValue
text **TL;DR:** LLMs can be used as active learning components because they are good at finding difficult or diverse examples - even outperforming few-shot learning methods. ​ While LLMs such as GPT-4 are commonly used in the training process of smaller BERT-like models (pseudo-labelling or data augmentation), we wondered if they could also be used as active learning agents. In contrast to conventional active learning, we see the opportunity that LLMs are really good at capturing the diversi…
label r/machinelearning
dataType post
communityName r/MachineLearning
datetime 2024-05-21
username_encoded Z0FBQUFBQm5Lak1FRWFjemJYNjhvNVM3VHNNaUM2YWVZZ19hcm8weHNiSThNZ1YzdTRqT1JHT2dYNlhpUm5EWWZ1bFlSTlBGemtJWmpMN0JEUHlNYl9ubmZVbEVubVhHZ0E9PQ==
url_encoded Z0FBQUFBQm5Lak9UOUszeWtKT0hmRWJMV3luQkFVUDNqOGJlTU55RkJhNk03ekFZbDEwaWtJalBLeVdYYVZCLWFLVjYyaVFyVlRtZk1NUVVUaUtRYVdtNDR1M3Fja05VRlpRZlhqd3JOcV9aV3VGS1k0dkc2bENESjUxVHp5STdRdWtPMlZNWW42UkpZZkVsaXpBTktDQTJONDE2OS1tTmppbkNZQ0pTMXp1eDRtbGp1ampzSU0tUzE3YmZFWHA2MzNlYktLSkgyOGo0

Raw Record

{
  "text": "**TL;DR:** LLMs can be used as active learning components because they are good at finding difficult or diverse examples - even outperforming few-shot learning methods.\n\n​\n\nWhile LLMs such as GPT-4 are commonly used in the training process of smaller BERT-like models (pseudo-labelling or data augmentation), we wondered if they could also be used as active learning agents. In contrast to conventional active learning, we see the opportunity that LLMs are really good at capturing the diversity and difficulty of examples and do not face a cold start problem. Conventional active learning often requires many seed instances, which makes it somewhat unattractive for many tasks where BERT models already achieve good performance in few-shot scenarios.\n\n​\n\nWe experiment with different LLMs as active learning components and indeed show that they can significantly improve performance in few-shot scenarios:\n\n[Few-shot results \\(32 instances\\) with different LLMs on GLUE tasks.](https://preview.redd.it/eyhmz4u0wr1d1.jpg?width=2306&format=pjpg&auto=webp&s=1d7a35b6e3c14d1c094283e90a9155fec3337bf1)\n\nWe also show that active learning with GPT-4 can outperform the few-shot learning method SetFit:\n\n[Comparison on AGNews \\(32 instances\\)](https://preview.redd.it/ji0a0nj8wr1d1.jpg?width=1651&format=pjpg&auto=webp&s=19a75253d7e0c966563af21a0de86b63b1d84387)\n\n**Paper:** [http://arxiv.org/abs/2405.10808](http://arxiv.org/abs/2405.10808)",
  "label": "r/machinelearning",
  "dataType": "post",
  "communityName": "r/MachineLearning",
  "datetime": "2024-05-21",
  "username_encoded": "Z0FBQUFBQm5Lak1FRWFjemJYNjhvNVM3VHNNaUM2YWVZZ19hcm8weHNiSThNZ1YzdTRqT1JHT2dYNlhpUm5EWWZ1bFlSTlBGemtJWmpMN0JEUHlNYl9ubmZVbEVubVhHZ0E9PQ==",
  "url_encoded": "Z0FBQUFBQm5Lak9UOUszeWtKT0hmRWJMV3luQkFVUDNqOGJlTU55RkJhNk03ekFZbDEwaWtJalBLeVdYYVZCLWFLVjYyaVFyVlRtZk1NUVVUaUtRYVdtNDR1M3Fja05VRlpRZlhqd3JOcV9aV3VGS1k0dkc2bENESjUxVHp5STdRdWtPMlZNWW42UkpZZkVsaXpBTktDQTJONDE2OS1tTmppbkNZQ0pTMXp1eDRtbGp1ampzSU0tUzE3YmZFWHA2MzNlYktLSkgyOGo0"
}

Entry Information