Row 11637
Content Data
This page contains data entry 11637 from the Axioma AXP content repository. The structured data below represents the complete record for this entry.
Hard to give a single answer, but--
> Especially when it comes to "prompt injection, adversarial attacks, profanity, off-topic content", one can home-bake some quick functions (such as regex on the less smart, and embeddings comparison for the smarter side).
Yes, you can home-bake, but if you are a business of meaningful scale, you probably really want "the best", or at least something close to it.
E.g., an airlines company really doesn't want their bot spouting racist nonsense that will end up on Twitter.
(Of course, the price has to be right, but the market is fairly competitive.)
To be clear, this is not to say that *everyone* does this, but it is an example of where strong market impetus can come from.
Now, in practice, the flip side is that the base behavior for, e.g., OAI is actually quite good, at least for trying to make sure you don't venture into NSFW. So, for this specific example (NSFW tooling), the long-term standalone market may not be huge.
> Or perhaps these tools target larger businesses (50+ employees) who would rather focus on their core expertise and offload all of this externally – like they do with grafana/prometheus/datadog, etc.
Or even startups--if you're a startup, you also should be focusing on your "core expertise" and offloading externally everything that you can.
(In practice, of course, the reason that you often don't is 1) a lot of these offerings have high minimum spend, 2) it can be harder for you to absorb integration risk, and 3) you're more likely to be able to absorb the "twitter post risk".)
| Field | Value |
|---|---|
| text | Hard to give a single answer, but-- > Especially when it comes to "prompt injection, adversarial attacks, profanity, off-topic content", one can home-bake some quick functions (such as regex on the less smart, and embeddings comparison for the smarter side). Yes, you can home-bake, but if you are a business of meaningful scale, you probably really want "the best", or at least something close to it. E.g., an airlines company really doesn't want their bot spouting racist nonsense that will end … |
| label | r/machinelearning |
| dataType | comment |
| communityName | r/MachineLearning |
| datetime | 2024-05-20 |
| username_encoded | Z0FBQUFBQm5Lakw2bmR4azBUREc4bmdIWHdWNDE3anRON2xTZ3dHQVVIWFRkeENOQTZfYnUzamZoenBGWHI3VU5GSVZrT1dNZG5JZkt2MnNTTDNndUgyLThSNFg5Y1A5akE9PQ== |
| url_encoded | Z0FBQUFBQm5Lak9KcmJLUWtHSUw3LUtpbkF0dkFhZWVQeWMxRklYZGRxeDRQQ2xmNWoybS1KeHJwU2pRSHlkc0FSWkZSUkxmRlhtV1ZIdV9mcVBMUXZFbHZCYVdZWEV6VDdfUEZ4Z255VVBmeFpIVzEwSUF0ZnA5ZzFBSGxGd0tRQ0FEWUMxWF9JRFVqWlhyWmx4RVk0SnFsMFAwUUpzWkNoNlJPRnRrS0ItUFN6UTdKdjc4STdtaDJ1TUFPMkdLSWpSUERPV0c0LXVYZ29NSUFnSkZFN09JeGsta2hJX2JJQT09 |
Raw Record
{
"text": "Hard to give a single answer, but--\n\n> Especially when it comes to \"prompt injection, adversarial attacks, profanity, off-topic content\", one can home-bake some quick functions (such as regex on the less smart, and embeddings comparison for the smarter side).\n\nYes, you can home-bake, but if you are a business of meaningful scale, you probably really want \"the best\", or at least something close to it.\n\nE.g., an airlines company really doesn't want their bot spouting racist nonsense that will end up on Twitter.\n\n(Of course, the price has to be right, but the market is fairly competitive.)\n\nTo be clear, this is not to say that *everyone* does this, but it is an example of where strong market impetus can come from.\n\nNow, in practice, the flip side is that the base behavior for, e.g., OAI is actually quite good, at least for trying to make sure you don't venture into NSFW. So, for this specific example (NSFW tooling), the long-term standalone market may not be huge.\n\n> Or perhaps these tools target larger businesses (50+ employees) who would rather focus on their core expertise and offload all of this externally – like they do with grafana/prometheus/datadog, etc.\n\nOr even startups--if you're a startup, you also should be focusing on your \"core expertise\" and offloading externally everything that you can.\n\n(In practice, of course, the reason that you often don't is 1) a lot of these offerings have high minimum spend, 2) it can be harder for you to absorb integration risk, and 3) you're more likely to be able to absorb the \"twitter post risk\".)",
"label": "r/machinelearning",
"dataType": "comment",
"communityName": "r/MachineLearning",
"datetime": "2024-05-20",
"username_encoded": "Z0FBQUFBQm5Lakw2bmR4azBUREc4bmdIWHdWNDE3anRON2xTZ3dHQVVIWFRkeENOQTZfYnUzamZoenBGWHI3VU5GSVZrT1dNZG5JZkt2MnNTTDNndUgyLThSNFg5Y1A5akE9PQ==",
"url_encoded": "Z0FBQUFBQm5Lak9KcmJLUWtHSUw3LUtpbkF0dkFhZWVQeWMxRklYZGRxeDRQQ2xmNWoybS1KeHJwU2pRSHlkc0FSWkZSUkxmRlhtV1ZIdV9mcVBMUXZFbHZCYVdZWEV6VDdfUEZ4Z255VVBmeFpIVzEwSUF0ZnA5ZzFBSGxGd0tRQ0FEWUMxWF9JRFVqWlhyWmx4RVk0SnFsMFAwUUpzWkNoNlJPRnRrS0ItUFN6UTdKdjc4STdtaDJ1TUFPMkdLSWpSUERPV0c0LXVYZ29NSUFnSkZFN09JeGsta2hJX2JJQT09"
}
Entry Information
- Entry ID: 11637
- Repository: Axioma AXP
- Dataset: arrmlet/reddit_dataset_36
- Total Entries: 100,000