Row 86490
Content Data
This page contains data entry 86490 from the Axioma AXP content repository. The structured data below represents the complete record for this entry.
Fraud Analysis Courses and other tips
Hello y'all!
I'm currently working as a data analyst for an auditing company that is starting to be proactive in their audits. The plan is to start using ML or any automated process to detect frauds or unusual transactions within the clients financial reports.
Researching I found different ways to deal with data to cluster them. KMeans, KMode, KPrototype. My data is all over the place and mostly mixed data, that's why k-prototypes might be a thing. Additionally, my data has millions of rows, but not necessarily highly dimensional, less than 50 columns/features.
I'm also aware of PyOC, which uses different methods to find outliers, mostly using only numerical fields. Whomp whomp.
I guess what I'm looking for are tips on courses or resources that I can consume to better prepare when the requests start coming. Any tips are appreciated!
Thanks!
| Field | Value |
|---|---|
| text | Fraud Analysis Courses and other tips Hello y'all! I'm currently working as a data analyst for an auditing company that is starting to be proactive in their audits. The plan is to start using ML or any automated process to detect frauds or unusual transactions within the clients financial reports. Researching I found different ways to deal with data to cluster them. KMeans, KMode, KPrototype. My data is all over the place and mostly mixed data, that's why k-prototypes might be a thing. Addit… |
| label | r/datascience |
| dataType | comment |
| communityName | r/datascience |
| datetime | 2024-05-24 |
| username_encoded | Z0FBQUFBQm5Lak1wTnkxTmZjQml5S0ZYUVdHRmh0aUhaVDdreGVUWlJQNnBmZnhxaTdyTVlhUHRPN21LTzc4NExMVW5ubFpoWUh1aWUxa1B3dVEzek1BVy05YXlKM0MzT1E9PQ== |
| url_encoded | Z0FBQUFBQm5Lak82SWU1SXBrMkJwOE9XMXZzSVI5SjdRYWlpcE10ajhmQWphcGRfTWh1aUlzdEZoNVQ2YWlUQ1ROeHk1Mklfb3B1T054cFo0VENpZkhBMzZkbVlFSjVrd1VzNEdnUVlwRmd1cUlxdnNmVFNGN2MxRm8xSGo1enA2b2lpVXh5Rkl4b2xRX29vZHViLXRQbGhfYndTc3BBQWdlOC0td3BmcTNaZUJmNDVtTHFrZlllTE1ScVd1UGpiR2xOdFZXdkRVMWtLajVrYkUxMWhJdlVXeC1xcG9XOXhWdz09 |
Raw Record
{
"text": "Fraud Analysis Courses and other tips\n\nHello y'all!\n\nI'm currently working as a data analyst for an auditing company that is starting to be proactive in their audits. The plan is to start using ML or any automated process to detect frauds or unusual transactions within the clients financial reports. \n\nResearching I found different ways to deal with data to cluster them. KMeans, KMode, KPrototype. My data is all over the place and mostly mixed data, that's why k-prototypes might be a thing. Additionally, my data has millions of rows, but not necessarily highly dimensional, less than 50 columns/features.\n\nI'm also aware of PyOC, which uses different methods to find outliers, mostly using only numerical fields. Whomp whomp. \n\nI guess what I'm looking for are tips on courses or resources that I can consume to better prepare when the requests start coming. Any tips are appreciated!\n\nThanks!",
"label": "r/datascience",
"dataType": "comment",
"communityName": "r/datascience",
"datetime": "2024-05-24",
"username_encoded": "Z0FBQUFBQm5Lak1wTnkxTmZjQml5S0ZYUVdHRmh0aUhaVDdreGVUWlJQNnBmZnhxaTdyTVlhUHRPN21LTzc4NExMVW5ubFpoWUh1aWUxa1B3dVEzek1BVy05YXlKM0MzT1E9PQ==",
"url_encoded": "Z0FBQUFBQm5Lak82SWU1SXBrMkJwOE9XMXZzSVI5SjdRYWlpcE10ajhmQWphcGRfTWh1aUlzdEZoNVQ2YWlUQ1ROeHk1Mklfb3B1T054cFo0VENpZkhBMzZkbVlFSjVrd1VzNEdnUVlwRmd1cUlxdnNmVFNGN2MxRm8xSGo1enA2b2lpVXh5Rkl4b2xRX29vZHViLXRQbGhfYndTc3BBQWdlOC0td3BmcTNaZUJmNDVtTHFrZlllTE1ScVd1UGpiR2xOdFZXdkRVMWtLajVrYkUxMWhJdlVXeC1xcG9XOXhWdz09"
}
Entry Information
- Entry ID: 86490
- Repository: Axioma AXP
- Dataset: arrmlet/reddit_dataset_36
- Total Entries: 100,000