Row 8477
Content Data
This page contains data entry 8477 from the Axioma AXP content repository. The structured data below represents the complete record for this entry.
I'm following some tutorials on doing some linear regression and as I was building my notebook, I'm working on outlier detection and amongst the techniques described for doing outlier detection, one of them involved calculating the Standard Deviation, but for this I need to know if my columns are of Guassian distribution. I'm aware that there are different techniques like:
* Histograms * KDE Plot * Q-Q Plot * Kolomogorov-Smirnov Test * Shapiro-Wilk Test * D'Agostino and Pearson's Test
And I bet there are a few more as well. So what is the best one to use? I guess Histograms just give a clue but do not show the real intention. What is the standard practice to identify if the dataset is Guassian or not?
| Field | Value |
|---|---|
| text | I'm following some tutorials on doing some linear regression and as I was building my notebook, I'm working on outlier detection and amongst the techniques described for doing outlier detection, one of them involved calculating the Standard Deviation, but for this I need to know if my columns are of Guassian distribution. I'm aware that there are different techniques like: * Histograms * KDE Plot * Q-Q Plot * Kolomogorov-Smirnov Test * Shapiro-Wilk Test * D'Agostino and Pearson's Test And I be… |
| label | r/machinelearning |
| dataType | post |
| communityName | r/MachineLearning |
| datetime | 2024-05-19 |
| username_encoded | Z0FBQUFBQm5Lakw0VHUwNWNaZnd3aEREMTdKdjUzQ01TaVYxQ2YwWXZ1UV85X3hlb0pSdlpQay1xQm44ZmpfaS1WR3VmTk5uRHJDVE1UTXFlV2E4WDFzYXU2XzNBblc0X3ZNSVJZVXBkQ1JjOTJERUhSaktZYnM9 |
| url_encoded | Z0FBQUFBQm5Lak9ITnpMQUxDWEljMVg5azZhQWxhUnRpeVBPMHluLWc0Tkk0VmxrWnZiVlNaT1BXUWM4UWx6eFdIdWkzTUc5UUxmbW9jWXZqQ1R6S0xWcGpQeDh6THBQcXhYZ0U1MVY0VThOSjRlS0hmaURjOG5lYk1EaVI2cVdsWXhwMmRtTFFDM3IySnlEeE0yNzNtMks0TUQwZS1iT3hidlBOTnMxMjFKUzZKWXJjVlpjZkVQMUZMSHhjdHhra1ZFdG5tamlCX3JjdDQwbFZzdVcwMnVMYTB1bFhNakZhUT09 |
Raw Record
{
"text": "I'm following some tutorials on doing some linear regression and as I was building my notebook, I'm working on outlier detection and amongst the techniques described for doing outlier detection, one of them involved calculating the Standard Deviation, but for this I need to know if my columns are of Guassian distribution. I'm aware that there are different techniques like:\n\n* Histograms\n* KDE Plot\n* Q-Q Plot\n* Kolomogorov-Smirnov Test\n* Shapiro-Wilk Test\n* D'Agostino and Pearson's Test\n\nAnd I bet there are a few more as well. So what is the best one to use? I guess Histograms just give a clue but do not show the real intention. What is the standard practice to identify if the dataset is Guassian or not?",
"label": "r/machinelearning",
"dataType": "post",
"communityName": "r/MachineLearning",
"datetime": "2024-05-19",
"username_encoded": "Z0FBQUFBQm5Lakw0VHUwNWNaZnd3aEREMTdKdjUzQ01TaVYxQ2YwWXZ1UV85X3hlb0pSdlpQay1xQm44ZmpfaS1WR3VmTk5uRHJDVE1UTXFlV2E4WDFzYXU2XzNBblc0X3ZNSVJZVXBkQ1JjOTJERUhSaktZYnM9",
"url_encoded": "Z0FBQUFBQm5Lak9ITnpMQUxDWEljMVg5azZhQWxhUnRpeVBPMHluLWc0Tkk0VmxrWnZiVlNaT1BXUWM4UWx6eFdIdWkzTUc5UUxmbW9jWXZqQ1R6S0xWcGpQeDh6THBQcXhYZ0U1MVY0VThOSjRlS0hmaURjOG5lYk1EaVI2cVdsWXhwMmRtTFFDM3IySnlEeE0yNzNtMks0TUQwZS1iT3hidlBOTnMxMjFKUzZKWXJjVlpjZkVQMUZMSHhjdHhra1ZFdG5tamlCX3JjdDQwbFZzdVcwMnVMYTB1bFhNakZhUT09"
}
Entry Information
- Entry ID: 8477
- Repository: Axioma AXP
- Dataset: arrmlet/reddit_dataset_36
- Total Entries: 100,000