Row 13380
Content Data
This page contains data entry 13380 from the Axioma AXP content repository. The structured data below represents the complete record for this entry.
Actually testing for normality is useless in practice. Really there are only two options
* The test will tell you that the dataset is not normal, since nothing in the real world is truely normally distributed * The test will tell you "I have too low of a p-value to confirm or deny that the dataset is normally distributed"
It is also the wrong question: You do not care at all whether you data is normally distributed. What you truely care about is "Is the dataset distributed in a way that I can use a normal distribution as a model".
Notice that for a variable to be modellable with a normal distribution, it does not have to be normal! It just has to be somewhat close. In practice, there are plenty of things you can slap a normal distribution on top of and not really loose anything.
In general, you have to distinguish modelling assumptions from tests!
I recommend the following: Just plot the empirical cumulative density function on top of the cdf of the normal distribution (use cdfs not pdfs, comparing pdfs is rarely useful). If they look similar enough, then you will be fine.
| Field | Value |
|---|---|
| text | Actually testing for normality is useless in practice. Really there are only two options * The test will tell you that the dataset is not normal, since nothing in the real world is truely normally distributed * The test will tell you "I have too low of a p-value to confirm or deny that the dataset is normally distributed" It is also the wrong question: You do not care at all whether you data is normally distributed. What you truely care about is "Is the dataset distributed in a way that I can … |
| label | r/machinelearning |
| dataType | comment |
| communityName | r/MachineLearning |
| datetime | 2024-05-20 |
| username_encoded | Z0FBQUFBQm5Lakw3UE9wTzF1ZnlzNXJUcEE2b19UcWpad1ljNmV6WXlKMWdBNmcyak9vQmIyVnZ0Sjc5YmJHUDdYN2dSRnNHSU1aVG9YbkdveFJvMXdraWJDOGFPUU1ZUnc9PQ== |
| url_encoded | Z0FBQUFBQm5Lak9LUTVZNzJkMk54akZ5WDl4dnYtQmsxRDBBX3pkNkp1UFdJTjZxdkVhcWh0MUtQT1BkeEUwNXhnUFczT3pmMW5PbWJaVmN5VUtSbU02QkI4Nm9VeE9UR2EweGZlanFCT3BJRDdtYjZ2YUNCY05BNzVhd2ZmNm55SDVjMjl1R0w1ZFBpY2RKVUJhTHJCb192ekdjMkRaam1CaEFyWlR1bmdRRVNqbnVMeVFISlllSkVQTHdSRFpxeTQ2YnMwY0phYUJuRHNuckNPbzJOT3IwSFQwaW9FRGhZSzgyUXZEMlhiaEtWMG12Mk1BT1ozTT0= |
Raw Record
{
"text": "Actually testing for normality is useless in practice. Really there are only two options\n\n* The test will tell you that the dataset is not normal, since nothing in the real world is truely normally distributed\n* The test will tell you \"I have too low of a p-value to confirm or deny that the dataset is normally distributed\"\n\nIt is also the wrong question: You do not care at all whether you data is normally distributed. What you truely care about is \"Is the dataset distributed in a way that I can use a normal distribution as a model\".\n\nNotice that for a variable to be modellable with a normal distribution, it does not have to be normal! It just has to be somewhat close. In practice, there are plenty of things you can slap a normal distribution on top of and not really loose anything.\n\nIn general, you have to distinguish modelling assumptions from tests!\n\nI recommend the following: Just plot the empirical cumulative density function on top of the cdf of the normal distribution (use cdfs not pdfs, comparing pdfs is rarely useful). If they look similar enough, then you will be fine.",
"label": "r/machinelearning",
"dataType": "comment",
"communityName": "r/MachineLearning",
"datetime": "2024-05-20",
"username_encoded": "Z0FBQUFBQm5Lakw3UE9wTzF1ZnlzNXJUcEE2b19UcWpad1ljNmV6WXlKMWdBNmcyak9vQmIyVnZ0Sjc5YmJHUDdYN2dSRnNHSU1aVG9YbkdveFJvMXdraWJDOGFPUU1ZUnc9PQ==",
"url_encoded": "Z0FBQUFBQm5Lak9LUTVZNzJkMk54akZ5WDl4dnYtQmsxRDBBX3pkNkp1UFdJTjZxdkVhcWh0MUtQT1BkeEUwNXhnUFczT3pmMW5PbWJaVmN5VUtSbU02QkI4Nm9VeE9UR2EweGZlanFCT3BJRDdtYjZ2YUNCY05BNzVhd2ZmNm55SDVjMjl1R0w1ZFBpY2RKVUJhTHJCb192ekdjMkRaam1CaEFyWlR1bmdRRVNqbnVMeVFISlllSkVQTHdSRFpxeTQ2YnMwY0phYUJuRHNuckNPbzJOT3IwSFQwaW9FRGhZSzgyUXZEMlhiaEtWMG12Mk1BT1ozTT0="
}
Entry Information
- Entry ID: 13380
- Repository: Axioma AXP
- Dataset: arrmlet/reddit_dataset_36
- Total Entries: 100,000