Row 6475
Content Data
This page contains data entry 6475 from the Axioma AXP content repository. The structured data below represents the complete record for this entry.
Hi all,
Hoping some folks here can fill in some gaps in my knowledge of multiple imputation, let me know if I'm generally using it correctly or not and whether I can use it in a specific case.
I'm in a relatively new role and working on a project where my boss wants rent predictions for *all* homes in our database. There are a few variables for where we're missing a handful of datapoints. In one case it was Zillow data for a single zip code. We found the houses in that zip code were clustered next to an adjacent zip and that said adjacent zip had similar values in years where both it and the one of interest were available. So we just substituted the values of the adjacent zip code. We have a pretty rich dataset we're working with so for most variables where we are missing a handful of observations I've been using multiple imputation.
However, there's one that is a measure of the value of the manufactured home that sits on top of a lot. It's essentially original price plus capital improvements minus depreciation. It's a fairly important variable we're using as it's a proxy for how nice a home is. Out of 18k someodd observations there are 500 and change that have either NA or implausible values for this metric. I found that among some subsets another measure of value was very close so for those subsets I substituted it. That left me with 38 NA or implausible values.
Up until this point I've been operating under two broad rules about how to use multiple imputation
1. Only use it when imputing a small number of observations compared to the population 2. There must be a good number of complete variables that can directly inform the one being imputed
Both are the case here. We have size of the home, age of the home, number of beds/baths and community that it is in (some are more upscale than others) all of which should give us a good idea of value. At the same time, we don't have variables that cover every aspect of this metric. Particularly situations where someone may have decked out a home with granite countertops and all the goodies or where there were atypically large capital improvements.
What say you people of r/datascience? Is my hacky understanding of how to use multiple imputation close enough? Can it be used in this situation?
| Field | Value |
|---|---|
| text | Hi all, Hoping some folks here can fill in some gaps in my knowledge of multiple imputation, let me know if I'm generally using it correctly or not and whether I can use it in a specific case. I'm in a relatively new role and working on a project where my boss wants rent predictions for *all* homes in our database. There are a few variables for where we're missing a handful of datapoints. In one case it was Zillow data for a single zip code. We found the houses in that zip code were clustered … |
| label | r/datascience |
| dataType | post |
| communityName | r/datascience |
| datetime | 2024-05-10 |
| username_encoded | Z0FBQUFBQm5LakwybkZZV0FtOG1VXzh6dkJXb1l5SzZkSU5NQ0J1bTVhNE5aNmpFYmVaVzNsdlVKS3FPUEtrM0ZWQWd1Rk9nb2JwVGJPSDhnZS1VZEZaWVZYYUVXUnVrbnc9PQ== |
| url_encoded | Z0FBQUFBQm5Lak9HZExoVDUyV00xU1JETEdodnNSQTdYUXBlYVluaUg0N0dzM2ZwRHhxYU56VTkxWDFRWkNFOHVtc180SEtDNUZWTElXSTlIZks4VEt2blFDZmgxdFpSVzEzNFRJSWRaRllkZU91em11Q21Fa0pRNDZ0QmltNEQxelNoMjFPR2k5dk15OEZEdmloTjRsdGUyVkxnaTQxQ3JQckFCcHc2V0p5RnM1c3FzbWNjYnNvYm5GLVRmRXpSbkJlajRxc1IwUnVjREotai1Gek1jMWhPS1dHc3Y3VUYxZz09 |
Raw Record
{
"text": "Hi all,\n\nHoping some folks here can fill in some gaps in my knowledge of multiple imputation, let me know if I'm generally using it correctly or not and whether I can use it in a specific case.\n\nI'm in a relatively new role and working on a project where my boss wants rent predictions for *all* homes in our database. There are a few variables for where we're missing a handful of datapoints. In one case it was Zillow data for a single zip code. We found the houses in that zip code were clustered next to an adjacent zip and that said adjacent zip had similar values in years where both it and the one of interest were available. So we just substituted the values of the adjacent zip code. We have a pretty rich dataset we're working with so for most variables where we are missing a handful of observations I've been using multiple imputation.\n\nHowever, there's one that is a measure of the value of the manufactured home that sits on top of a lot. It's essentially original price plus capital improvements minus depreciation. It's a fairly important variable we're using as it's a proxy for how nice a home is. Out of 18k someodd observations there are 500 and change that have either NA or implausible values for this metric. I found that among some subsets another measure of value was very close so for those subsets I substituted it. That left me with 38 NA or implausible values.\n\nUp until this point I've been operating under two broad rules about how to use multiple imputation\n\n1. Only use it when imputing a small number of observations compared to the population\n2. There must be a good number of complete variables that can directly inform the one being imputed \n\nBoth are the case here. We have size of the home, age of the home, number of beds/baths and community that it is in (some are more upscale than others) all of which should give us a good idea of value. At the same time, we don't have variables that cover every aspect of this metric. Particularly situations where someone may have decked out a home with granite countertops and all the goodies or where there were atypically large capital improvements.\n\nWhat say you people of r/datascience? Is my hacky understanding of how to use multiple imputation close enough? Can it be used in this situation?",
"label": "r/datascience",
"dataType": "post",
"communityName": "r/datascience",
"datetime": "2024-05-10",
"username_encoded": "Z0FBQUFBQm5LakwybkZZV0FtOG1VXzh6dkJXb1l5SzZkSU5NQ0J1bTVhNE5aNmpFYmVaVzNsdlVKS3FPUEtrM0ZWQWd1Rk9nb2JwVGJPSDhnZS1VZEZaWVZYYUVXUnVrbnc9PQ==",
"url_encoded": "Z0FBQUFBQm5Lak9HZExoVDUyV00xU1JETEdodnNSQTdYUXBlYVluaUg0N0dzM2ZwRHhxYU56VTkxWDFRWkNFOHVtc180SEtDNUZWTElXSTlIZks4VEt2blFDZmgxdFpSVzEzNFRJSWRaRllkZU91em11Q21Fa0pRNDZ0QmltNEQxelNoMjFPR2k5dk15OEZEdmloTjRsdGUyVkxnaTQxQ3JQckFCcHc2V0p5RnM1c3FzbWNjYnNvYm5GLVRmRXpSbkJlajRxc1IwUnVjREotai1Gek1jMWhPS1dHc3Y3VUYxZz09"
}
Entry Information
- Entry ID: 6475
- Repository: Axioma AXP
- Dataset: arrmlet/reddit_dataset_36
- Total Entries: 100,000