Row 6159
Content Data
This page contains data entry 6159 from the Axioma AXP content repository. The structured data below represents the complete record for this entry.
Hello my fellow DS/stats peeps,
I am working on a new problem where I am dealing with 15 years worth of hourly data of average website clicks. On a given day, I am interested in estimating the peak volume of clicks on a website with a 95% confidence interval. The way I am going about this is by bootstrapping my data 10,000 times for each day but I am not sure if I am doing this right or it might not even be possible.
Procedure looks as follows:
* Group all Jan 1, Jan 2,… Dec 31 into daily buckets. So I have 15 years worth of hourly data for each of these days, or 360 data points (15\*24). * For a single day bucket (take Jan 1), I sample 24 values (to mimic the 24 hour day) from the 1/1 bucket to create a resampled day, store the max during each resampling. I do this process 10,000 times for each day. * At this point, I have 10,000 bootstrapped maxes for all days of the year.
This is where I get a little lost. If I take the .975 and .025 of the 10,000 bootstrapped maxes for each day, in theory these should be my 95% bands of where the max should live. When I bootstrap my max point estimate by taking the max of the 10,000 samples, it’s the same as my upper confidence band.
Am I missing something theoretical or maybe my procedure is off? I’ve never bootstrapped a max or maybe it is not something that is even recommended/possible to do.
Thanks for taking the time to reading my post!
| Field | Value |
|---|---|
| text | Hello my fellow DS/stats peeps, I am working on a new problem where I am dealing with 15 years worth of hourly data of average website clicks. On a given day, I am interested in estimating the peak volume of clicks on a website with a 95% confidence interval. The way I am going about this is by bootstrapping my data 10,000 times for each day but I am not sure if I am doing this right or it might not even be possible. Procedure looks as follows: * Group all Jan 1, Jan 2,… Dec 31 into daily buc… |
| label | r/datascience |
| dataType | post |
| communityName | r/datascience |
| datetime | 2024-05-07 |
| username_encoded | Z0FBQUFBQm5LakwySFBnVjdFMGt1eWxPaElTRnNUOUYxQWs0NEhjTEtiY2JUQlhBS2FMZC1QMGJvRWY3OEtWZUZhU2xPUXU1ZC13RXQzbU1DTy1hbGlHZXI0bUlvNXlibGhIV3YxTnNvYm56QTJlRlB4WW54MFE9 |
| url_encoded | Z0FBQUFBQm5Lak9HUUNyR3p6bnZ0Zm1VbFdKQXFFNVNlMk41VHJmVXlscmJ0UTlfckdnb2NTb0lUWWtfa3hJbzVmOVRuN2FpSDBsLTlVQS1sdC04LVZpMGZxTU9BbW9aVnN1MzNkZnpBcmxDbzhST3h4UmxZUVVkTUFwSXB4aVFla2J0XzBBOExPRURpOGNqU0dhUHk3TkpfdWhWNUVrbUtWdjgwMDNRaEx5X3VKUVhSSU5aanJNRUxiZUZpVkVVMUZBTjI5Z0JPelEx |
Raw Record
{
"text": "Hello my fellow DS/stats peeps,\n\nI am working on a new problem where I am dealing with 15 years worth of hourly data of average website clicks. On a given day, I am interested in estimating the peak volume of clicks on a website with a 95% confidence interval. The way I am going about this is by bootstrapping my data 10,000 times for each day but I am not sure if I am doing this right or it might not even be possible.\n\nProcedure looks as follows:\n\n* Group all Jan 1, Jan 2,… Dec 31 into daily buckets. So I have 15 years worth of hourly data for each of these days, or 360 data points (15\\*24).\n* For a single day bucket (take Jan 1), I sample 24 values (to mimic the 24 hour day) from the 1/1 bucket to create a resampled day, store the max during each resampling. I do this process 10,000 times for each day.\n * At this point, I have 10,000 bootstrapped maxes for all days of the year.\n\nThis is where I get a little lost. If I take the .975 and .025 of the 10,000 bootstrapped maxes for each day, in theory these should be my 95% bands of where the max should live. When I bootstrap my max point estimate by taking the max of the 10,000 samples, it’s the same as my upper confidence band.\n\nAm I missing something theoretical or maybe my procedure is off? I’ve never bootstrapped a max or maybe it is not something that is even recommended/possible to do.\n\n \nThanks for taking the time to reading my post!",
"label": "r/datascience",
"dataType": "post",
"communityName": "r/datascience",
"datetime": "2024-05-07",
"username_encoded": "Z0FBQUFBQm5LakwySFBnVjdFMGt1eWxPaElTRnNUOUYxQWs0NEhjTEtiY2JUQlhBS2FMZC1QMGJvRWY3OEtWZUZhU2xPUXU1ZC13RXQzbU1DTy1hbGlHZXI0bUlvNXlibGhIV3YxTnNvYm56QTJlRlB4WW54MFE9",
"url_encoded": "Z0FBQUFBQm5Lak9HUUNyR3p6bnZ0Zm1VbFdKQXFFNVNlMk41VHJmVXlscmJ0UTlfckdnb2NTb0lUWWtfa3hJbzVmOVRuN2FpSDBsLTlVQS1sdC04LVZpMGZxTU9BbW9aVnN1MzNkZnpBcmxDbzhST3h4UmxZUVVkTUFwSXB4aVFla2J0XzBBOExPRURpOGNqU0dhUHk3TkpfdWhWNUVrbUtWdjgwMDNRaEx5X3VKUVhSSU5aanJNRUxiZUZpVkVVMUZBTjI5Z0JPelEx"
}
Entry Information
- Entry ID: 6159
- Repository: Axioma AXP
- Dataset: arrmlet/reddit_dataset_36
- Total Entries: 100,000