{"id":15735,"date":"2011-06-20T15:58:49","date_gmt":"2011-06-20T21:58:49","guid":{"rendered":"http:\/\/rankexploits.com\/musings\/?p=15735"},"modified":"2011-06-20T15:58:49","modified_gmt":"2011-06-20T21:58:49","slug":"whats-uncorrelated-with-what-for-paulk","status":"publish","type":"post","link":"https:\/\/rankexploits.com\/musings\/2011\/whats-uncorrelated-with-what-for-paulk\/","title":{"rendered":"What&#8217;s uncorrelated with what? For PaulK"},"content":{"rendered":"<p>This post is mostly for Paul_K; he and I discussed on of my claims in my previous article <a href=\"http:\/\/rankexploits.com\/musings\/2011\/noaa-may-cooler-than-april\/\">discussing the NOAA anomaly. <\/a> Specifically, I claimed to compute two possible test variables from linear fit to observed data which in the current post I will call d*<sub>a<\/sub> and d*<sub>m<\/sub> and said the errors in these tests parameters were uncorrelated.   I probably worded it in a confusing way, and I&#8217;d like to clarify for Paul_K.  After clarification, it may turn out Paul_K has a suggestion, but for the time being, we seem to be talking at cross purposes.   To do so, I&#8217;m going to discuss some observations based on trendless synthetic data.<\/p>\n<p>Suppose I generate a series of 137 months of Gaussian white noise with standard deviation 1 and mean zero, and call that &#8220;Y<sub>i<\/sub>&#8220;.   I can now compute the mean of Y<sub>i<\/sub> in the normal way.  We know the <I>true<\/I> mean of this data should be 0.  But instead, we will get some sample mean value Y<sub>bar<\/sub>, which we recognize as <I>an error<\/i>.  That is: it&#8217;s different from 0, the known correct value if we repeat the experiment over and over and over and take an average over the ensemble of all possible experiments.  Let&#8217;s call this \u00ce\u00b5<sub>Y<\/sub>.<\/p>\n<p>Next, I can also compute the best fit trend through <\/p>\n<p>Y<sub>i<\/sub>=m i + b + u<sub>i<\/sub><\/p>\n<p>where the &#8216;i&#8217; values are (1,2,3,&#8230;. 137), m is the best fit trend, and b is the intercept at i=0, and u<sub>i<\/sub> are the residuals. The best fit values of m and b are those which minimize the sum of the squares of the 137 residuals.  <\/p>\n<p>Because I&#8217;ve generated the Y<sub>i<\/sub> as gaussian white noise, we know the <I>true<\/I> value of m for the process is m=0. Nevertheless, when I generate 137 data points, I&#8217;ll nearly always get a non-zero value for the trend based on the sample, m.  The difference between the sample value of the trend, m, and the true trend is the error in the trend.  Let&#8217;s call this \u00ce\u00b5<sub>m<\/sub>.<\/p>\n<p>My claim is that the errors \u00ce\u00b5<sub>Y<\/sub> and  \u00ce\u00b5<sub>m<\/sub> are uncorrelated.   I also make a separate claim that if we do the least squares fit to the <\/p>\n<p>Y<sub>i<\/sub>=m (i &#8211; i<sub>mean<\/sub>)  + a + u<sub>i<\/sub> <\/p>\n<p>where i<sub>mean<\/sub> is the mean value for the indices of the data. So, if the indices were 1 through 137, i<sub>mean<\/sub>=69.<\/p>\n<p>We&#8217;ll find that the value of &#8216;a&#8217; is equal to the sample mean of the Y<sub>i<\/sub>; I leave showing this as an exercise to the reader. \ud83d\ude42  But it&#8217;s obvious from this if &#8216;a&#8217; is equal to the sample mean Y<sub>i<\/sub>; then the error in a just the error \u00ce\u00b5<sub>Y<\/sub>.  So, I am now claiming the errors in the determination of &#8216;a&#8217;, the intercept in the second form of the fit is uncorrelated from the error in the determination of the trend &#8216;m&#8217;. <\/p>\n<p>To provide an empirical demonstration of this when the residuals are white noise, I created 10<sup>5<\/sup> 137 month sequences of white noise and computed a the sample mean and trend for each series. This resulted in 10<sup>5<\/sup>  pairs of trends and means.  I then took the step of normalizing with the estimate of the uncertainty based on the least squares fit. This results scaled errors with mean equal to zero and standard deviation equal to 1.  I then computed the cross correlation between the scaled error for the mean and the trend which for the most recent case tested happened to be -0.0004674765; and also made a scatter plot of the two:<br \/>\n<a href=\"http:\/\/rankexploits.com\/musings\/wp-content\/uploads\/2011\/06\/ScatterPlot_dstars.png\"><img loading=\"lazy\" decoding=\"async\" src=\"http:\/\/rankexploits.com\/musings\/wp-content\/uploads\/2011\/06\/ScatterPlot_dstars-500x500.png\" alt=\"\" title=\"ScatterPlot_dstars\" width=\"500\" height=\"500\" class=\"aligncenter size-medium wp-image-15742\" srcset=\"https:\/\/rankexploits.com\/musings\/wp-content\/uploads\/2011\/06\/ScatterPlot_dstars-500x500.png 500w, https:\/\/rankexploits.com\/musings\/wp-content\/uploads\/2011\/06\/ScatterPlot_dstars-300x300.png 300w, https:\/\/rankexploits.com\/musings\/wp-content\/uploads\/2011\/06\/ScatterPlot_dstars.png 1008w\" sizes=\"auto, (max-width: 500px) 100vw, 500px\" \/><\/a><br \/>\nThe more or less circular appearance is typical of uncorrelated data.<\/p>\n<p>Since the main purpose of this post was to clarify a point for PaulK, and I <i>think<\/i> this might be sufficient for the clarification, I&#8217;m going to stop here so we can continue our conversation on my attempting to develop a sort of &#8220;portmanteau&#8221; test involving testing the combination of the trend and the constant term from a least squares fit. I think such a test would have the properties of being a bit more difficult to cherry pick and under some circumstances can be more powerful than testing either the trend alone or the constant term alone. (There are circumstances where I think it may be  worse&#8211; specifically, if the baseline was too short in time, it can become less powerful. But testing power is something that can be done after we figure out if I&#8217;m missing something that makes my portmanteau test untenable. )<\/p>\n<p><b>R File<\/b>: This R file (<a href=\"http:\/\/rankexploits.com\/musings\/wp-content\/uploads\/2011\/06\/MonteCarlo_d_star_simpler1.txt\">MonteCarlo d_star<\/a>.) was used to generate the plot. The demonstration with AR1 is also included. So are various cryptic and obscure comments. <\/p>\n","protected":false},"excerpt":{"rendered":"<p>This post is mostly for Paul_K; he and I discussed on of my claims in my previous article discussing the NOAA anomaly. Specifically, I claimed to compute two possible test variables from linear fit to observed data which in the current post I will call d*a and d*m and said the errors in these tests &hellip; <a href=\"https:\/\/rankexploits.com\/musings\/2011\/whats-uncorrelated-with-what-for-paulk\/\" class=\"more-link\">Continue reading <span class=\"screen-reader-text\">What&#8217;s uncorrelated with what? For PaulK<\/span> <span class=\"meta-nav\">&rarr;<\/span><\/a><\/p>\n","protected":false},"author":2,"featured_media":0,"comment_status":"closed","ping_status":"closed","sticky":false,"template":"","format":"standard","meta":{"footnotes":""},"categories":[3],"tags":[],"class_list":["post-15735","post","type-post","status-publish","format-standard","hentry","category-statistics"],"_links":{"self":[{"href":"https:\/\/rankexploits.com\/musings\/wp-json\/wp\/v2\/posts\/15735","targetHints":{"allow":["GET"]}}],"collection":[{"href":"https:\/\/rankexploits.com\/musings\/wp-json\/wp\/v2\/posts"}],"about":[{"href":"https:\/\/rankexploits.com\/musings\/wp-json\/wp\/v2\/types\/post"}],"author":[{"embeddable":true,"href":"https:\/\/rankexploits.com\/musings\/wp-json\/wp\/v2\/users\/2"}],"replies":[{"embeddable":true,"href":"https:\/\/rankexploits.com\/musings\/wp-json\/wp\/v2\/comments?post=15735"}],"version-history":[{"count":0,"href":"https:\/\/rankexploits.com\/musings\/wp-json\/wp\/v2\/posts\/15735\/revisions"}],"wp:attachment":[{"href":"https:\/\/rankexploits.com\/musings\/wp-json\/wp\/v2\/media?parent=15735"}],"wp:term":[{"taxonomy":"category","embeddable":true,"href":"https:\/\/rankexploits.com\/musings\/wp-json\/wp\/v2\/categories?post=15735"},{"taxonomy":"post_tag","embeddable":true,"href":"https:\/\/rankexploits.com\/musings\/wp-json\/wp\/v2\/tags?post=15735"}],"curies":[{"name":"wp","href":"https:\/\/api.w.org\/{rel}","templated":true}]}}