{"id":20115,"date":"2012-07-09T11:41:21","date_gmt":"2012-07-09T17:41:21","guid":{"rendered":"http:\/\/rankexploits.com\/musings\/?p=20115"},"modified":"2012-07-12T07:59:46","modified_gmt":"2012-07-12T13:59:46","slug":"methods-suppose-we-had-5000-proxies","status":"publish","type":"post","link":"https:\/\/rankexploits.com\/musings\/2012\/methods-suppose-we-had-5000-proxies\/","title":{"rendered":"Methods: Suppose we had 5000 proxies."},"content":{"rendered":"<p>I continue to be rather interested in the possible ways one might create a reconstruction of historic temperatures based on a time series of temperatures T measured over a finite calibration period with N years and a finite number of proxies M that are thought to respond temperature.   There appear to be an assortment of methods which vary in at least the following ways:<br \/>\n1) The method regressing proxy response and temperature during the calibration period,<br \/>\n2) the method selecting proxies in the first place,<br \/>\n3) the method of down-selecting proxies from an initial selection,<br \/>\n4) the method of calculating the reconstructed historic temperature based on the down-selected proxies.<\/p>\n<p>We&#8217;ve flung around a bit of math to discuss now various methods might be biased, but I finally decided just to create a script that will permit me to compare results of different &#8220;method of calculating the reconstructed historic temperature based on the down-selected proxies&#8221; (i.e. 4) and seeing how the bias and noisiness of each  method can be affected by the method of &#8221; own-selecting proxies from an initial selection&#8221;.    <\/p>\n<p>I think the examples I&#8217;ve chosen are illustrative&#8211; but readers should be cautioned that these were chose to <i>highlight<\/i> biases and noisiness that might occur along with the differences bewteen the various methods.   Figuring out how the potential biases might affect a <i>honest to goodness<\/i> reconstruction would require numerical experiments. (And of course any estimates of potential errors in a reconstruction would be based on the assumption that proxies really do respond to temperature, that that temperature response persist in the historic period and so on.)<\/p>\n<p>With this in mind, I&#8217;m going to give a cursory description of the methods.<\/p>\n<p><b>1) The method regressing proxy response and temperature during the calibration period<\/b><\/p>\n<p>All methods will use a forward equation fitting some proxy value P to <em>some<\/em> temperature, T. Generically, it will be assumed that the <i>expected value<\/i> of a proxy variable &#8220;P&#8221; at location &#8216;i&#8217; is a linear function of temperature anomaly. That is: <\/p>\n<p><equation><eqnumber>(1)<\/eqnumber>$latex \\displaystyle E[P_{i}](T) = \\lambda_{i} E[T]  $ <\/equation><br \/>\nwhere $latex \\lambda_{i} $ is a function of proxy &#8216;i&#8217;.  T is some temperature.  <\/p>\n<p>In &#8220;the litrachure&#8221; some fit proxy values to a regional value (i.e. Northern Hemispheric Temperature, or Southern Hemispheric Temperature) others fit proxy values to an estimate of the temperature near the proxy itself. This choice can affect potential bias and noisiness, so to explores this, my synthetic experiments will use two types of proxies called &#8220;global&#8221; and &#8220;local&#8221;.  <\/p>\n<p>M &#8220;Global&#8221; proxies $latex P_{G,i} $ will be generated from a &#8216;known&#8217; global temperature baselined to the calibration period  (which is the item we really wish to reconstruct) using an equation of this form:<\/p>\n<p><equation><eqnumber>(1G)<\/eqnumber>$latex \\displaystyle P_{G,i}(T_{G}) = \\lambda_{G,i} T_{G} + s_T \\sqrt{1-\\lambda_{G,i}^2 } w_{G,i}  $ <\/equation> <\/p>\n<p>where T_{G} represents the &#8216;global&#8217; temperature which will have an imposed known value and  $latex w_{G,i} $ is gaussian white noise with standard deviation 1 and  s_T is the sample standard deviation of the temperature over the calibration period.  In the current exercise, our goal is to examine how well different methods reproduce this global temperature outside the calibration periods so this temperature will also be referred to as the &#8220;Target&#8221; temperature. <\/p>\n<p>In the current synthetic experiments <i>all<\/I> global proxies will share the same value of $latex \\lambda_{G,i} = \\lambda_{G} $ which will be 0.25 computed over the <i>calibration period<\/I> of N years. <\/p>\n<p>The second type of proxy will be generated based on &#8220;local&#8221; temperatures using a very similar equation.<\/p>\n<p><equation><eqnumber>(1L)<\/eqnumber>$latex \\displaystyle P_{L,i}(T_{L,i}) = \\lambda_{L,i} T_{L,i} + s_T \\sqrt{1-\\lambda_{L,i}^2 } w_{L,i}  $ <\/equation> <\/p>\n<p>where $latex T_{L,i} $ represents the &#8216;local&#8217; temperature at proxy location &#8216;i&#8217; and  $latex w_{L,i} $ is gaussian white noise with standard deviation 1.   (Note: the sample standard deviation of these proxies will equal the sample standard deviation of the local temperature.  However: These s.d. at proxy &#8216;i&#8217; <i>will not<\/i> match that at proxy &#8216;j&#8217; This will be important later.  I would also suggest that this is the <i>right<\/i> way to weight for what I call &#8220;method 1&#8221; later on.  An alternate way that I think may sometimes be done would be totally c*appy.)<\/p>\n<p>The local temperature $latex T_{L,i} $ will be generated such that <em>on average<\/em>, the correlation between the <i>local<\/i> temperature $latex T_{L,i} $ and the global (or target) temperature $latex T_{G} $ is 0.7.  (Note 0.25\/0.7 = 0.357.  However, each local temperature will also differ from the global (or target) temperature.    I think it&#8217;s useful to clarify this with this graph showing the synthetic temperature that will be used:<\/p>\n<p><a href=\"http:\/\/rankexploits.com\/musings\/wp-content\/uploads\/2012\/07\/SyntheticTemperatures.png\"><img loading=\"lazy\" decoding=\"async\" src=\"http:\/\/rankexploits.com\/musings\/wp-content\/uploads\/2012\/07\/SyntheticTemperatures-500x500.png\" alt=\"\" title=\"SyntheticTemperatures\" width=\"500\" height=\"500\" class=\"aligncenter size-medium wp-image-20128\" srcset=\"https:\/\/rankexploits.com\/musings\/wp-content\/uploads\/2012\/07\/SyntheticTemperatures-500x500.png 500w, https:\/\/rankexploits.com\/musings\/wp-content\/uploads\/2012\/07\/SyntheticTemperatures-300x300.png 300w, https:\/\/rankexploits.com\/musings\/wp-content\/uploads\/2012\/07\/SyntheticTemperatures.png 1008w\" sizes=\"auto, (max-width: 500px) 100vw, 500px\" \/><\/a><\/p>\n<p>The global (i.e. target) temperature is illustrated in red. This temperature will be a constant for 80 years and then exhibit a ramp up for 80 years. Each expected value of local temperature at a proxy will have a qualitatively similar shape, but with a different trend during the final 80 years; these are shown with dashed blue lines. The actual temperature at each proxy will be generated by adding noise to the expected value&#8211; these actual temperatures are showing in grey.  The point of all of this is that the average temperature at proxy &#8216;i&#8217; $latex T_{L,i} $ will differ from the average temperature at proxy &#8216;j&#8217; when i\u00e2\u2030\u00a0j.<\/p>\n<p>Before proceeding: In case you haven&#8217;t guessed, the final 80 years will be the calibration periods; the first 80 years will be the reconstruction period.  <\/p>\n<p>In the current synthetic experiments each proxies, i, will be assigned <em>unique<\/em> values of $latex  \\lambda_{L,i} $.  These values will be drawn from a Normal distribution with mean $latex  E[\\lambda_{L}] = 0.357  $ and standard deviation of $latex \\sigma_{L} =0.1$.  <\/p>\n<p>Moreover, to show a potential bias in some of the methods of creating proxies, I will assign the $latex  \\lambda_{L,i} $ such that they are negatively correlated with the <del datetime=\"2012-07-09T21:25:22+00:00\">average<\/del> temperature trend for a at proxy &#8216;i&#8217; during the calibration period set to -0.5. What this means is that in the figure above, the proxy whose temperature corresponds to the lowest dark blue trace is <i>likely<\/i> to have a <del datetime=\"2012-07-09T21:25:22+00:00\">higher<\/del> lower than average value of $latex  \\lambda_{L,i} $ while the proxy whose temperature corresponds to the highest blue trace is <i>likely<\/i> to have the <del datetime=\"2012-07-09T21:25:22+00:00\">lowest<\/del> highest value of  $latex  \\lambda_{L,i} $.  <font color=\"blue\">Language corrected because  I forgot the script sets the correlation based on the <em>trend<\/em>, not the value during the reconstruction period. See blog comments. <\/font><\/p>\n<p>Given the distribution of $latex \\lambda_{G} $, $latex  \\lambda_{L,i} $, $latex  T_{G} $, $latex T_{L,i} $  $latex P_{G,i} $  and $latex P_{L,i} $ above, I will use three basic methods of estimating the temperature in during the first 80 &#8216;years&#8217; based on calibrating proxies to temperature over the final 80 &#8216;years&#8217; in the figure above.  Each basic method will then be &#8216;tweaked&#8217;! <\/p>\n<p> The basic methods are:<\/p>\n<ol>\n<li>Method I: Compute reconstructed temperature using $latex T_{rec,I} = \\sum P_i\/ \\sum m_i $ where the sums are over &#8220;M&#8221; proxies $latex m_i $ are the best fit coefficients for $latex P_i = m_i T_i $, with $latex (P_i,T_i)  $ from the calibration period (i.e. final 80 years.)  This method can be applied with either local or global proxies fit to either  $latex T_i $. (I&#8217;ll show global proxies fit to global temperature and local proxies fit to local temperature.) <\/li>\n<li>Method II: Compute reconstructionusing  $latex T_{rec,II} = \\sum [ P_i\/ m_i ] $ Here, the $latex m_i $ are exactly as above. <\/li>\n<li>Method III :Compute reconstruction using  $latex T_{rec,III} = \\sum [ P_i]\/ m_G $ where  m_Gi is computed by fitting $latex \\sum [ P_i]=  m_G  T_G  $.  That is: The we compute sum over all proxies and the find the best fit trend with the global temperature.  <\/li>\n<\/ol>\n<p><b>Results!<\/b><br \/>\nI&#8217;ll now show reconstructions using M=5000 proxies. This number is selected to demonstrate what we might hypothetically get if we have a a fairly short calibration periods (~80 years) but could somehow&#8211; quite miraculously&#8211; be able to obtain many, many proxies. (I&#8217;d use even more but I don&#8217;t want to take 10 minutes to generate a graph!)  Reality is we are stuck with very few proxies. But for now, my goal is to merely see whether the results of a method would converge on the correct value <I>if only<\/i> we could get enough decent proxies.  (Bo will want me to look at fewer proxies and I can do that afterwards.)<\/p>\n<p>I am going to let most images speak for themselves. (I&#8217;ll admit some of the later ones will require some discussion&#8211; but I want to let people look at them and interpret them for themselves first.)<\/p>\n<p><b>Fit global proxies to global temperatures<\/b><br \/>\nReconstructions using Methods I-III with all trends obtained using fits to <I>global<\/i> temperature (i.e. the Target) and applying no proxies removed for low correlation:<\/p>\n<p><a href=\"http:\/\/rankexploits.com\/musings\/wp-content\/uploads\/2012\/07\/GlobalProxiesNoScreening.png\"><img loading=\"lazy\" decoding=\"async\" src=\"http:\/\/rankexploits.com\/musings\/wp-content\/uploads\/2012\/07\/GlobalProxiesNoScreening-500x500.png\" alt=\"\" title=\"GlobalProxiesNoScreening\" width=\"500\" height=\"500\" class=\"aligncenter size-medium wp-image-20146\" srcset=\"https:\/\/rankexploits.com\/musings\/wp-content\/uploads\/2012\/07\/GlobalProxiesNoScreening-500x500.png 500w, https:\/\/rankexploits.com\/musings\/wp-content\/uploads\/2012\/07\/GlobalProxiesNoScreening-300x300.png 300w, https:\/\/rankexploits.com\/musings\/wp-content\/uploads\/2012\/07\/GlobalProxiesNoScreening.png 1008w\" sizes=\"auto, (max-width: 500px) 100vw, 500px\" \/><\/a> <\/p>\n<p>Reconstructions using Methods I-III with trends obtained fitting global proxies to <i>global<\/i> temperature and eliminating those proxies whose &#8220;R&#8221; values were not significant to at the 95% level. (Note this eliminates roughly 40% of proxies.)<br \/>\n<a href=\"http:\/\/rankexploits.com\/musings\/wp-content\/uploads\/2012\/07\/GlobalProxiesScreened.png\"><img loading=\"lazy\" decoding=\"async\" src=\"http:\/\/rankexploits.com\/musings\/wp-content\/uploads\/2012\/07\/GlobalProxiesScreened-500x500.png\" alt=\"\" title=\"GlobalProxiesScreened\" width=\"500\" height=\"500\" class=\"aligncenter size-medium wp-image-20147\" srcset=\"https:\/\/rankexploits.com\/musings\/wp-content\/uploads\/2012\/07\/GlobalProxiesScreened-500x500.png 500w, https:\/\/rankexploits.com\/musings\/wp-content\/uploads\/2012\/07\/GlobalProxiesScreened-300x300.png 300w, https:\/\/rankexploits.com\/musings\/wp-content\/uploads\/2012\/07\/GlobalProxiesScreened.png 1008w\" sizes=\"auto, (max-width: 500px) 100vw, 500px\" \/><\/a><br \/>\n(Note: the rms and average difference in the legend are based on comparison between the reconstruction and the <I>target<\/I> during the reconstruction period only.)<\/p>\n<p><b>Fit to Local Temperatures<\/b><br \/>\nReconstructions using Methods I-III with trends obtained fitting local proxies to <i>local<\/i> temperature:<br \/>\n<a href=\"http:\/\/rankexploits.com\/musings\/wp-content\/uploads\/2012\/07\/LocalProxiesNoScreening.png\"><img loading=\"lazy\" decoding=\"async\" src=\"http:\/\/rankexploits.com\/musings\/wp-content\/uploads\/2012\/07\/LocalProxiesNoScreening-500x500.png\" alt=\"\" title=\"LocalProxiesNoScreening\" width=\"500\" height=\"500\" class=\"aligncenter size-medium wp-image-20149\" srcset=\"https:\/\/rankexploits.com\/musings\/wp-content\/uploads\/2012\/07\/LocalProxiesNoScreening-500x500.png 500w, https:\/\/rankexploits.com\/musings\/wp-content\/uploads\/2012\/07\/LocalProxiesNoScreening-300x300.png 300w, https:\/\/rankexploits.com\/musings\/wp-content\/uploads\/2012\/07\/LocalProxiesNoScreening.png 1008w\" sizes=\"auto, (max-width: 500px) 100vw, 500px\" \/><\/a><\/p>\n<p>(Note: I think method II is noisier than previously because now the individual proxies display a <I>range<\/i> of responsiveness to temperature.)<\/p>\n<p><a href=\"http:\/\/rankexploits.com\/musings\/wp-content\/uploads\/2012\/07\/LocalProxiesLocalScreening.png\"><img loading=\"lazy\" decoding=\"async\" src=\"http:\/\/rankexploits.com\/musings\/wp-content\/uploads\/2012\/07\/LocalProxiesLocalScreening-500x500.png\" alt=\"\" title=\"LocalProxiesLocalScreening\" width=\"500\" height=\"500\" class=\"aligncenter size-medium wp-image-20150\" srcset=\"https:\/\/rankexploits.com\/musings\/wp-content\/uploads\/2012\/07\/LocalProxiesLocalScreening-500x500.png 500w, https:\/\/rankexploits.com\/musings\/wp-content\/uploads\/2012\/07\/LocalProxiesLocalScreening-300x300.png 300w, https:\/\/rankexploits.com\/musings\/wp-content\/uploads\/2012\/07\/LocalProxiesLocalScreening.png 1008w\" sizes=\"auto, (max-width: 500px) 100vw, 500px\" \/><\/a> <\/p>\n<p>There is a lot going on in the graph above.  But I&#8217;m going to yank my fingers away from the keyboard and let people have fun looking at it and explaining the offsetting biases. That will help people understand just how pesky picking the best method to create the reconstruction might be! <\/p>\n<p>The code: <a href='http:\/\/rankexploits.com\/musings\/wp-content\/uploads\/2012\/07\/HockeyStick_Lines_July9.txt'>HockeyStick_Lines_July9<\/a><\/p>\n","protected":false},"excerpt":{"rendered":"<p>I continue to be rather interested in the possible ways one might create a reconstruction of historic temperatures based on a time series of temperatures T measured over a finite calibration period with N years and a finite number of proxies M that are thought to respond temperature. There appear to be an assortment of &hellip; <a href=\"https:\/\/rankexploits.com\/musings\/2012\/methods-suppose-we-had-5000-proxies\/\" class=\"more-link\">Continue reading <span class=\"screen-reader-text\">Methods: Suppose we had 5000 proxies.<\/span> <span class=\"meta-nav\">&rarr;<\/span><\/a><\/p>\n","protected":false},"author":2,"featured_media":0,"comment_status":"closed","ping_status":"closed","sticky":false,"template":"","format":"standard","meta":{"footnotes":""},"categories":[3],"tags":[273,420],"class_list":["post-20115","post","type-post","status-publish","format-standard","hentry","category-statistics","tag-reconstructions","tag-screening"],"_links":{"self":[{"href":"https:\/\/rankexploits.com\/musings\/wp-json\/wp\/v2\/posts\/20115","targetHints":{"allow":["GET"]}}],"collection":[{"href":"https:\/\/rankexploits.com\/musings\/wp-json\/wp\/v2\/posts"}],"about":[{"href":"https:\/\/rankexploits.com\/musings\/wp-json\/wp\/v2\/types\/post"}],"author":[{"embeddable":true,"href":"https:\/\/rankexploits.com\/musings\/wp-json\/wp\/v2\/users\/2"}],"replies":[{"embeddable":true,"href":"https:\/\/rankexploits.com\/musings\/wp-json\/wp\/v2\/comments?post=20115"}],"version-history":[{"count":0,"href":"https:\/\/rankexploits.com\/musings\/wp-json\/wp\/v2\/posts\/20115\/revisions"}],"wp:attachment":[{"href":"https:\/\/rankexploits.com\/musings\/wp-json\/wp\/v2\/media?parent=20115"}],"wp:term":[{"taxonomy":"category","embeddable":true,"href":"https:\/\/rankexploits.com\/musings\/wp-json\/wp\/v2\/categories?post=20115"},{"taxonomy":"post_tag","embeddable":true,"href":"https:\/\/rankexploits.com\/musings\/wp-json\/wp\/v2\/tags?post=20115"}],"curies":[{"name":"wp","href":"https:\/\/api.w.org\/{rel}","templated":true}]}}