Response to Keenan

Doug,
Thanks for posting your email to me. As I said, I will respond to whatever you are willing to post in public. I am taking the liberty to respond here where I can more easily use blockquotes.

Lucia,

Your recent blog post “BEST data: Trend looks statistically significant so far” includes the following statement.

Since smoothing is discussed in Keenan’s letters to the economist, I will note that this data is computed by averaging over 12 months. So, it is “smoothed” relative to monthly data.

If we have a series of n monthly values, and obtain from that a series of n/12 annual values, then we are not doing smoothing in the sense I intended; rather, we are doing aggregation. Aggregation is fine. Smoothing is problematic, and the problem is easy to understand.

Suppose that our original series is a1, a2, a3, a4, …, an, and we are taking a 3-point moving average. Denote the smoothed series by b1, b2, b3, b4, …, b n-2. Then b1 = (a1 + a2 + a3)/3, and b2 = (a2 + a3 + a4)/3, etc. Notice that both b1 and b2 depend upon a2 and a3. Hence b1 and b2 are correlated with each other. In other words, the smoothed series is more autocorrelated than the original series. That is the problem that smoothing (via a moving average) introduces.

Suppose, on the other hand, that we simply aggregate the original series, to obtain c1, c2, c3, c4, …, cn/12. Then c1 = (a1+ a2+ a3+ … + a12)/12 and c2 = (a13+ a14+ a15+ … + a24)/12, etc. Thus c1 and c2 do not depend upon the same elements of the original series. Ergo, the aggregation does not introduce auto-correlation.

I’m not entirely sure how whatever point you are making is supposed to relate to what I wrote in the context I wrote it. I was simply explaining what I did. I took unsmoothed data and stuffed it in your algorithm, computed the AICc criterion.

Moreover, since, when criticizing Muller, you recently complained about smoothing, I didn’t smooth it. I stuffed the unsmoothed data into a process, computed AICc criterion and “Presto!”

It seems to me you are now explaining that it was ok for you to smooth (by averaging) because you understand the problem introduced by smoothing and use it appropriately. In your email, you are describing this as a valid process:

  1. smoothing over ‘n’ values
  2. pulling out only every “nth
  3. give the full process of smoothing followed by culling the name of aggregation so as to distinguish it from smoothing and not culling.

Of course, many types of aggregates exist: Aggregating functions include giving the count, finding a miximum or minimum, a sum, a mode and so on. Yours happens to be an aggregate of “n” sequential things– that is you created an aggregate obtained by averaging — which is smoothing.

You are also explaining that auto-correlation is introduced by performing a moving average: I believe many of us understood this long before you explained it. Maybe your point is to explain that when you, Doug Keenan smooth, you are careful to avoid pitfalls involved in smoothing so as to avoid fooling yourself.

If so, that is not the message one would take home from the criticism you sent Richard Muller [BEST Scientific Director]; Charlotte Wickham [BEST Statistical Scientist]
Cc: James Astill; Elizabeth Muller. In that letter, you quoted Briggs and wrote:

Unless the data is measured with error, you never, ever, for no reason, under no threat, SMOOTH the series!

It seems to me that Muller made a point very similar to the one you are making now, when in his response, he wrote

“He [Keenan] is, of course, being illogical. Just because smoothing can increase the probability of our fooling ourselves doesn’t mean that we did. There is real value to smoothing data, and yes, you have to beware of the traps, but if you are then there is a real advantage to doing that.”

Maybe in future, when you post your criticisms of BEST, you can be a little more nuanced and explain the actual problem that arose when BEST’s averaged and then did something not-permissible after averaging. That is: Explain how they used averaging and then failed to compensate.

In your email to me you continue.

Your post also says this.

Doug’s claim that we might conclude the “IPCC assumption is insupportable” might be convincing to me if I bought for one second that it makes sense to use d=1 in ARIMA(3,1,0). I don’t. I think that arguments for statistical models with d=1 tend to violate the 1st law of thermodynamics.

The physical plausibility of an ARIMA(p,1,q) could be questioned, at least on long time scales; on the other hand, ARIMA(p,1,q) might be a reasonable approximation, on time scales of interest here, to a more physically-plausible process. As an analogy, Earth is approximately spherical, but if someone is drawing a map of England, it is reasonable to assume flatness. Similarly, given the shortness of the time series, ARIMA(p,1,q) could be reasonable.

First: I’m not saying “the physical plausibility of an ARIMA(p,1,q) could be questioned”, I am questioning it based on the laws of thermodynamics, heat transfer. (I have admitted to providing only a heuristic explanation. )

Second: Just because the surface of the earth can be approximated as flat even when it is spherical doesn’t mean something that is implausible at long time steps can magically become plausible at short time steps. You actually have to be able to give an explanation why an approximation might hold over a particular range.

I question ARIMA(p,1,q) for natural forcings at all time scales. Moreover, to the extent that driftless ARIMA(p,1,q) is put forward as contradicting that an apparent trend is due to anthropogenic forcings, I question d=1 even if you have decided to provide a flimsy reason to to consider “driftless ARIMA(p,1,q) a candidate over the time span your consider doesn’t mean that anyone have to buy it.

I don’t buy your argument: I consider it flimsy.

Doug, you continue

In any case, the comparison of the ARIMA(p,1,q) model and the IPCC model strongly indicates that the IPCC model is failing to explain some substantial structural variation in the data–and that is the sole purpose of the comparison. Thus the comparison gives insight, regardless of physical plausibility.

If, as you now claim, you were to limit your claim to saying that ARIMA(1,0,0) + trend is failing to explain the structure of the time series you would find few to contradict you. Many people were saying so at blogs, in journal articles and elsewhere long before you wrote your WSJ article.

However, if the “the sole purpose of the comparison” is to show the AR(1)+ trend model is “failing to explain some substantial structural variation in the data”, then the following claim in your WSJ article is rather overblown.

“… but the improved fit does tell us that until more research is done on the best assumptions to apply to global average temperature series, the IPCC’s conclusions about the significance of the temperature changes are unfounded. “

The improved fit shows research is needed? Lots of people are doing research and not because of your demonstration that some other ARIMA model fits the data better.

But beyond this, it is hardly the case that the IPCC’s conclusions about the significance of temperature changes rests solely, or even principally on the AR(1)+ trend model. Much of the IPCC’s confidence in the significance is based on non-statistical arguments– including comparisons of outcomes between models driven by GHG’s and those not. It’s true some readers here might not accept the IPCC conclusions and don’t favor modeling approaches. But showing that the AR(1)+trend model is fails to explain structural variation barely puts a dent in their argument that the trend is attributable to anthropogenic forcings because that’s not their main argument.

Moreover, that you (or many others) can find and present a model a driftless ARIMA model with a better AICc coefficient than ‘AR(1)+trend’ doesn’t overturn the IPCC’s conclusions of the significance of temperature changes. We can easily find other ARIMA models with either drift or trend (e.g. ARIMA(0,1,4)+drift) with better AICc’s than the one you highlighted in the supporting materials for your WSJ. If the argument is to pick out based on AICc, (and I see little other in your WSJ article) the IPCC’s conclusions seems to remain founded. The only criticism would be that the discussion of the AR(1)+ trend model isn’t particularly useful but we would still conclude there was either drift or trend if we look for the ARIMA(p,d,q) model with the best AICc criteria.

You, Doug, go on:

Your post further says the following.
That is: using the preliminary BEST data going back to 1800, the IPCC AR(1) blows the model favored by Doug Keenan in his Wall Street Journal out of the water. The model that says “Statistically significant warming” wins.

Update(May 25): If I use annual averaged best data, the model that wins reverses. As already promised below, I’ll be discussing other models.

In the first case, you were using seasonal (monthly) data, which cannot be directly compared like that–as your Update effectively showed.

Correct. I was using monthly data — as would be required if it were actually true that one should “never, ever, for no reason, under no threat, SMOOTH the series!”

Naturally, since the admonition to never smooth is invalid, I acknowledged the difficulty when people requested I used the smoothed — i.e. averaged data. In fact, I expect people would ask me to do so because it’s actually ok to smooth!.

You, Doug go on:

Additionally, the ARIMA model was not “favored” by me; it was solely used for comparison with the IPCC model. Indeed, the WSJ piece ended by saying that more, and difficult, research was required.

First: With regard to your current claim that the ARIMA model was “solely used for comparison with the IPCC model”, it appears that your criticism in the letter to the Economist makes a much more expansive claims for your article in the WSJ where you write

To summarize, most research on global warming relies on a statistical model that should not be used. This invalidates much of the analysis done on global warming. I published an op-ed piece in the Wall Street Journal to explain these issues, in plain English, this year.

In this, you certainly appear to be claiming that your op-ed piece in the Wall Street Journal communicates the fact that much of the analysis done on global warming is invalid!

Either– as you now claim– the only thing your article shows is that the AR(1)+trend model “is failing to explain some substantial structural variation in the data”, or your op-ed piece does something that seems to indicate that “much of the analysis on global warming [is invalid]”. (I should note that, in my opinion, the wording of your WSJ claim gives the reader the impression that the better AICc for the driftless ARIMA model does more than merely point out a structural inadequacy in the AR(1)+trend model. )

Second: I an mystified by your objection my using the verb “favor” to describe your characterization of the ARIMA(3,1,0) model over the ARIMA (1,0,0)+trend model.

You elected to use the ARIMA(3,1,0) model rather than some other as the example in the supporting materials for your WSJ article. You bring up AICc as a criteria to pick on model over another. You guide the reader to the notion that the ARIMA(3,1,0) is better by this metric. That fits the definition of “favoring” it.

At least as far as supplying the reader with information, you use ARIMA(3,1,0) model as “the” model to support this statement in your WSJ article

“A fairly elementary alternative assumption that some researchers and I have tested fits the actual temperature data better than the IPCC’s AR1 assumption?so much better that we can conclude that the IPCC’s assumption has no support. Under the alternative assumption, the data do not show a significant increase in global temperatures.”

So, even though alternatives exist that have a better fit than driftless ARIMA(3,1,0), you happened to pick this one- to highlight in your WSJ article and supporting materials. And you use the definite article “the” not “this” or “an”– which suggests to naive readers that there might only be two alternatives. But even if you’d used “this” or “an”, you focused on the existence of a particular alternative one and advise it is better than AR(1)+trend. You don’t mention the possiblity of ARIMA(0,1,4) with drift which means there is a secular warming trend– and you don’t mention this despite the fact that it’s whose AICc beats ARIMA(3,1,0).

I would say the choices you made when writing your WSJ article and defending ARIMA(p,d=1,q) constitutes your favoring ARIMA(3,1,0) both over AR(1)+trend but even over other alternatives with d=1.

But beyond that in your email to me (posted here) you defended the ARIMA(p,1,q) model against criticism that it is physically implausible thus giving the impression that you do, indeed, “favor” — that is show a preference– for this model above a number of other models.

I want to add this: I realize in your WSJ you write this,

We don’t know whether the alternative assumption itself is reasonable?other assumptions might be even better?but the improved fit does tell us that until more research is done on the best assumptions to apply to global average temperature series, the IPCC’s conclusions about the significance of the temperature changes are unfounded.

Yet, for some reason, you object to my saying that we do know whether “the” alternative referred to in your WSJ article and in your supporting materials is reasonable. What we know is that physical arguments would lead us to say that it implausible if we are going to attribute warming to natural forcings.

That is: if d=1, then it must be because the anthropogenic forcings are d=1, because “natural forcings” cannot be. Of course, if d=1, because anthropogenic forcings are driving temperature changes, this tends to support the notion that current temperature are high because ghg’s have increased. In this case your the closing sentence of your article “the IPCC’s conclusions about the significance of the temperature changes are unfounded.” is false. Because finding d=1– if that was “the” only alternative to AR(1)+ trend would indicate that the rise cannot result from natural forcings.

It may well be true that my argument isn’t based on statistics but on physics, but there is absolutely no reason why those trained in the physical sciences and who are familiar with thermodynamics cannot say that they have “issues” with someone presenting an analysis that highlights a statistical model that appears to violate physics.

To summarize, your criticisms are invalid.

Sincerely, Doug

I think it’s pretty silly to decree one’s own arguments victorious. I happen to think most of what you write is pretty weak. But that’s my opinion and I will leave it to others to decide what they think of your criticism of best, your claims that you don’t “favor” the model you highlighted, and your defense of putting forward ARIMA(p,1,q) models. I’ve said I don’t like d=1, I think averaging can be used– carefully– and I you certainly give the impression that you favor–i.e. prefer– ARIMA(3,1,0) to AR(1)+trend.

Update: 3:30 pm. Doug Keenan emailed requesting I edit to create subscripts in the first quote. I did so.

67 thoughts on “Response to Keenan”

  1. Lucia,

    Nice summary.
    I would emphasize that any statistical model which is clearly in conflict with known physical reality should automatically be excluded from consideration, even if the ‘fit’ to the data is good…. if you are really trying to understand how a process works. It seems to me that Doug is willing to look past the rather reasonable expectation that increased forcing should raise average temperature. I don’t at all understand his reasoning, and he has (AFAIK) not offered any meaningful explanation.

  2. Lucia, I think the point has been made by you and others in various threads at your blog that the ARIMA models are not appropriate, physically, for modeling the global temperature. I sometimes have problems keeping track of your arguments and those of Carrick, but I think Carrick has indicated that an ARIMA model is inappropriate at any time scale because the issues of unboundedness presented by random walks and the nonlinear nature of the physical processes effecting temperature that no amount of trend extraction, linear of nonlinear, can make amenable for the use of an ARIMA model.

    I think your arguments, Lucia, have been more directed to a piecemeal argument against the Keenan reasoning in his WSJ article. An integer d in the ARIMA model could be used to extract a linear trend, but it is not appropriate (unphysical) to use differencing to make the series stationary from a non stationary random walk series. I think what follows immediately is that the Keenan ARIMA model with an integer d (differencing) has to admit to a trend in the series under discussion.

    I think the discussions in these threads here have shown that the question of temperature modeling is far more complicated than either the Keenan WSJ article or the IPCC literature cited by Keenan have indicated. Further to the point you are making, Lucia, or at least in my view of it, is that simply providing an alternative model cannot be implied to indicate that a trend does not exist in the instrumental temperature series for the globe. (I have not read the Keenan article in the whole so I do not know what it implied and what parts of it might have been taken out of context by others who have commented on it).

    What I see that happens all too often in the MSM form of journalism is to oversimplify and take analyses/evidence a step or two too far in conclusions or conjecture and often to make an advocacy point. It can happen on all sides of the AGW issue and it is important to keep the factual evidence or extent of an analysis in a paper/article separate from the implications that might be surmised from it. All of this leads to my view that (good) blogging will continue to show the limitations and weaknesses of modern day journalism.

  3. The excerpt below would appear to give Keenan cover on his alternative claim and thus we are left with his model merely being incorrect. Now that we have had the discussions about ARIMA models and temperature maybe we can get Keenan to comment on those points.

    “Under the alternative assumption, the increase in global temperatures is not significant. We do not know, however, whether the alternative assumption itself is reasonable—other assumptions might be even better. Determining how viable the alternative assumption is would require study. There have been studies that consider other assumptions and thereby reach different conclusions about the temperature data. The IPCC report nods toward such studies, but without acknowledging that the soundness of its conclusions rests upon its choice of assumption—or that making a good choice, one that well corresponds with physical reality, requires further, difficult research.”

  4. “It may well be true that my argument isn’t based on statistics but on physics, but there is absolutely no reason why those trained in the physical sciences and who are familiar with thermodynamics cannot say that they have “issues” with someone presenting an analysis that highlights a statistical model that appears to violate physics.”

    Ouch! 🙂

  5. In discussing “smoothing” at Jeff Id’s, Doug Keenan makes the point that the process introduces autocorrelation to a time series. That is for “smoothing” as he is defining it; see the “b” instance of his “a, b, c” illustration.

    Creating an annual average out of 12 sequential monthly points (averages) doesn’t qualify as “smoothing” by Keenan’s definition (his “c” instance).

    This strikes me as a reasonable distinction. It strikes me as prudent to be very wary of smoothing as Keenan describes it, for the reasons he gives.

    Lucia, do you agree?

  6. Amac–

    In discussing “smoothing” at Jeff Id’s, Doug Keenan makes the point that the process introduces autocorrelation to a time series. That is for “smoothing” as he is defining it; see the “b” instance of his “a, b, c” illustration.

    I believe that exact explanation is in the bit I quoted above.

    When you ask me “do you agree”, I’m not sure precisely what you are asking me if I agree with.

    Moving averages and many types of smoothing introduce autocorrelation. One needs to be aware of this and know how it affects downstream computations. Muller acknowledged the latter when he responded

    “He [Keenan] is, of course, being illogical. Just because smoothing can increase the probability of our fooling ourselves doesn’t mean that we did. There is real value to smoothing data, and yes, you have to beware of the traps, but if you are then there is a real advantage to doing that.”

    I agree with Muller. Smoothing can cause difficulties, but one must be wary. In the particular case of moving averages described by Keenan, one of the difficulties is the introduction of autocorrelation. If you are asking me if I agree one must be wary of that: Yes.

    But that’s not the idea Keenan communicated using all caps and the words never and ever and so forth.

  7. Kenneth

    (I have not read the Keenan article in the whole so I do not know what it implied and what parts of it might have been taken out of context by others who have commented on it).

    To read the article in whole, you need to google “Keenan WSJ” The article itself contains very little analysis, but points to technical details at Keenan’s web site.

    You will find this in Keenan’s WSJ article

    The IPCC report nods toward such studies, but without acknowledging that the soundness of its conclusions rests upon its choice of assumption—or that making a good choice, one that well corresponds with physical reality, requires further, difficult research.

    I’d suggest Doug’s claim may rest on his not processing the contents of chapter 3 of the WG1 (which is the chapter he mentions in the WSJ article.) The appendix 3A of chapter 3 states

    “Hence, the statistical significances of REML AR1-based linear
    trends could be overestimated (Zheng and Basher, 1999;
    Cohn and Lins, 2005). Nevertheless, the results depend on the
    statistical model used, and more complex models are not as
    transparent and often lack physical realism. Indeed, long-term
    persistence models (Cohn and Lins, 2005) have not been shown
    to provide a better fit to the data than simpler models.”

    That is: The IPCC openly acknowledges that the AR1+ linear trend is not perfect, can lead to over-estimation of statistical significance. If Doug’s main point is to demonstrate this– well “Yawn”.

    Moreoover: my criticism of Keenan’s example choice (dare I use the word “favored”?) driftless ARIMA(3,1,0) is precisely that it “lacks physical realism. Doug response to this is rather flimsy.

    (Cohn and Lins doesn’t have the problem that Keenan’s choice does– I just wish they would extend what they did to include the fact that — if there is any preferred ‘deterministic trend’ in the AR4, it is the multi-model mean, not a linear trend. I don’t think I managed to get Cohn to understand what I’m on about.)

  8. “I don’t think I managed to get Cohn to understand what I’m on about.”

    I suspect some of those whom you show a need to “adjust or widen” their thinking do understand and do get your point. That part should be rather easy. The admitting might be the difficult part. I think Cohn (who really knows how to code in R) keep going back to their paper addressing only the issue of a linear trend and never dealt with the issues you raised about a non linear trend having smaller residual errors.

    I liked what Wu did in his paper that I linked in another thread on detrending a non stationary temperature series with a non linear trend. I was wondering if, after detrending in the Wu mode, one could legitimately apply an ARIMA model or actually an ARMA model (no differencing).

    Let us see what Doug Keenan says. And I did read the WSJ article right after I made my disclaimer above.

  9. Isn’t everyone making too big a deal about this quote?

    “Unless the data is measured with error, you never, ever, for no reason, under no threat, SMOOTH the series!”

    Isn’t that just signal to noise principle. I know several RF Engineers and Technicians and I think there is an analogy here. When they test a cable, they rely heavily on signal to noise ratios. If the noise is too high they don’t put on a filter. The noise may be part of the signal or the noise can help find a problem with the cable or data source. The same I assume is true for stats, if the data is bad (this is “unless the data is measured with error” part) then you need to check data and the transmission to clean up the signal/data. After you are satisfied then you can apply the filter or smooth the data. I believe SNR is an important concept in statistics too.

  10. Jeff Id (Comment #84966),
    Remember, never get Luica angry.
    Actually, that suggestion can be applied to a lot of smart people I know.. they tend not to suffer fools gladly. 🙂

  11. Kermit–
    Depends what you mean by too big a deal.

    Keenan was sent an email and asked his comments on BEST. He replied. There were some exchanges. Keenan then posted his exchange, and it was re-run at Bishop Hill. Part of Keenan’s criticism of BEST is to quote Briggs writing ““Unless the data is measured with error, you never, ever, for no reason, under no threat, SMOOTH the series!”

  12. Oh… I’m not angry. But I had some email back and forth when the WSJ article came out, and I’ve decided that having the back and forth in private email wastes my time.

    The Keenan WSJ article is in public. His comments on BEST are in public. He clearly wants to engage the comments I made at my blog– that is, what I said in public. He wants to announce that my criticisms are invalid over at Jeff’s — and to do that without himself having first put what he wrote in public.

    I would rather respond to that in public. I can’t speak for Tamino or Judy, but if Doug thinks what Tamino or Judy wrote about his BEST comments is not valid, I would like to read Doug’s rebuttal– in public. Until that time, as far as I am concerned, Doug has not supported his claim that what they wrote was invalid.

  13. Lucia,

    Oh… I’m not angry.

    .
    OK. Still, I hope you can forgive those who might imagine otherwise.

  14. SteveF–

    Still, I hope you can forgive those who might imagine otherwise.

    Sure. I can understand why someone might think otherwise.

  15. Re: lucia (Nov 1 17:30),

    Sure. I can understand why someone might think otherwise.

    That was hardly a “Hulk smash!” comment, more a relatively dispassionate dissection or vivisection.

  16. I didn’t mean to detract from your post. It didn’t read angry but it is tough love at best. Blogland is home to the geek, not the weak, tri-lambdas unite!

    haha.

    Anyway, I unfortunately agree with Lucia’s critiques. To make the claim though that temp changes might not be significant (i.e. real and detectable) based on some modeling of noise, is a stretch for the engineer in me.

  17. Jeff Id,
    “I unfortunately agree with Lucia’s critiques.”
    What is unfortunate about that?

  18. FWIW monthly data is smoothed daily data is smoothed hourly data is . . ., well as the late and sorely missed Jon Swift would say

    “So scientists divide, a record
    Into smaller ones that we display,
    And these have smaller yet that cite ’em,
    And so proceed ad infinitum.”

    The monthly records should have even stronger autocorrelation.

  19. Eli–
    Monthly records do have stronger auto-correlation. 🙂

    Jeff–
    I didn’t think you were detracting from my post. I just thought I would mention I’m not angry.

  20. Eli there is a difference between “down sampling” versus “smoothing”. In “down sampling’ ( resampling in general) you’re applying a low-pass filter then decimating at the Nyquist frequency. Generally the monthly data is a downsampled version of the daily data and so forth.

    Smoothing is low-pass filtering data (by whatever algorithm) while retaining more data points than warranted by the Nyquist limit.

    It’s of course OK to do this if you know what you’re doing, Briggs opinions aside. Keenan apparently even admits to this now (at least for today).

    Regarding Muller’s comment though:

    . Just because smoothing can increase the probability of our fooling ourselves doesn’t mean that we did. There is real value to smoothing data, and yes, you have to beware of the traps, but if you are then there is a real advantage to doing that.”

    Other than for display purposes, what possible value is there to retaining the extra data points in smoothing?

    I’ll point out that the Nyquist-Shannon’s Sampling Theorem basically states that if you have a band-limited signal that is sampled at at-least twice the maximum frequency present in the signal, you can without error recreate a series with any arbitrary sampling rate higher than the Nyquist limit of the data series.

    So literally… there is no more information present in the oversampled version of the series than is already there in the Nyquist limited version.

    Again, where’s the “real value” in this?

  21. SteveF

    Well it’s unfortunate that the stuff which is most highlighted by the media isn’t as accurate as it could be. It’s like this huge temp reconstruction by BEST which is probably oppositely titled from reality goes out and all skeptics are idiots. Then the biggest critiques are off base.

    I have come to believe that with the step slicing, lack of identification of UHI and what I now beleive are useless CI’s that this is my least favorite temp product. They overprocessed the data horribly. Just keeping everything and kriging would have been preferable.

  22. Guess who wrote this:

    A statistical model should be plausible on both statistical and scientific grounds. Statistical grounds typically involve comparing the model with other plausible models or comparing the observed values with the corresponding values that are predicted from the model. Discussion of scientific grounds is largely omitted from texts in statistics (because the texts are instructing in statistics), but it is nonetheless crucial that a model be scientifically plausible. If statistical and scientific grounds for a model are not given in an analysis and are not clear from the context, then inferences drawn from the model should be regarded as unfounded.

    http://www.bishop-hill.net/blog/2011/10/21/keenans-response-to-the-best-paper.html

  23. JeffId–
    I think part of the media problem is that the journalists want answers fast. But no one can give an informed answer without first spending some time on the papers. Unless the papers were first presented at conferences, or the authors circulated pre-prints people who comment immediately generally can’t have dug into the details.

    This case is even more difficult: there are 4 papers. People commenting have to tease out what they think about:
    1) BEST best estimate of temperatures themselves.
    2) What they think of the methodology relative to other methodologies.
    3) What they think about BEST’s stated measurement uncertainties on temperature themselves.
    4) What they think of BESTs specific conclusions about UHI.
    5) What they think about the AMO stuff.
    6) What they think about BEST’s claims about statistical significance of various items. (This can be distinct from measurement uncertainty– but uncertainty affects it. Bounds on bias errors also can matter a lot here.)

    With 4 papers out there, journalists may not be able to craft questions that result in focused answers.

    But meanwhile, at blogs, forums, in news paper articles and in various “reviews” conversations about a whole bunch of things are getting intertwined. (Worse, the whole Curry vs. Muller narrative took over for a while. I doubt if they ever disagreed on much. Judy’s response to Rose’s questions should likely have been, “Hang on. Let me call Muller and find out what he really said or meant.”)

  24. I would say that the Lucia’s tone with Keenan is not different than I have seen her use with others. Lucia responds to almost all reasonable questions whereas I have seen other blog visitors and owners ignore what I thought were excellent questions.

    I am very much interested in hearing what Keenen has to add to the discussion here on modeling temperature series, and particularly in light of the quote from him immediately above and the direction of the discussion here.

    I forgot to note in the Wu detrending that I mentioned above that there is a residual cyclical content after detrending that would be required to be removed before comtemplating an ARMA model. I am wondering what that ARMA model would show with regards to auto correlation after trend (by Wu’s methods) and cycle removal.

  25. Kenneth–
    Isn’t the the case that ARIMA models can exhibit things that look a bit like cycles? I’m pretty sure there are examples in my stats book.

  26. “I think part of the media problem is that the journalists want answers fast.”

    I think you are being too generous here. I think what I see is that the journalists let their own biases, or perhaps better would be to say views, get in the way of their reporting. Throw in an inclination to make a controversial headline that might attract reader attention and I think you have what I see. An informed and attentive reader, in my estimation, can separate the factual evidence, or lack thereof, from these other distractions, but I would suppose a partisan reader looking for ready quips to quote would not. It is the dissemination of those quips that can turn the evidence/content from a paper/article on its head.

    Of course, the remaining element in these media articles is the primary source which usually is an author of a paper or an article. How much of the slant that evolves can be attributed to what the quoted author original wrote or subsequently quoted on? For example, how much is Keenan to blame for the misconceptions coming out of his WSJ article. And Muller for quotes on the BEST papers.

    I have not been reading the discussions between Muller and Judy Curry, but I see in this link below that perhaps Curry thinks Muller has gone a step or two too far in her interpretation of the BEST results. I would think in Keenan’s case one would want to know more about his intentions in throwing doubt on any AGW and Muller on her intentions of appearing to be a skeptic turned member of the consensus.

    http://www.express.co.uk/features/view/280948/Is-global-warming-over-

  27. “Lucia responds to almost all reasonable questions”

    I know. She does an excellent job interacting, whereas I often tire out and hope others will answer. Blogging can suck up huge time.

  28. Jeff–
    One of the problems at blogs is also that people will drop in and say “I’d really like to see you discuss ‘my pet topic B’ “. The topic isn’t utterly irrelevant– but discussing that would be a lot of work.
    I sometimes try to make it clear that certain questions directed at me aren’t “my” topic and tell people they are free to discuss it.

    There are entire threads where other people are talking to each other.

  29. “Isn’t the the case that ARIMA models can exhibit things that look a bit like cycles? I’m pretty sure there are examples in my stats book.”

    But not as regularly occurring as I saw in the Wu residuals.

  30. “I know. She does an excellent job interacting, whereas I often tire out and hope others will answer. Blogging can suck up huge time.”

    I was not referencing you in the other blog owners, Jeff. You have certainly countenanced this old guy on many occasions.

  31. Re: Carrick (Nov 1 22:11),

    Carrick, in the Muller quote on “smoothing” in #84989, it seemed to me that he was using that word with respect to BEST as you used “downsampling” rather than as, say, Briggs talks about the concept of smoothing. (But I could be wrong, and don’t have time to check my recollection.)

  32. AMac, then Muller needs to learn what “smoothed” means. 😉

    Seriously, normally when people use it, they’re just low-pass filtering (often just a moving average), while retaining the original sampling rate.

  33. Re: Carrick (Nov 1 22:11),

    I’ll point out that the Nyquist-Shannon’s Sampling Theorem basically states that if you have a band-limited signal that is sampled at at-least twice the maximum frequency present in the signal, you can without error recreate a series with any arbitrary sampling rate higher than the Nyquist limit of the data series.

    CD players can use oversampling when doing the digital to analog conversion. It helps to minimize phase shift, I think.

  34. DeWitt:

    CD players can use oversampling when doing the digital to analog conversion. It helps to minimize phase shift, I think.

    Perhaps related, but if you look at how DACs work, typically they have a voltage latch, so you’re keeping the voltage constant during each clock cycle.

    This introduces high-frequency noise that can potentially cause interference. By upsampling the signal before running it through the ADC, you’ve reduced the magnitude of these harmonics and also upshifted them even farther out of the audio band.

    (If you push the artifactual harmonics high enough in frequency, your audio cable will filter it for you via the skin effect.)

    Of course by increasing the sampling frequency, the width of the pulse is smaller, which means you also have a smaller delay. When you have an audio interface that records a signal (converts to digital), applies a digital effects filter, then plays it back out, you don’t want the delay to be overly long.

    This might be what you are referring to.

    That’s why high end audio interfaces will let you sample at ultra high frequencies such as 320ksps (google the MOTU 896 for an example). They also have a high speed FIREWIRE bus that allows the host computer to transfer data to and from the audio interface with minimal latency.

  35. DeWitt,

    The sampling theorem gives you theoretical conditions for perfect sampling and reconstruction. Oversampling A/D conversion allows practical implementation of the antialiasing filter which ensures that you stay within the Nyquist limit. Your analog filter only needs to satisfy Nyquist for the oversampled rate, and then you can apply digital filtering before downsampling (with appropriate noise shaping if desired) to the target rate.

    Oversampling at the other end, e.g., in CD player D/A converters, is usually there to permit a lower resolution D/A part to achieve a higher-resolution analog output.

  36. AMac–
    http://www.berkeleyearth.org/resources.php

    In the AMO paper, I think BEST used boxcar averaging.

    The land temperature data were smoothed with a 12-­‐month running average (boxcar smoothing); this removes high frequency (e.g. monthly) changes. The data prior to 1950 were noisier than the subsequent data, primarily because the number of stations was smaller, and for that reason we restricted the period for our analysis to 1950-­‐2010.

    The question is: Did they fool themselves?

    They also fit a 5th order polynomial to the yearly temperature data and subtracted that. The goal is “large signal that is not being studied in order to reduce bias in the remainder”. I would guess the goal is mostly to reduce what one might consider the “deterministic signal”. I sympathize with this goal– in fact, one of my criticisms of some papers is not only don’t they try to do this, but they don’t even seem to recognize that part of the deviation from a linear trend might be deterministic.

    Whether you can identify and remove the deterministic climate trend in this way , I can’t say. It’s difficult to know because we can’t necessarily claim there is a separation in scales between the slowly varying and rapidly varying components in climate. But presumably one can remove “large signal that is not being studied in order to reduce bias in the remainder.”

    After this, they compare the shape of the residual temperature to the AMO & ENSO visually and computing correlations.

    They find higher correlations with AMO than with ENSO.

    Now one of the criticisms with smoothing is you will often get a higher correlation than without smoothing. This is certainly true. So, we do expect the correlation between T&AMO will be higher after smoothing and the correlation between T&ENSO will be higher after smoothing. So, if they were “discovering” these correlations exist, we would need to think about whether we believe that claimed based on use of smoothed data. Arguments could ensue. One definitely needs to consider smoothing. But the mere fact that running a box car smooth introduced auto-correlation into AMO, T, and ENSO themselves doesn’t tell us whether this will result in BEST getting wrong answers for the questions they asked!.

    But here, BEST aren’t identifying the existence of correlation after smoothing.

    They are reporting R_T&AMO is larger than R_T&ENSO. So, recognizing that computed correlation are larger than would be the case if they didn’t smooth, the question is: Can we say that the correlation between temperature and AMO R_T&AMO really is higher than the correlation between temperature and enso (R_T&ENSO)? That is: if we tested the hypothesis the correlations were equal, could we say one is greater than the other?

    I don’t know the answers to these.It seems to me that asking whether the major claim can be supported based on the observed difference in computed correlations — both of which are inflated would be better than saying something like “never smooth”.

    Also, with respect to Carrick’s question: Why would one retain all 12 months worth of data? I’m not sure.

    But certainly, since the group was trying to detect a lag retaining them facilitates searching to see if the correlation between smoothed ENSO and Temperature peaked at 1 month, 2 months and so on. If, after smoothing, you retain monthly data, finding the lag associated with the peak correlation is computationally simple. You do it, you are done.

    You could probably still search by just annual averages of temperature starting in Jan-Dec and finding the correlation with the Jan-Dec AMO. Next create an annual average temperature series using Feb-Jan findi the correlation with the Jan-Dec AMO and so on. (Then, if you want to get fancy, repeat this using Feb-Jan AMO to see if it makes any difference.)

    It’s not clear to me whether the methods are liable to get different answers for the ideal lag. I imagine if you did this with synthetic data both methods would mostly work. If I had to guess, if you ran a whole bunch of synthetic data, keeping all the monthly data will have a very slight edge in getting better answers. But this is a guess. Still, if my guess is correct, smoothing and retaining all monthly values might make the test more sensitive– possibly letting one detect a feature that would be undetectable by dropping the intermediate values.

    My bigger concern with the smoothing is smoothing doesn’t just chop off all frequencies higher 1 year^-1 while leaving those below 1 year^-1 untouched. There is a roll off. And the 5th order polynomial might have affected the higher frequencies. So, one might want to know whether the results are very sensitive to the whole pre-whitening process. (I don’t know. )

    I haven’t read the spectral stuff… But the process must also affect spectra. (I’m going to the gym before looking at that more.)

  37. Lucia,
    “And the 5th order polynomial might have affected the higher frequencies. ”

    Yes, that is the first thing I thought. Subtracting a fifth order polynomial sure would seem to allow some of the honest-to-goodness ‘non-GHG driven’ variation to be artificially removed. So I would guess their “box-car smoothed, polynomial subtracted” residual data would have less variation than real.

  38. BEST:

    “The data prior to 1950 were noisier than the subsequent data, primarily because the number of stations was smaller, and for that reason we restricted the period for our analysis to 1950-­‐2010.”

    The increase in temp in 30s and 40s and then following slow down is therefore dropped from the analysis. This could be construed as convenient. Do they detail what noise they saw, other than the few stations/data pt comment? Even if there is more noise wouldn’t that arguably add more data points to whole analysis and improve the paper. Or they could compare a with and without report. There could be larger cycles involved that they miss or they could at least state that is a possibility for future study and improvement (& maybe they will add that in the final).

  39. I erroneously referred to Richard Muller above as her. That might have created confusion as I was talking about Muller and Judith Curry (who is a her).

  40. Kenneth Fritsch (Comment #85006) I think Lucia’s tone with Keenan, Monkton and Kimoto has been markedly different from her usual. There is a pattern.

  41. Anyway, “…BEST data: Trend looks statistically significant so far..”

    Trend, statistically significant….

    I expect you should have the numbers?

  42. Personally, I think climate data should only be smoothed enough so that it contains useful information. If it already useful, do not smooth it any further.

    Berkeley smoothed the Land temperatures with a 12 month moving average (5 months prior, 1 month current and 6 months post-current) because the Land data series is so variable.

    The Land data is more variable than I think most of us understood before (I’d run into it with US Land temperatures but I guess I didn’t see that the variability extended to the global Land series as well even though it contains up to 39,000 datapoints).

    The Troposphere and the Ocean temperatures, by contrast, have much less variability. This even extends into the diurnal cycle as well which is only about 0.2C for the Troposphere and the Ocean but is up to 20C for the Land.

    Generally, we shouldn’t smooth climate data beyond just getting rid of a noisy signal because valuable information will be lost.

    I note the ENSO works its magic with a 3 month lag. When you are smoothing Land temperatures over 12 months, that 3 month lag impact is lost. Berkeley has very little ENSO signal and they even wrote a paper saying the AMO has more impact than the ENSO. Well obviously, the long smoothing cycle (and an unexplainable 5th-order polynomial) destroys the ENSO information which is available.

    That is just an example of why we want to dampen the smoothing as much as possible.

    Further, Berkeley’s data only extends to September 2009 because the 12 month smooth information ends on this date (noting the last two months of April 2010 and May 2010 were based on Antarctica alone).

    So the 10-year temperature trend of the 12 month moving average can only be analyzed from September 1999 to Sept 2009.

  43. @Lucia

    Personally, I detect a ton of jealousy in Lucia’s tone. I think she’s upset because nobody offers her a guest editorial in a major news outlet.

    I find you to have…no arguments in some cases. You have no comment on the fact that there is no way to distinguish between natural and man made forcings. The blockbuster fact is, the earth naturally averages higher global temperatures and more carbon dioxide in the atmosphere.

    Furthermore, you inexplicably accept Michael Mann’s fake graph. Shame on you.

  44. Lucia,
    Just a thought on the physics vs stats conversation. I was musing on the question of what sort of model the physics might lead us to in terms of temperature time-series.
    Let’s start with the single box linear feedback model for the mixed layer, which as we know matches GCM model history data well:
    CdT/dt = F(t) – lamda*T
    Where F(t) is cumulative forcing
    C is heat capacity of the mixed layer (scaled by surface area)
    T is temperature change due to F(t)
    t is time
    lamda is the inverse climate sensitivity (watts/m2/deg K).
    The solution to the above equation for a single step forcing, Fo, is given by
    T = Fo*(1 – exp(-t/tau)) where tau = C/lamda.
    We can solve for T for a more general F(t) using superposition, since the equation is linear in T. I will discretize everything into one-year timesteps to keep the maths simple. Let tk = time after k incremental steps, i.e. tk = k years.
    Let f1, f2, f3 etc refer to annual forcing increments. The sum of these over k years gives the cumulative forcing F(tk).
    The superposition solution for temperature after k years is given by:
    T(tk) = sigma [fi *(1 – exp{-(k+1-i)/tau} )/lamda ]
    If we write the equivalent expression for the (k+1)th year and then take the difference, we obtain, with a bit of manipulation a recursive form of the general solution:-
    T(tk+1) – T(tk) = alpha* F(tk+1) – alpha*T(tk)
    Where alpha = 1 – exp(-1/tau) = constant only dependent on timestep length.
    The LHS of this solution is the first difference term in any annual temperature time series. The RHS shows the characteristics of a clearly autoregressive function. For a tau value of around 4, which gives a reasonable match to realworld data, as well as tested GCM results, the value of alpha works out to be 0.22. (NB Recall that this value applies only to annualised data.) This model form reaffirms the requirement for co-integration of temperature and forcing for any “trend” in temperature series to be meaningful.
    I would suggest that this form supports Lucia in that it provides little support for an ARIMA (p,1,q) structure in any statistical model, however short the duration of the data. On the other hand, it reaffirms Keenan’s point that linear OLS plus AR(1) noise should not be readily assumed. In fact, the linear temperature series model only makes sense under an assumption of linearly increasing radiative plus non-radiative forcing over a long time-frame (>> tau). This seems like a stretch to me, given the annual to decadal (ENSO) and the multi-decadal cycles (Solar, AMO, PDO) evident . A major factor in the noise term in the temperature series is the unaccounted-for variation in F(t) from whatever source. This may be due to non-radiative flux entering the mixed layer (aka natural variation?) or error in radiative forcing estimation.

    I am not suggesting that this offers a resolution to the issue, but hopefully it might help clarify the issue for some.

  45. Lucia,

    ” That is: The IPCC openly acknowledges that the AR1+ linear trend is not perfect, can lead to over-estimation of statistical significance. ”

    Here’s the puzzling part (Table 3.2):

    The Durbin Watson D-statistic (not shown) for the residuals, after allowing for first-order serial correlation, never indicates significant positive serial correlation.

    I don’t know how this was calculated, but it seems to be a claim that “AR1+ linear trend” is valid. This is at odds with the Beran’s example discussed in another thread. Thus, it would interesting to verify this result.

  46. UC–
    They may claim it’s valid by some criteria without also claiming it’s perfect or possibly even best.

    I’ll admit that they don’t discuss this at length– but by the same token, use of AR(1)+trend is such a tiny part of their reason for justifying believing warming is real that I suspect the authors just judge a long discussion unimportant. Going through the bulk of their argument, it’s not really based on “linear trend” is statistically significant. I honestly don’t think anyone thinks the ‘climate response’ (i.e. deterministic response to forcings over time) is linear in time over the past century. The forcings certainly haven’t been.

  47. Carrick–
    Jay is ‘slowed down’. I think he’s supposed to be moderated…. better check on that.

    The one who is really left off on a raft ‘off the island’ is the person various IP from ‘.btcentralplus.com who I suspect is a sockpuppet using ‘ constantly changing his emails! (I suspect it’s the same sock-puppet who first logged on using a proxy-server IP with a fake email.) The filter is catching him/her/it. The majority of comments are rhetorical questions– though the most recent isn’t.

    Anon. is ok. But sock-puppetry to argue by rhetorical question? That’s way out of line.

  48. Hi Carrick,
    This is not really “my” model. It was the model of choice selected by Breusch and Vahid. See here:
    http://www.buseco.monash.edu.au/ebs/pubs/wpapers/2011/wp4-11.pdf
    (The paper is very readable and worth taking a bit of time over, even if, for philosophical or other reasons, you don’t like the idea of having a unit root in the statistical model.)
    Quite a while ago, I used the same structural form to fit to GISSTEMP 1880-2009 data, because I wanted to make some specific tests. I am not wedded to the model.
    I will dig out my original code and run it out in prediction for a while. Note that this model does have a positive drift term, which sneaks in with a weak (one-sided) test of significance, discussed in the B&V paper above.

  49. Thanks Paul_K The reason I ask is you can fit any data to almost any model, given enough parameters (and good enough fitting software).

    The question is what happens in the “verification period” (ranges of dates for which you have data, but aren’t being constrained by the model). Polynomial based models diverge as x^n (n=highest order in polynomial) when you evaluate them outside of the fitted range for example.

    I’m just curious what happens with the ARIMA based approaches when you do that.

    [As I understand it, if you add a seasonal component to your model, ARIMA(p,d,q)x(P,D,Q), that of course forces a periodic structure onto your noise. ]

  50. Let me say this again, and please maybe get it right:

    The question is what happens in the “verification period” (ranges of dates for which you have data, but [the model isn’t] being constrained by the [data]).

  51. Hi Carrick,
    Here are a couple of realizations of the ARIMA (0,1,2) in predictive mode. I generated 10 or so (manually), and can’t see a striking difference between them. They all show well bounded stochastic excursion around a drift of 0.006 deg K/year. However, it is interesting that they also all show a flattish period upto about 2030, which is a hangover from the recent observational data. I am on the wrong platform to run out Monte Carlo on them which would be the necessary step to yield the assembly estimate. This would give me a mean prediction as well as being a necessary step to answer your verification question. In any event, I hope that these give you some idea(s).
    http://img412.imageshack.us/img412/5039/arimarealisations.jpg

  52. Hi Lucia, you probably missed my earlier question, “Anyway, “…BEST data: Trend looks statistically significant so far..”
    Trend, statistically significant….
    I expect you should have the numbers?

    Can you enlighten us?

  53. MarkR-
    I didn’t miss your question. I don’t even know what idea you are trying to express or what your questions are supposed to be asking. Lately, you have been posting numerous nearly incomprehensible comments. To avoid wasting my time, I am ignoring them.

  54. My apologies for not being clear Lucia, for the sake of clarity, is the BEST data Trend statistically significant so far? If not you may want to correct your preceding post. Do beg pardon for taking up your important time.

  55. MarkR– I take it then that in your dialect

    I expect you should have the numbers?

    translates to

    is the BEST data Trend statistically significant so far?

    The answer to the later question is: Yes. It is.

    Thanks for taking the time to express your question in a form where I could understand it.

  56. Carrick,
    Interesting – but pure coincidence of course. It just happens that this particular statistical model fit seems to take somewhere around 30 years to “resolve” a departure from the apparent trend line – which is actually caused by a drift of 0.006 degrees per year. It is actually the decaying effect of historic noise terms in the model. I guess that that is why it does a good job of capturing the 60-year cycles in the historic data.

  57. Lucia,
    In my comment #85055, I have just realised that I made an error in transcribing the key equation.

    It should have been:

    T(tk+1) – T(tk) = alpha* F(tk+1)/lamda – alpha*T(tk)

    This is a perfectly valid solution routine to emulate all of the GCM results tested to date (over the instrument period). The more I think about this, the more I am convinced that it provides a credible method of formulating statistical models of, and hypotheses about, the temperature series which bear a direct relationship to the structure of the total (radiative and non-radiative forcing) assumptions. In other words, it allows a direct bridge between the statistics and the physics.

Comments are closed.