In a previous post, I reported trend analysis of the difference between GISSTemp observations and the model mean projection for surface temperatures from simulation from IPCC models extended into the 21st century using the A1B scenario tell us we should reject the hypothesis that the model mean simulation agrees with the observations reported by GISS. Specifically, trend analysis beginning with start dates prior to 1980 all indicate we should reject the multi-model mean trend as reproducing the observed trend.
In that post, I also said I would apply the same analysis to Hadley when the December 2009 data arrived. Below, I have a figure showing mean trends for the difference between observations and the multi-model mean and uncertainty intervals based on HadCrut NH&SH through Dec. 2009.

Because the trend is fit to the difference between observations and the multi-model mean simulation, we expect a trend of 0C/century if the multi-model mean simulation faithfully reproduced the mean trend. (The IPCC projections in the AR4 were based on this assumption.)
Examining the figure above, we see the notion the multi-model mean of simulations for the earth’s surface trend faithfully reproduce the HadCRUT observations should be rejected at a confidence of 95% if we elect to do the analysis with start years beginning in 1950, 1960, 1970, 1980, and 2000 or 2001. The uncertainty intervals are based on the assumption the residuals from a straight line can be described by an ARMA(1,1) noise model, and the influence of El Nino has been accounted for using the MEI index.
Some readers may be curious as to whether this result differs from the similar analysis using data through Nov. 2009. Of course, they don’t differ much, but interestingly enough, the trend computed from1980 just kicked over into “rejection” territory; it was in “fail to reject” before the December anomaly was reported. (For earlier results see previous post. Of course, this is a mostly rhetorical issue, since in both cases, the upper range of the uncertainty intervals just grazes the key “0 C/decade” value.
What if the IPCC had used different models?
The analysis above is just fine if we consider the IPCC projection to be based on the specific ensemble of models they actually used. However, when I presented a similar analysis for GISS, Nick Stokes objected that I did not consider how the results might have varied if the IPCC had used a different collection of models.
Fortunately, it is very easy to extend the method to account for this. To account for this, I modified the standard error in the trend by taking the square-root of the sum of the square in the standard error in trends across all 22 models and the standard error in the difference in the mean. The latter is used to form the uncertainty intervals in the figure shown above.
The standard error in trends across all 22 models was computed by finding the variance in the trends across all 22 models, subtracting an estimate of contribution of “weather” from this, and then dividing by 21, which is the number of degrees of freedom for the models and then taking the square root of the resulting value. (The contribution of “weather” was subtracted to avoid double counting as it is already contained in the difference between observations and models. I estimated by assuming the “weather noise” for models and observations are identical, obtained this value from the uncertainty in the trend based on the linear fit to the difference between the model mean and the observation.)
Of course, these algebraic manipulations did widen the uncertainty intervals somewhat, particularly for start dates beginning at earlier times. The new results are shown below:
As you can see, accounting for the possibility that IPCC might have used other some other set of models selected randomly from the universe of all possible models, the notion that the model mean of simulations from “IPCC-like” models faithfully reproduces HadCRUT should be rejected at a confidence of 95% if we being trend analysis in 1950, 1960, 1970, 2000 and 2001. However, we fail to reject if we begin trend analysis in 1980 or 1990.
That is to say: If we are counting rejections based on analyses beginning in January of each decade, we dial back one ‘reject’ to a ‘fail to reject’. This is the rejection from 1980. However, for the most part, accounting for the variability due to the choice of individual models or runs makes only a small difference in the general assessment of whether the multi-model mean surface temperature from of some hypothetical set of models selected from a universe of all possible models would faithfully reproduce the earth’s surface temperature.
Why are we getting so many rejections?
Readers always ask me what these rejections mean. The answer is: based on the choice of statistical model for the ‘weather noise’ (i.e. ARMA), the simulations probably do not faithfully predict observations of the earth’s surface temperature. This statistical analysis cannot tell us anymore than that.
Reasons why the difference exists could range from a) the weather noise might not be ARMA, b) the forcings under the SRES were higher than on earth, c) some physics in the models are inadequate, d) the observations do not reflected the change in surface temperature of the earth’s surface, e) other I haven’t even thought of, f) the MEI index is nothing more than a fiddle factor that overfit the data and so causes spurious shrinkage of the uncertainty intervals. The current analysis can’t distinguish this.
That said, the current analysis tells us: If the assumptions are correct, the trend in the difference between the multi-model mean and the earth’s surface temperature is statistically significant. That is, the multi-model mean is overpredicting warming and that difference appears statistically significant.
Posts discussing some analytical choices
- Why taking the difference deals with volcanoes and other non-linearities better than other methods.
- How I estimate the uncertainty intervals for ARMA(1,1) noise.

Don’t the modelers do these sorts of anlayses in the PRL with discussions? Or do they still claim that the temporal span is still too short, i.e. weather, not climate?
There is a typo: the last date on the abscissa should be 2010. Your posts are a daily must. Keep up the good work and civil discourse.
MCS
The models either predict warming or they don’t.
There is no error margin that should include Zero or even a modest warming. They are not accurate enough to be relied on if the trends to date do not predict a dangerous level of warming to come. A simple mathematical +/- does not test the theory at the level it should be tested on.
Redefine the error margins so that they result in at least 2.0C of warming by 2100. That would be the logical error margin as opposed to the propaganda margin.
Morley–
The dates on the abscissa are the start date. All trends are computed from that start date through Dec. 2009.
And what about interval years? 1981, 1982, 1983…. the climate is so full of chaotic behavior that one could excuse this graph as possibly cherry picking the somewhat bad years for the models, gaming with the coincidence of them ending all in “0”. But the difference between 2000 and 2001 is enough to see that one year makes a lot of difference in there.
Is that graph too hard to do? I know you’ve done something similar, but not with the IPCC models in the baseline = 0….
Great work anyways, as always.
Luis–
These tests don’t flip around that much. The graph isn’t hard to do, but I had programmed this in EXCEL so I could see all the various corrections etc. as I do them. I plan to rewrite using a different tool to make it easier to spit out all the years.
EXCEL is nice for some things, but clumsy for others.
I think you are confirming what people like Roy Spencer are saying – the forcings assumed for the models are wrong, for reasons not yet fully understood.
When I do modelling (in chem eng, not climate) I have to fit the model to the data. If it doesn’t fit, I have to pick apart the assumptions and the code until I understand the reason for the variance. Then, once the model is seated on a solid basis, I have a chance to recommend actions defensibly. Money rides on this (too).
I had such a case not long ago, fitting a model I built to a clunkier preexisting one using older package, and also to the real data. I found a couple of errors in the previous model, in one case a wrong thermodynamic value was used. When I corrected this one process tank temperature went off scale. So I cautiously ask a plant guy ‘does that tank ever boil?’, sure enough he said: yes it does. Yay! Model fits reality.
GIGO is alive and well, it likes living in computer models. I do hope the climate modellers can kick their thermodata and temperature records into shape, the real data is increasingly making them look like idiots, er, hide the decline?!
Bruce of Newcastle (Comment#31680),
For sure GIGO is alive and well in lots of places, including climate models. But in the case of climate models, there are two areas that stand out as likely homes for the garbage data: aerosol forcings and ocean heat uptake. People worry a lot about the accuracy of historical and recent temperature trends, but I think there is perhaps too much focus on these data. While there may be some overstatement of increasing temperature trends, the total error is not likely to be very large. After all, the satellite data do confirm considerable warming over the last 30+ years, and there is only a small discrepancy with surface based measurements. Substantial discrepancies between the predicted temperature profile of the troposphere and the measured profile suggest the models have moist convective heat transport wrong, but these discrepancies have been waved away (eg. Santer el al 2008) as less than 95% certain.
.
Large aerosol cooling effects and long ocean lag values, including enormous heat transport to the deep ocean, are used in the models (and indeed required by the models) to make high climate sensitivity fit the temperature history. These are the most uncertain of the data that go into climate models. Until improved measurements of these parameters force modelers to adopt much shorter ocean lags (ARGO) and much lower aerosol forcings (NASA’s Glory mission), the models will continue to indicate high sensitivity and extreme future warming.
.
Lucia’s continuing efforts to show the IPCC models deviate significantly from the temperature data is helpful of course, if only because it makes fools like Tamino go apoplectic (which is at least entertaining!). But so long as aerosols and ocean lags remain poorly defined, modelers can (and will) tweak the assumed values to explain the recent temperature discrepancies, or just dismiss discrepancies as the result of extreme weather noise. I think you can count on many of the next round of IPCC model releases to accurately hind-cast the lack of post-2000 warming, but continue to predict extreme future warming.
Lucia,
“The standard error in trends across all 22 models was computed by finding the standard deviation in the trends across all 22 models, subtracting an estimate of contribution of “weather†to this, and then dividing by the square root of 21, “
Shouldn’t this be done with variances? ie you should get the variance in trends, subtract the weather variance, divide by 21 and then take the square root? Or is that what you did?
Re: Nick Stokes (Feb 2 08:33),
Actually, I did do it with variances. I’ll fix the text.
Lucia,
“the MEI index is nothing more than a fiddle factor that overfit the data and so causes spurious shrinkage of the uncertainty intervals”
I think it would be better if the MEI were de-trended before it is used to account for short-term temperature variation, since the rather obvious trend since 1950 in the index may well be a result of warming rather than a cause of warming. But outside of the influence of a trend in the MEI, why would the MEI be considered a fiddle factor? It sure does correlate strongly with short term variation.
SteveF–
The correlation between MEI and time in the sample is included in the computation of the uncertainty intervals for a multi-linear regression. This is true provided that you don’t use it in this order:
1) Compute the best fit to predict temperature based on MEI. (Call mmei)
2) Treat mmei as a known constant and correct the temperature data.
3) Then find best fit trend, m, between “mei corrected” temperature and time, Tempcorr_mei= m*time + b.
There can be reasons why you might what to do it this way, but if you do, you would need to formally compute the uncertainty in mmei, and then figure out how to include this uncertainty in the estimated uncertainty for m.
In contrast, if you just do a multi-linear regression between T = f(MEI, time), then the formalism for computing the uncertainty in the parameters involves computing the correlation between MEI and time. If that correlation is strong, the uncertainty intervals for paraterns are widened relative to data sets where the correlation between MEI and time is small.
So.. the correlation that bothers is taken care of– it’s in the uncertainty intervals.
FWIW: I actually did this both ways to see if it makes much of a difference. However, when I did it the first way, I didn’t figure out how to expand the uncertainty intervals properly. If I showed those results, the uncertainty intervals would look smaller than the ones I show–but I know they are too small.
Lucia,
Thanks for the explanation. But does detrending the MEI essentially do the same thing, that is, yield a ‘clean’ (time independent) regression T = f(detrended MEI)?
SteveF–
I don’t see how detrending MEI helps. Supposed that MEI does have a positive correlation with time. Also, supposed the cause or the temperayure rise during your sample really is due the fact that it ends with a high MEI and begins with a low MEI. If you detrend MEI, you are basically assuming the rise in MEI over your sample does not explain any of the trend over time. The regression will now attribte that increase to “time”– the other exogenous variable. This will result in a biased trend with time.
Mind you, what we expect is MEI should be trendless in the long run, and we expect no correlation between MEI and time if our sample is long enough. But I don’t see how mathematically removing the trend prior to fitting removes any uncertainty. I think all that happens is that if you do that to MEI at the outset and then use the corrected MEI in a multiple-regression, the regression will report uncertainty intervals that are too small. You will trick yourself into thinking you have shown something significant when all you did was set a term that contributes to the uncertainty intervals to zero– when, in reality, that term should be finite.
Lucia,
There is I think one other interesting thing in the first graphic: the best estimate for the trend in the difference (independent of uncertainty intervals for each) with starting date is pretty strong. This seems to me consistent with model errors in calculated sensitivity, since the GHG forcing (which is pretty accurately known) has increased most rapidly in the post 1980 period. Too high a climate sensitivity in the models ought to produce this kind of increasing trend in the difference between models and data as forcing rises.
lucia (Comment#31690),
OK, now I am confused. Maybe I do not understand what you are doing after all.
.
The temperature trend could be regressed against time and other variables (forcing, lagged forcing, MEI) I suppose. But I do not see why the temperature should be an explicit function of time in any case… we don’t expect the temperature to rise just because time passes. But we certainly might expect a variable like MEI to have along term trend that is caused by (rather than causes) a long term trend in temperature. If the MEI is detrended before regression against measured temperature, T= f(detrended MEI), then I would expect the regression constant would be the same as would be calculated regressing measured temperature against both time and MEI together, T = f(MEI, time), and that the regression constant that pops out for ‘time’ ought to have the same absolute value as the de-trending constant, but with the opposite sign. Am I missing something here?
If you detrend MEI then you are going to increase the values in the early part of the MEI data and reduce them in the later part, which will have the effect of reducing the contribution of MEI to the overall temperature trend, so yes, the net will be that you assign less of the measured temperature rise to MEI and more to other factors; the correlation with MEI gets weaker. I do not see how this would artificially narrow the confidence limits. If you do not detrend the MEI, then it seems to me you are implicitly saying that you believe the MEI really is not being driven by a long term temperature trend, and that the trend is in fact being partly driven by the long term MEI trend.
Lucia,
How do you put italic print in your comments?
Re: SteveF (Feb 2 11:10),
Be careful there. The trends are not linearly independent. After all, the trend since 1950 uses the same data after 1980 as we use to compute the trend after 1980. Using different start dates mostly shows that a) the result isn’t particularly sensitive to the analysts choice of start year. b) The uncertainty intervals are much larger for short periods of time.
I regress the difference in observations and the multi-model mean using (Time, MEI) as exogenous variables.
It’s parametric. There is a trend in time but the cause is not, strictly speaking time.
But if the multi-model mean correctly predicts the evolution, there should be no trend in time in the difference between simulations and observations. If they didn’t use the anomaly method, we could just compare instantaneous temperatures. But we can’t do that because those agree during the baseline by defintion (or, more specifically, the mathemagic of subtraction.) So, instead, I test to see if there is a trend in the difference over.
I think there are missing words in this. I’m adding in green.
I hope that when you do T=F(MEI, Time), the parameter describing Temp ~ m* time doesn’t become the parameter that used to describe Temp ~ mmei* MEI. So that can’t be what you mean. Could you clarify?
But, whatever you expect, you could always just either look at the actual formulas or just check a few cases in a statistics package.
Either way, if you detrend, your estimate of the uncertainty intervals for the parameters will not include the magnitude of the correlation in the sample. Detrending does not magically make the fact of correlation in your data disappear. It can set the apparent value to zero when you stuff the detrended data into a regression later on. But if you detrended data before using it in a multiple regression, then you formally include the uncertainty from the detrending operation for MEI in the full uncertainty calculation. This would be an algebraic operation to add the uncertainty back into whatever your stats package spit back out at you.
If you don’t include it,you are fooling yourself.
Lucia,
“Do you mean the fitting parameter fir MEI you got from fitting MEI = f(time) only?) but with the opposite sign. Am I missing something here?”
Yup, that is what I meant.
“This would be an algebraic operation to add the uncertainty back into whatever your stats package spit back out at you.”
OK, now I understand. You need a way to carry the uncertainty in the slope of the de-trend operation into the final regression.
SteveF–
Yes. If you don’t detrend MEI, and MEI really does affect temperatures, the multiple regression gives you unbiased estimates for the parameters in the fit, and provides the correct uncertainty in the regression coefficients. If you do detrend, I don’t know if estimates of the parameters remain unbiased but I do know the uncertainty estimates are computed as if the original sample did not exhibit a correlation between MEI and Time. So, the computed uncertainty intervals are smaller, and this is incorrect.
Lucia,
I expect the regression coefficient is reduced slightly, since de-trending the MEI would remove some of the MEI driven trend, independent of reducing the uncertainty estimate.
Lucia,
“The uncertainty intervals are much larger for short periods of time.”
Sure, but the best estimate of difference between the models and the data increases even faster than the uncertainty increases; the difference both grows and becomes more statistically significant during the period when the GHG forcing is the greatest. Not a good sign for the models.
Re: SteveF (Feb 2 13:24),
Yes. That’s why the smaller difference in trend computed from 1950 is statistically significant. But what I mean is that you need to be careful when trying to interpret the apperant trend in “trend” as a function of start date. If I repeated starting in 2004 or something, the trend you see will suddenly change.
lucia (Comment#31707),
Didn’t some of your earlier posts along the same line include every year as a starting date? (Or is my memory failing me?) Nobody can claim cherry-picking starting years if the calculation and corresponding uncertainty is done for every year.
SteveF–
Yes. I did the same thing with all years– but back then I used red-noise. I also didn’t do trend analysis based on the difference. Don’t know why I didn’t think of the simplest way to deal with the volcanoes early on. It had been bugging me. Then… I thought of it.
So, I need to reorganize the script that does all that. I didn’t initially because I needed to figure out which way of actually implenting the ARMA(1,1) uncertainty intervals gave about 5% false rejections when I claim that’s what I’m getting. (I tried fitting a logarithm, taking a ration and doing a bunch of other things. The method I’m using gives about the correct rate, and is easy to implement.)
Dear Lucia,
Thank you for a splendid post. This is very clear and very useful. All best wishes,
Dale McIntyre
More Lukewarmer Science
“Carbon-cycle feedback smaller, but still positive
The team found that this feedback coefficient is about five times smaller than previously expected – which suggests that the amplification of manmade global warming by carbon-cycle feedback will be less than previously thought.”
http://environmentalresearchweb.org/cws/article/research/41536
SteveF (Comment#31682)
“.. I think you can count on many of the next round of IPCC model releases to accurately hind-cast the lack of post-2000 warming, but continue to predict extreme future warming.”
(Comment#31706)
“…the difference both grows and becomes more statistically significant during the period when the GHG forcing is the greatest. Not a good sign for the models.”
SteveF,
Thanks for the good commentary and thanks to you and Lucia for the analysis. Was stretching my rusty stats. I can’t talk variables unfortunately as this is not my field. (Incidentally talking of MEI the SOI is usually on our nightly news here :). Despite el Nino the IOD has rustled up a cyclone for us, which sometimes happens, so our farmers in Queesland are all celebrating and its cool and rainy here. One more residual in MLR T = f(MEI, time))
Regrettably I think you are talking in the right direction, which is one reason for my concern. If you write a model *to* fit a dataset *then* predict what you want it to predict, you don’t have a model you have a fantasy. If I did that I could be put in jail in my business.
Keep up the auditing! Last time I was statistically audited I passed, it was a nice feeling and we got a paper out of it.
Lucia,
Do you have a list of the 22 models that you studied?
Nick-
The data are available here:
http://climexp.knmi.nl/selectfield_co2.cgi?someone@somewhere
I used the ones that indicate “sresa1b”.
Lucia,
A lot of those have several runs. Did you use just the first one of each?
Nick-
No. I downloaded all of them. To create the multi-model mean I
1) Averaged over fund for an individual model to get t he model mean then
2) Averaged over the 22 models.
This is the same method Santer et. Al used in their paper criticizing Douglas. The formulas are provided in equations (7)& (8) I think it’s also described in the AR4. Do you need the specific citation to Santer’s paper on tropospheric trends? .
Lucia,
I did download the model data, and I’ve tried my own modelling on my blog here. I found a couple of the model runs gave error messages (advise Geert, it said, but I settled for just 20 models that downloaded easily). I used annual means, and didn’t make an MEI correction. And I didn’t regress the model mean differences; I computed the means separately and subtracted, but added the variances.
The results were rather similar to yours, although I think that for me adding the variance due to model selection had a bigger effect. I’ve uploaded the R code.
What do you mean by
“I did not correct for AR4 correlation”? Do you mean serial autocorrelation for lagged residuals? It is lower for annual averages. But you do know that testing annual averages instead of monthly averages has lower power– right?
It also looks like you are examining tests of individual models rather than tests of the multi-model mean.
Nick
This is a conceptual error on your part. The “noise” in the multi-model mean is correlated between the two time series. Pinatubo causes a depression in temperature in both. When the noise is correlated, you need to use the difference in the two time series. You can easily do the proof in the limit of infinite series and see that cross correlation screws things up if you don’t take the difference. Unrecognized cross correlation between time series can make the uncertainty intervals too large or too small depending on whether the cross correlation is positive or negative. In our case, it makes the uncertainty intervals too large.
You should look this up in a non-tendentious general stats book– the problem of cross-correlation is well understood. If you want to stubbornly pretend it is not and present inflated uncertainty intervals, welll… then…well and good. You also seem to have some fuzzy notions about waht Santer’s test (eq. 12) of multi-model means accounts for. It accounts for the scatter in trends from the multi-model means and you seem to suggest it does not. I have no idea why you are trying to then add even more noise into the test– the test in equation (12) of Santer is just fine– provided the noise in the two series are linearly independent, and you really expect residuals from a linear trend from repeat samples of a realization to be statistically independent form each other.
Lucia,
It’s early morning here (and it was late when I was absent-mindedly writing AR4 for AR(1)). But I think you’re right that I have mis-described Santer’s test. Like Douglass, he accounts for what I have called the third kind of variance, between model trends, and has omitted the uncertainty of those trends (second kind) resulting from regression (the whiskers in my Fig 1).
On subtracting the trends, I don’t agree that not doing it is a conceptual error. If the series are independent, the way I did it should give the same result. I did it that way to make the sources of variance clearer. However, I agree that subtraction could be useful in subtracting out common spurious influences, so I’ll try a modification that does that. Unfortunately, I probably won’t get that done today.
Going back to Santer, he seems to acknowledge that there would be an effect of the regression uncertainty. He says:
“The first assumption (which was also made by DCPS07) is that the uncertainty in <<bm>> is entirely due to inter-model differences in forcing and response, and not due to differences in variability and ensemble size.”
Nick
Santer’s equation (12) is correct for testing whether model means trends drawn from a universe of models whose mean trend matches the earth’s provided the standard errors are estimated correctly. There is absolutely no controversy over the standard error for the model means including all correct contributions, including residual weathernoise from the models. If you don’t think so, ask Santer. If each model had a bajillion runs, the numerical value in the application Santer did would be smaller than the one he computed across his sample. You absolutely don’t get to add in more “weather noise” for a test of the model mean.
There may be some other hypothesis you are trying to test, but in which case, you should figure out how to describe that hypothesis in words, and afterwards we can discuss whether or not your algebra achieves the goal of your hypothesis test.
This question has been thrashed to death– and is closely related to the one I sent Gavin and which appears in the climategate letter.
On the one hand, you’ve agreed to subtract… so that’s good. But I thinks it’s important that you understand that doing it the other way is a conceptual error.
What matters is whether that which you consider “noise” is independent across repeat realizations (i.e. runs.)
If you don’t grasp the conceptual error, we are probably going to have to take baby steps for you to get this. So, instead of starting by arguing about the subtraction, let’s start by looking at the problem in computing your “whiskers”.
In your current post, you treat all residuals from a linear fit as “noise”. The term is vague, but we’ll use that. Now consider two runs of a model– say Model E. In both runs, the temperature drops after “dip” after Pinatubo erupts. So, let’s consider two runs. The residuals for month ‘i’ from run “a” are, ra,i; those from run “b” are rb,i.
If you wished to, you could compute the correlation Ra,b < ra,i rb,i >/ < ra,i 2ra,i2 > If the effect of pinatubo (or any non-linearity in the underlying trend) is sufficiently large you will find this correlation is discernable.
This means the “noise” is correlated across separate model runs.
When this is the case, the method of estimating the uncertainty for an individual run used by Santer (and you to create your whiskers) is inappropriate. Assuming your choice of statistical model is otherwise ok, then if R is positive, your whiskers will be too large; if it is negative; your whiskers will be too small.
If you do not believe this, I invite you to do the following:
Generate three time series of random white noise. wo,i, w1,i,w2,i. Make all three have the same standard devaiton and a mean of zero. To focus on the specific argument of interest, make all of these white.
Now create two time series. Create two time series 1 & 2using
T1,i =m ti+ wo,i+w1,i and
T2,i =m ti+ wo,i+w2,i
where “t” is time, and m is the shared trend. Note noise from the two is autocorrelated.
Compute trends and do a t-test your way where you find the noise for T1 and T2 individually, ignore the shared autocorrelatio and diagnose rej/fail to reject with α=5%. Do this a bajillion times and see if you reject 5% of the time– as you should. (You won’t; you’ll reject less than you ought to. This mean your test isn’t α=5% it’s α= some smaller value and you ought to admit it.)
Now, create a series that is dT= T1-T1. To a t-test to see if it’s mean trend for dT=0 rejects or fails at the proper rate. (It will.)
The reason the non-subtraction method will not give you the proper false positive rate is you are using to estimate the variance in the difference in trends is based on the assumption that what you call “noise” in run 1 is not correlated with “noise” in run 2.
In the simple example I invite you to perform, the shared noise will look white. But that’s not important to the general effect. What matters is whatever you treat as “noise” is correlated across repeat samples. If you want to see test for red noise, have at it. If you want to add a shared non-linearity in the underlying trend for the two cases, add that. Etc.
Interestingly enough, if you are trying to find out of the trend in T1 and T2 are the same, subtracting and analysing the difference will rarely do much harm, and will often prevent you from fooling yourself.
Now, back to models and observations: the Pinatubo noise is correlated across repeat runs of models like GISS-ER. It would would be correlated across repeat runs we could imagine re-running the 20th century but having been able to scramble the IC’s for earth weather slightly. (The ability to perform these repeat runs underlies the computation of the uncertainty intervals for the observations, so the fact that it can’t actually be done is irrelevant. Your analysis already assumes this hypothetical.)
Subtracting the multi-model mean and the observations takes care of this issue when comparing the observations and the models because the model-mean claims to capture the part that would be explained by Pinatubo, and which subsequently would be correlated with what Pinatubo caused to happen to the earth’s realization. (Doing this also means you can’t make bar and whiskers super-imposed on a band figures quite like the ones in Santer anymore because you’ll only have one series, not two.)
First: Ideally variability due to inter-model differences in forcing is the uncertainty you want in the t-test.
When testing the multi-model mean, you don’t want the rest of the stuff.
Santer is saying he is assuming he’s got that number, but actually he knows he doesn’t. It happens that Santer’s estimate based on his sample is larger than the uncertainty due to inter-model differences in forcing and response because none of the models have enough runs to remove the “weather noise” from his estimate of the inter-model differences.
But that’s ok with me because there is no way to take out the weather component out given the limited number of model runs that exist. The only way to take it out would be for each modeling group to have provided many many model runs– which they did not.