Showing posts with label uncertainty. Show all posts
Showing posts with label uncertainty. Show all posts

Sunday, 23 July 2017

Is nitpicking a climate doomsday warning allowed?


Journalist and amateur mass-psychic David Wallace-Wells published an article in the New York Magazine titled: "The Uninhabitable Earth - Famine, economic collapse, a sun that cooks us: What climate change could wreak — sooner than you think."

Michael Mann responded on Facebook as one of the first scientists. He disliked the "doomist framing" and noted several obvious inaccuracies at the top of the article that all exaggerated the problem.

Some days later seventeen climate scientists of Climate Feedback (including me) reviewed the NY Mag article.

In my previous blog post I argued that this is a difficult, but important topic to talk about. The dangers of unfettered climate change are huge. Also if we do not act faster than we did in the past we are taking serious risks with our civilisation and existence. We should seriously consider that the situation can become worse than what we expect on average. Such cases are a big part of the total risk.

While the danger is there and should be discussed, the article contained many inaccuracies, which typically exaggerated the problem. Thus I have rated the article as having "low scientific credibility", which was the most selected rating of the other scientists as well.

Both the critique of Mann and of Climate Feedback produced quite some controversy. This was exceptional; normally the people who see climate change as an important problem trust the scientists who told them about the problem.

Like other climate scientists I correct both sides when I see something and have enough expertise. Although it is much rarer to have to correct people who see climate change as a problem (let's call them the "concerned"). I guess the real problem is big enough, there is not much need to exaggerate it. The overwhelming amount of nonsense comes from people playing down the problem. Normally the deniers really dislike contrary evidence, but the concerned are mostly happy to be corrected and to be able see the problem more clearly.

It is interesting that this time the corrections were much more controversial. Why was it different this time? Was our "nitpicking" a case of the science police striking again? Could Climate Feedback give more useful ratings?

Doomsday scenarios are as harmful as climate change denial

Michael Mann followed up his Facebook post with an article in the Washington Post together with communication expert Susan Joy Hassol that "doomsday scenarios are as harmful as climate change denial". The key argument was:
Some seem to think that people need to be shocked and frightened to get them to engage with climate change. But research shows that the most motivating emotions are worry, interest and hope. Importantly, fear does not motivate, and appealing to it is often counter-productive as it tends to distance people from the problem, leading them to disengage, doubt and even dismiss it.
This is an argument climate activists who want to be effective need to be aware of. The well-known groups already tend to stick to the science. Climate change is bad enough as it is and they rightly value their credibility.

The argument makes me uncomfortable, however, when connected to science. The situation is what it is. Also if that provokes fear, scientists should stick to the evidence and not tone it down to be "effective". It is the job of an activist to be effective. It is our job is to be honest.

If the population would have to fear we are not honest, that would produce additional uncertainty and give more room for the prophets of doom. Thus toning it down fearing fear can also produce fear.


David Wallace-Wells uses the fear that scientists hide the severity to make his story more scary. He is claiming throughout that article that scientists are not giving it straight: scientist are technocrats who are too optimistic that the problem can be solved, "climate denialism has made scientists even more cautious", he talked to many scientists, but does not name most in the original article, thus suggesting they are only willing to tell the truth anonymously, "the many sober-minded scientists I interviewed over the past several months ... have quietly reached an apocalyptic conclusion", "Pollyannaish plant physiologists", "Climatologists are very careful when talking about Syria", "But climate scientists have a strange kind of faith: We will find a way to forestall radical warming, they say, because we must."

This framing may have made the article more attractive to readers who expect scientists not to be honest and to understate the problems. This in turn may have provoked a more allergic reaction to the Climate Feedback critique than an article mostly read by people who love and respect science.

That the audience matters is also suggested by clear difference in the responses on Reddit sceptic and Reddit Collapse. The people on the Reddit of the real sceptics (not the fake climate "sceptics") were interested, while the people preparing for the collapse of civilisation were more often unhappy about Climate Feedback.

In the Climate Feedback reviews the doomist tone is often not appreciated, but if you look at the details, at the annotations of the scientists, it is clear that the problem is that the NY Mag article contains errors that exaggerate the problem, not the bad news that is accurate. In the summaries spreading doom and exaggerating were sometimes used interchangeably. So let me say clearly: I have never talked to a scientists who was more worried in private than in public.

While scientists say what they think, they do tend to be careful in what they claim. The more careful the claim, the more confident a scientist can be that the evidence is sufficient to support it. We like strong and thus careful claims. This is justified when it comes to the question whether there is a problem. You do not want to cry wolf too often when there is none. However, as I have argued on this blog before we should not be careful about the size of the wolf. Saying the wolf is a Chihuahua is not good advice to the public.

My advice to the public would be to expect problems to be somewhat worse than the scientific mainstream claims, but not to go to prophets of doom and especially to avoid sources with a history of inaccurate information. (At least for mature problems, in case of fresh problem it can go both ways.)

Nitpicking

Looking at the high risk tails is uncomfortable for everyone, also for scientists. That provokes more critical reading and an unfortunate claim in such a story will get more comments than a similar one hidden in a middle of the road most accurate story.

However, unfortunately the article also often made statements are clearly inaccurate, wrong or are missing important context. The biggest error in the article — from my perspective as someone who works on how accurately we know how much the Earth is warming — was this line:
there are alarming stories every day, like last month’s satellite data showing the globe warming, since 1998, more than twice as fast as scientists had thought.
This has been updated by David Wallace-Wells to now read:
there are alarming stories in the news every day, like those, last month, that seemed to suggest satellite data showed the globe warming since 1998 more than twice as fast as scientists had thought (in fact, the underlying story was considerably less alarming than the headlines).
This was a report on a satellite upper air dataset that was the favourite of the climate "sceptics" because it showed the laast warming. Scientists have always warmed that that dataset was unreliable. Now an update has brought it in line with the other temperature datasets. The "twice as fast" is just for a cherry picked period. That is about as bad as mitigation sceptical claiming that global warming has stopped by cherry picking a specific period.


The scientific assessment for the actual warming did not change, certainly not become twice as much. If anything we now understand the problem better, which would mean less risk. The change is also not that much compared to the warming we had over the last century.

Because the actual scientific assessment did not change one could also argue that the mistake is inconsequential for the main argument of the story and the comment thus nitpicking. At least for me it matters. I hope more people feel this way.

There were many more mistakes like this and cases where missing context will give the reader the wrong impression. I do not want to go through them all in this already long post; you can read the annotations.

There were also cases which were also nitpicking from a scientific viewpoint. Where these comments were mine, I included them for completeness. They also did not influence my rating much.

Apparently I have to add that parts of the text without annotations are not automatically accurate. Especially for such a long article annotating is a lot of work. At a certain moment there are enough annotations to make an assessment. In addition even with 17 scientists it will happen that none of the scientists has relevant expertise for specific claims.

Climate Feedback rating system

We may want to have another look at the rating system used by Climate Feedback; see below. Normally finding a grade is quite straight forward. In this case I had to think long and was still not really satisfied.


Suggested guidelines for the overall scientific credibility rating
Remember that we do not evaluate the opinion of the author, but instead the scientific accuracy of facts contained within the text, and the scientific quality of reasoning used.
  • +2 = Very High: No inaccuracies, fairly represents the state of scientific knowledge, well argumented and documented, references are provided for key elements. The article provides insights to the reader about climate change mechanisms and implications.
  • +1 = High: The article does not contain major scientific inaccuracies and its conclusion follows from the evidence provided.
  • 0 = Neutral: No major inaccuracies, but no important insight to better explain implications of the science.
  • -1 = Low: The article contains significant scientific inaccuracies or misleading statements.
  • -2 = Very Low: The article contains major scientific inaccuracies for key facts supporting the author’s argumentation and/or omits to mention important information and/or presents logical flaws in using information to reach his or her conclusion.
  • n/a = Not Applicable: The article does not build on scientifically verifiable information (e.g., it is mostly about politics or opinions).


The scale is not symmetrical in the sense that if you get X facts wrong and X facts right you are in the middle. It might be that some people expect that. I would argue that getting only 50% right is pretty bad for a science article.

A problem in this case was that we can only give integer grades. So I gave a -1. I thought about neutral, but decided against it because that would mean "no major inaccuracies" and there were. For the part I could judge I found several mistakes and cases of missing important context (that is the description of -1). Had it been possible, I would have given the article a -0.5., because the tag "low scientific credibility" sounds a bit too harsh.

The rating of the article and the summary is made independently by all scientists and most made the same consideration. Had I been able to see the other "low" rating, I might have opted for "neutral" for balance. (We can see the annotations of the other scientists and can also respond to them. Sometimes when I am one of the first to make annotations, I wait with my rating to see what problems the others find.)

The relativity of wrong is very important. This same week Climate Feedback reviewed a Breitbart story about the accuracy of the instrumental warming estimate (my blog post on it). That was a complete con job and got "very low" rating. Giving the New York Magazine piece half the Breitbart rating does nor feel right. A scale from 0 to 4 may work better than one from -2 to 2. A zero sounds a lot worse than one and "low" would not be half of "very low".

The actual problem may be only -2 means inaccuracies influencing the main line of the story. It is more a scale for a science nerd looking for a high quality article than a scale for a citizen wanting to know how reliable the main line of the story is. Up to now that was mostly correlated, for this piece is was not, which made grading hard.

The more concrete a claim is, the more objective it can be assessed. Thus I would personally prefer to keep it a scale for science nerds and not go to a more vague and subjective assessment whether the main line is accurate.

A previous Feedback on a climate nightmare article by climate journalist Eric Holthaus got plus and minus ones. That shows that such an article can get positive ratings. That the article was never rated "neural" suggests that we may have to reconsider its description. That it states "no important insight" may make neutral almost worse than -1. It sounds like the famous quote: "not even wrong."



Based on the above discussion of the NY Mag review my suggestion for a new rating system would be the one below. It should be seen if it also fits well to other articles. I changed the numbers, the short descriptions and the long description for neutral.


Suggested guidelines for the overall scientific credibility rating
Remember that we do not evaluate the opinion of the author, but instead the scientific accuracy of facts contained within the text, and the scientific quality of reasoning used.
  • 4 = Excellent science reporting: No inaccuracies, fairly represents the state of scientific knowledge, well argumented and documented, references are provided for key elements. The article provides insights to the reader about climate change mechanisms and implications.
  • 3 = Very good science reporting: The article does not contain major scientific inaccuracies and its conclusion follows from the evidence provided.
  • 2 = Good science reporting: Mostly accurate statements and only minor inaccurate ones.
  • 1 = Some problems: The article contains significant scientific inaccuracies or misleading statements.
  • 0 = Major errors: The article contains major scientific inaccuracies for key facts supporting the author’s argumentation and/or omits to mention important information and/or presents logical flaws in using information to reach his or her conclusion.
  • n/a = Not Applicable: The article does not build on scientifically verifiable information (e.g., it is mostly about politics or opinions).


Why climate feedback?

There were people asking why we, Climate Feedback, were doing this. This could be interpreted in two ways:
1) Why do you nitpick this article I feel is an important wake-up call?
2) Why do you do this at all?

First of all, we do not know in advance what the outcome will be. Many articles on climate change are also very good and get great ratings. That is the kind of feedback journalists appreciate and which may help them in their careers and stimulate them to write better articles. Some journalists have even asked for reviews of important pieces to showcase the quality of their work.

We had a few authors who updated their article. David Wallace-Wells also did so and added more sources and transcripts. As far as I know such updates have only happened on the side that accepts the science. Also in that way our work improves science journalism, although such updates will come too late for most readers.

Most people will likely only get a general impression of how reliable news sources are when it comes to climate change. When we have enough reviews Climate Feedback will also make "official" assessments of the reliability of media sources. For more prolific writers also their individual credibility starts to become clear.

The reviews have already made clear that people who accept that climate change is real typically write accurate articles, while writers who do not want to solve the problem typically write error-ridden and misleading articles. That is good to know.

Only criticising the climate "sceptics" is not in the nature and in the training of good scientists. More utilitarian: solving climate change is a marathon, the energy transition will not be completed before 2050. Adaptation to limit the consequences of climate change will also be a job for generations to come. It is thus important that scientists are seen as trustworthy and only picking on one group would damage our reputation.

There are people who want to understand the details and when they meet misinformation being able to explain what is wrong with it. On reddit these people gather in /r/skeptic/. They normally accept climate change is real and like details/nitpicking, quality arguments and a rational world.


I do worry that there is also a downside to science policing in that people are less comfortable speaking about climate change fearing to be corrected. That is one reason to let minor cases slip and in bigger cases be gracious when it is the first time making a mistake. David Wallace-Wells responded graciously to the critiques, it were others that objected.

Making a mistake is completely different from the industrial production of nonsense on WUWT & Co.


It would be progress if scientists had a smaller role in this weird US "debate". It was forced on us by a continual stream of misinformation on the science from climate "sceptics". There should be a debate what to do about it and that is a debate for everyone. If you would see less scientists in US debates around climate change that would probably mean that the important questions are finally being addressed.

It would also be progress because scientists are typically not very good communicators. Partially that is because most scientists are introverted. Partially that is the nature of the problem, some things are simply not true, some arguments are simply not valid and there is little room for negotiation and graciousness.

On the other hand, science communication works pretty well in countries without the systemic corruption in Washington and the US media. So I do not think scientists are the main problem.

Related reading

Part I of this blog series: "How to talk about climate doomsday scenarios."

The updated New York Magazine piece By David Wallace-Wells: The Uninhabitable Earth - Famine, economic collapse, a sun that cooks us: What climate change could wreak — sooner than you think. (The reviewed original, the version with annotations.)

The Climate Feedback Feedback: Scientists explain what New York Magazine article on “The Uninhabitable Earth” gets wrong.

New York Magazine now also published extended interviews with the scientists interviewed for the piece: James Hansen, Peter Ward, Walley Broker, Michael Mann, and Michael Oppenheimer.

Introduction to Climate Feedback: Climate scientists are now grading climate journalism

Tuesday, 6 December 2016

Scott Adams: The Non-Expert Problem and Climate Change Science



Scott Adams, the creator of Dilbert, wrote today about how difficult it is for a non-expert to judge science and especially climate science. He argues that it is normally a good idea for a non-expert to follow the majority of scientists. I agree. Even as a scientist I do this for topics where I am not an expert and do not have the time to go into detail. You cannot live without placing trust and you should place your trust wisely.

While it is clear to Scott Adams that a majority of scientists agree on the basics of climate change, he worries that they still could all be wrong. He lists the below six signals that this could be the case and sees them in climate science. If you get your framing from the mitigation sceptical movement and only read the replies to their nonsense you may easily get his impression. So I thought it would be good to reply. It would be better to first understand the scientific basis, before venturing into the wild.

The terms Global Warming and Climate Change are both used for decades

Scott Adams assertion: It seems to me that a majority of experts could be wrong whenever you have a pattern that looks like this:

1. A theory has been “adjusted” in the past to maintain the conclusion even though the data has changed. For example, “Global warming” evolved to “climate change” because the models didn’t show universal warming.


This is a meme spread by the mitigation sceptics that is not based on reality. From the beginning both terms were used. One hint is name of the Intergovernmental Panel on Climate Change, a global group of scientists who synthesise the state of climate research and was created in 1988.

The irony of this strange meme is that it were the PR gurus of the US Republicans who told their politicians to use the term "climate change" rather than "global warming", because "global warming" was more scary. The video below shows the historical use of both terms.



Global warming was called global warming because the global average temperature is increasing, especially in the beginning there were still many regions were warming was not yet observed, while it was clear that the global average temperature was increasing. I use the term "global warming" if I want to emphasis the temperature change and the term "climate change" when I want to include all the other changes in the water cycle and circulation. These colleagues do the same and provide more history.

Talking about "adjusted", mitigation sceptics like to claim that temperature observations have been adjusted to show more warming. Truth is that the adjustments reduce global warming.

Climate models are not essential for basic understanding

Scott Adams assertion: 2. Prediction models are complicated. When things are complicated you have more room for error. Climate science models are complicated.

Yes, climate models are complicated. They synthesise a large part of our understanding of the climate system and thus play a large role in the synthesis of the IPCC. They are also the weakest part of climate science and thus a focus of the propaganda of the mitigation sceptical movement.

However, when it comes to the basics, climate model are not important. We know about the greenhouse effect for well over a century, long before we had any numerical climate models. That increasing the carbon dioxide concentration of the atmosphere leads to warming is clear, that this warming is amplified because warm air can contain more water, which is also a greenhouse gas, is also clear without any complicated climate model. This is very simple physics already used by Svante Arrhenius in the 19th century.

The warming effect of carbon dioxide can also be observed in the deep past. There are many reasons why the climate changes, but without carbon dioxide we can, for example, not understand the temperature swings of the past ice ages or why the Earth was able to escape from being completely frozen (Snowball Earth) at a time the sun was much dimmer.

The main role of climate models is trying to find reasons why the climate may respond differently this time than in the past or whether there are mechanisms beyond the simply physics that are important. The average climate sensitivity from climate models is about the same as for all the other lines of evidence. Furthermore, climate models add regional detail, especially when in comes to precipitation, evaporation and storms. These are helpful to better plan adaptation and estimate the impacts and costs, but are not central for the main claim that there is a problem.

Model tuning not important for basic understanding

Scott Adams assertion: 3. The models require human judgement to decide how variables should be treated. This allows humans to “tune” the output to a desired end. This is the case with climate science models.

Yes, models are tuned. Mostly not for the climatic changes, but to get the state of the atmosphere right, the global maps of clouds and precipitation, for example. In the light of my answer to point 2, this is not important for the question whether climate change is real.

The consensus is a result of the evidence

Scott Adams assertion: 4. There is a severe social or economic penalty for having the “wrong” opinion in the field. As I already said, I agree with the consensus of climate scientists because saying otherwise in public would be social and career suicide for me even as a cartoonist. Imagine how much worse the pressure would be if science was my career.

It is clearly not career suicide for a cartoonist. If you claim that you only accept the evidence because of social pressure, you are saying you do not really accept the evidence.

Scott Adams sounds as if he would like scientists to first freely pick a position and then only to look for evidence. In science it should go the other way around.

This seems to be the main argument and shows that Scott Adams knows more about office workers than about the scientific community. If science was your career and you would peddle the typical nonsense that comes from the mitigation sceptical movement that would indeed be bad for your career. In science you have to back up your claims with evidence. Cherry picking and making rookie errors to get the result you would like to get are not helpful.

However, if you present credible evidence that something is different, that is wonderful, that is why you become a scientist. I have been very critical of the quality of climate data and our methods to remove data problems. Contrary to Adams' expectation this has helped my career. Thus I cannot complain how climatology treats real skeptics. On the contrary, a lot of people supported me.

Another climate scientist, Eric Steig, strongly criticized the IPCC. He wrote about his experience:
I was highly critical of IPCC AR4 Chapter 6, so much so that the [mitigation skeptical] Heartland Institute repeatedly quotes me as evidence that the IPCC is flawed. Indeed, I have been unable to find any other review as critical as mine. I know "because they told me" that my reviews annoyed many of my colleagues, including some of my [RealClimate] colleagues, but I have felt no pressure or backlash whatsoever from it. Indeed, one of the Chapter 6 lead authors said “Eric, your criticism was really harsh, but helpful "thank you!"
If you have the evidence, there is nothing better than challenging the consensus. It is also the reason to become a scientist. As a scientist wrote on Slashdot:
Look, I'm a scientist. I know scientists. I know scientists at NOAA, NCAR, NIST, the Labs, in academia, in industry, at biotechs, at agri-science companies, at space exploration companies, and at oil and gas companies. I know conservative scientists, liberal scientists, agnostic scientists, religious scientists, and hedonistic scientists.

You know what motivates scientists? Science. And to a lesser extent, their ego. If someone doesn't love science, there's no way they can cut it as a scientist. There are no political or monetary rewards available to scientists in the same way they're available to lawyers and lobbyists.

Scientists consider and weigh all the evidence

Scott Adams assertion: 5. There are so many variables that can be measured – and so many that can be ignored – that you can produce any result you want by choosing what to measure and what to ignore. Our measurement sensors do not cover all locations on earth, from the upper atmosphere to the bottom of the ocean, so we have the option to use the measurements that fit our predictions while discounting the rest.

No, a scientist cannot produce any result they "want" and an average scientist would want to do good science and not get a certain result. The scientific mainstream is based on all the evidence we have. The mitigation sceptical movement behaves in the way Scott Adams expects and likes to cherry pick and mistreat data to get the results they want.

Arguments from the other side only look credible

Scott Adams assertion: 6. The argument from the other side looks disturbingly credible.

I do not know which arguments Adams is talking about, but the typical nonsense on WUWT, Breitbart, Daily Mail & Co. is made to look credible on the surface. But put on your thinking cap and it crumbles. At least check the sources. That reveals most of the problems very quickly.



For a scientist it is generally clear which arguments are valid, but it is indeed a real problem that to the public even the most utter nonsense may look "disturbingly credible". To help the public assess the credibility of claims and sources several groups are active.

Most of the zombie myths are debunked on RealClimate or Skeptical Science. If it is a recent WUWT post and you do not mind some snark you can often find a rebuttal the next day on HotWhopper. Media articles are regularly reviewed by Climate Feedback, a group of climate scientists, including me. They can only review a small portion of the articles, but it should be enough to determine which of the "sides" is "credible". If you claim you are sceptical, do use these resources and look at all sides of the argument and put in a little work to go in depth. If you do not do your due diligence to decide where to place your trust, you will get conned.



While political nonsense can be made to look credible, the truth is often complicated and sometimes difficult to convey. There is a big difference between qualified critique and uninformed nonsense. Valuing the strength of the evidence is part of the scientific culture. My critique of the quality of climate data has credible evidence behind it. There are also real scientific problems in understanding changes of clouds, as well as the land and vegetation. These are important for how much the Earth will respond, although in the long run the largest source of uncertainty is how much we will do to stop the problem.

There are real scientific problems when it comes to assessing the impacts of climate change. That often requires local or regional information, which is a lot more difficult than the global average. Many impacts will come from changes in severe weather, which are by definition rare and thus hard to study. For many impacts we need to know several changes at the same time. For droughts precipitation, temperature, humidity of the air and of the soil and insolation are all important. Getting them all right is hard.

How humans and societies will respond to the challenges posed by climate change is an even more difficult problem and beyond the realm of natural science. Not only the benefits, but also the costs of reducing greenhouse gas emissions are hard to predict. That would require predicting future technological, economic and social development.

When it comes to how big climate change itself and its impacts will be I am sure we will see surprises. What I do not understand is why some are arguing that this uncertainty is a reason to wait and see. The surprises will not only be nice, they will also be bad and all over increase the risks of climate change and make the case for solving this solvable problem stronger.




Related reading

Older post by a Dutch colleague on Adams' main problem: Who to believe?

How climatology treats sceptics

What's in a Name? Global Warming vs. Climate Change

Fans of Judith Curry: the uncertainty monster is not your friend

Video medal lecture Richard B. Alley at AGU: The biggest control knob: Carbon Dioxide in Earth's climate history

Just the facts, homogenization adjustments reduce global warming

Climate model ensembles of opportunity and tuning

Journalist Potholer makes excellent videos on climate change and true scepticism: Climate change explained, and the myths debunked


* Photo Arctic Sea Ice by NASA Goddard Space Flight Center used under a Creative Commons Attribution 2.0 Generic (CC BY 2.0) license.
* Cloud photo by Bill Dickinson used under a Creative Commons Attribution-NonCommercial-NoDerivs 2.0 Generic (CC BY-NC-ND 2.0) license.

Wednesday, 3 August 2016

Climate model ensembles of opportunity and tuning



Listen to grumpy old men.

As a young cloud researcher at a large conference, enthusiastic about almost any topic, I went to a town-hall meeting on using a large number of climate model runs to study how well we know what we know. Or as scientists call this: using a climate model ensemble to study confidence/uncertainty intervals.

Using ensembles was still quite new. Climate Prediction dot Net had just started asking citizens to run climate models on their Personal Computers (old big iPads) to get the computer power to create large ensembles. Studies using just one climate model run were still very common. The weather predictions on the evening television news were still based on one weather prediction model run; they still showed highs, lows and fronts on static "weather maps".

During the questions, a grumpy old men spoke up. He was far from enthusiastic about his new stuff. I see a Statler or Waldorf angrily swing his wooden walking stick in the air. He urged everyone, everyone to be very careful and not to equate the ensemble with a sample from a probability distribution. The experts dutifully swore they were fully aware of this.

They likely were and still are. But now everyone uses ensembles. Often using them as if they sample the probability distribution.

Before I wrote about the problems confusing model spread and uncertainty made in the now mostly dead "hiatus" debate. That debate remains important: after the hiatus debate is before the hiatus debate. The new hiatus is already 4 month old.* And there are so many datasets to select a "hiatus" from.


Fyfe et al. (2013) compared the temperature trend from the CMIP ensemble (grey histogram) to observations (red something) implicitly assuming that the model spread is the uncertainty. While the estimated trend is near the model spread, it is well within the uncertainty. The right panel is for a 20 year period: 1993–2012. The left panel starts in the cherry picked large El Nino year: 1998–2012.

This time I would like to explain better why the ensemble model spread is typically smaller than the confidence interval. These reasons suggest other questions where we need to pay attention: It is also important for comparing long-term historical model runs with observations and could affect some climate change impact studies. For long-term projections and decadal climate prediction it is likely less relevant.

Reasons why model spread is not uncertainty

One climate model run is just one realisation. Reality has the same problem. But you can run a model multiple times. If you change the model fields you begin with just a little bit, due to the chaotic nature of atmospheric and oceanic flows a second run will show a different realisation. The highs, lows and fronts will move differently, the ocean surface is consequently warmed and cooled at different times and places, internal modes such as El Nino will appear at different times. This chaotic behaviour is mainly found at the short time scales and is one reason for the spread of an ensemble. And it is one reason to expect that model spread is not uncertainty because models focus on getting the long term trend right and differ strongly when it comes to the internal variability.

But that is just reason one. The modules of a climate model that simulate specific physical processes have parameters that are based on measurements or more detailed models. We only know these parameters within some confidence interval. A normal climate model takes the best estimate of these parameters, but they could be anywhere within the confidence interval. To study how important these parameters are special "perturbed physics" ensembles are created where every model run has parameters that vary within the confidence interval.

Creating a such an ensemble is difficult. Depending on the reason for the uncertainty in the parameter, it could make sense to keep its value constant or to continually change it within its confidence interval and anything in between. It could make sense to keep the value constant over the entire Earth or to change it spatially and again anything in between. The parameter or how much it can fluctuate may dependent on the local weather or climate. It could be that when parameter X is high also parameter Y is high (or low); these dependencies should also be taken into account. Finally, also the distributions of the parameters needs to be realistic. Doing all of this for the large number of parameters in a climate model is a lot of work, typically only the most important ones are perturbed.

You can generate an ensemble that has too much spread by perturbing the parameters too strongly (and by making the perturbations too persistent). If you do it optimally, the ensemble would still show too little spread because not all physical processes are modelled because they are thought not to be important enough to justify the work and the computational resources. Part of this spread can be studied by making ensembles using many different models (multi-model ensemble), which are developed by different groups with different research questions and different ideas what is important.

That is where the title comes in: ensembles of opportunity. These are ensembles of existing model runs that were not created to be an ensemble. The most important example is the ensemble of the Coupled Models Intercomparison Project (CMIP). This group coordinates the creating of a set of climate model runs for similar scenarios, so that the results of these models can be compared with each other. This ensemble will automatically sample the chaotic flows and it is a multi-model ensemble, but it is not a perturbed physics ensemble; these model runs are model aiming at the best possible reproduction of what happened. For this reason alone the spread of the CMIP ensemble is expected to be too low.

The term "ensembles of opportunity" is another example the tendency of natural scientists to select neutral or generous terms to describe the work of colleagues. The term "makeshift ensemble" may be clearer.

Climate model tuning

The CMIP ensemble also has too little spread when it comes to the global mean temperature because the model are partially tuned to it. There is just an interesting readable article out on climate model tuning in BAMS**, which is intended for a general audience. Tuning has a large number of objectives, from getting the mean temperature right to the relationship between humidity and precipitation. There is also a section on tuning to the magnitude of warming the last century. It states about the historical runs:
The amplitude of the 20th century warming depends primarily on the magnitude of the radiative forcing, the climate sensitivity, as well as the efficiency of ocean heat uptake. ...

Some modeling groups claim not to tune their models against 20th century warming, however, even for model developers it is difficult to ensure that this is absolutely true in practice because of the complexity and historical dimension of model development. ...

There is a broad spectrum of methods to improve model match to 20th century warming, ranging from simply choosing to no longer modify the value of a sensitive parameter when a match is already good for a given model, or selecting physical parameterizations that improve the match, to explicitly tuning either forcing or feedback both of which are uncertain and depend critically on tunable parameters (Murphy et al. 2004; Golaz et al. 2013). Model selection could, for instance, consist of choosing to include or leave out new processes, such as aerosol cloud interactions, to help the model better match the historical warming, or choosing to work on or replace a parameterization that is suspected of causing a perceived unrealistically low or high forcing or climate sensitivity.
Due to tuning models that have a low climate sensitivity tend to have stronger forcings over the last century and model with a high climate sensitivity a weaker forcing. The forcing due to greenhouse gasses does not vary much, that part is easy. The forcings due to small particles in the air (aerosols) that like CO2 stem from the burning of fossil fuels and are quite uncertain and Kiehl (2007) showed that high sensitivity models tend to have more cooling due to aerosols. For a more nuanced updated story see Knutti et al. (2008) and Forster et al. (2013).


Kiehl (2007) found an inverse correlation between forcing and climate sensitivity. The main reason for the differences in forcing was the cooling by aerosols.
This "tuning" initially was not an explicit tuning of model parameters, but mostly because modellers keep working until the results look good. Look good compared to observations. Bjorn Stevens talks about this in an otherwise also recommendable Forecast episode.

Nowadays the tuning is often performed more formally and an important part of studying the climate models and understanding their uncertainties. The BAMS article proposes to collect information on tuning for the upcoming CMIP. In principle a good idea, but I do not think that that is enough. In a simple example of climate sensitivity and aerosol forcing, the groups with low sensitivity and forcing and the ones with high sensitivity and forcing are happy with their temperature trend and will report not to have tuned. But that choice also leads to too little ensemble spread, just like the groups that did need to tune. Tuning makes it complicated to interpret the ensemble, it is no problem for a specific model run.

Given that we know the temperature increase, it is impossible not to get a tuned result. Furthermore, I mention several additional reasons why the model spread is not the uncertainty above that complicate the interpretation of the ensemble in the same way. A solution could be to follow the work in ensemble weather prediction with perturbed-physics ensembles and to tune all models, but to tune them to cover the full range of uncertainties that we estimate from the observations. This should at least cover the the climate sensitivity and ocean heat uptake, but preferably also other climate characteristics that are important for climate impact and climate variability studies. Large modelling centres may be able to create such large ensembles by themselves, the others could coordinate their work in CMIP to make sure the full uncertainty range is covered.

Historical climate runs

Because the physics is not perturbed and especially due to the tuning, you would expect that the CMIP ensemble spread is too low for global mean temperature increase. That the CMIP ensemble average fits well to the observed temperature increase shows that with reasonable physical choices we can understand why the temperature increased. It shows that known processes are sufficient to explain it. That is fits so accurately, does not say much. I liked the title of an article from Reto Knutti (2008): "Why are climate models reproducing the observed global surface warming so well?" Which implies it all.

Much more interesting to study how good models are, are spatial patterns and other observations. New datasets are greeted with much enthusiasm by modellers because they allow for the best comparison and are more likely to show new problems that need fixing and lead to a better understanding. Also model results for the deep past are important tests, which models are not tuned for.


That the CMIP ensemble mean fits to the observations is no reason to expect that the observations are reliable


When the observations peak out of this too narrow CMIP ensemble spread that is to be expected. If you want to make a case that our understanding does not fit to the observations, you have to take the uncertainties into account, not the spread.

Similarly, that the CMIP ensemble mean fits to the observations is no reason to expect that the observations are reliable. Because of the overconfidence in the data quality also many scientists took the recent minimal deviations from the trend line too seriously. This finally stimulated more research into the accuracy of temperature trends, into inhomogeneities in the ERSST sea surface temperatures, into the effect of coverage and how we blend sea, land and ice temperatures together. There are some more improvements under way.

Compared to the global warming of about 1°C up to now, these recent and upcoming corrections are large. Many of the problem could have been found long ago. It is 2016. It is about time to study this. If funding is an issue we could maybe sacrifice some climate change impact studies for wine. Or for truffles. Or caviar. The quality of our data is the foundation of our science.

That the comparison of the CMIP ensemble average with the instrumental observation is so central to the public climate "debate" is rather ironic. Please take a walk in the forest. Look at all the different changes. The ones that go slower as well as the many that go faster than expected.

Maybe it is good to emphasise that for the attribution of climate change to human activities, the size of the historical temperature increase is not used. The attribution is made via correlations with the 3-dimensional spatial patterns between observations and models. By using the correlations (rather than root mean square errors), the magnitude of the change in either the models or the observations is no longer important. Ribes (2016) is working on using the magnitude of the changes as well. This is difficult because of inevitable tuning, which makes specifying the uncertainties very difficult.

Climate change impact studies

Studying the impacts of climate change is hard. Whether dikes break depends not only on sea level rise, but also on the changes in storms. The maintenance of the dikes and the tides are important. It matters whether you have a functioning government that also takes care of problems that only become apparent when the catastrophe happens. I would not sleep well if I lived in an area where civil servants are not allowed to talk about climate change. Because of the additional unnecessary climate dangers, but especially because that is a clear sign of a dysfunctional government that does not prioritise protecting its people.

The too narrow CMIP ensemble spread can lead to underestimates of the climate change impacts because typically the higher damages from stronger than expected changes are larger than the reduced damages from smaller changes. The uncertainty monster is not our friend. Admittedly, the effect of the uncertainties is rather modest. This is only important for those impacts we understand reasonably well already. The lack of variability can be partially solved in the statistical post-processing (bias correction and downscaling). This is not common yet, but Grenier et al. (2015) proposed a statistical method to make the natural variability more realistic.

This problem will hopefully soon be solved when the research programs on decadal climate prediction mature. The changes over a decade due to greenhouse warming are modest, for decadal prediction we thus especially need to accurately predict the natural variability of the climate system. An important part of these studies is assessing whether and which changes can be predicted. As a consequence there is a strong focus on situation specific uncertainties and statistical post-processing to correct biases of the model ensemble in the means and in the uncertainties.

In the tropics decadal climate prediction works reasonably well and helps farmers and governments in their planning.



In the mid-latitudes, where most of the researchers live, it is frustratingly difficult to make decadal predictions. Still even in that case, we would still have an ensemble where the ensemble can be used as a sample of the probability distribution. That is important progress.

When a lack of ensemble spread is a problem for historical runs, you might expect it to be a problem for projecting for the rest of the century. This is probably not the case. The problem of tuning would be much reduced because the influence of aerosols will be much smaller as the signal of greenhouse gasses becomes much more dominant. For long term projections the main factor is that the climate sensitivity of the models needs to fit to our understanding of the climate sensitivity from all studies. This fit is reasonable for the best estimate of the climate sensitivity, which we expect to be 3°C for a doubling of the CO2 concentration. I do not know how well the fit is for the spread in the climate sensitivity.

However, for long-term projections even the climate sensitivity is not that important. For the magnitude of the climatic changes in 2100 and for the impact of climate change in 2100, the main source of uncertainty is what we will do. As you can see in the figure below the difference between a business as usual scenario and strong climate policies is 3 °C (6 °F). The uncertainties within these scenario's is relatively small. Thus the main question is whether and how aggressively we will act to combat climate change.





Related information

Is it time to freak out about the climate sensitivity estimates from energy budget models?

Fans of Judith Curry: the uncertainty monster is not your friend

Are climate models running hot or observations running cold?

Forecast: Gavin Schmidt on the evolution, testing and discussion of climate models

Forecast: Bjorn Stevens on the philosophy of climate modeling

The Guardian: In a blind test, economists reject the notion of a global warming pause

Reference

Forster, P.M., T. Andrews, P. Good, J.M. Gregory, L.S. Jackson, and M. Zelinka, 2013: Evaluating adjusted forcing and model spread for historical and future scenarios in the CMIP5 generation of climate models. Journal of Geophysical Research, 118, 1139–1150, doi: 10.1002/jgrd.50174.

Fyfe, John C., Nathan P. Gillett and Francis W. Zwiers, 2013: Overestimated global warming over the past 20 years. Nature Climate Change, 3, pp. 767–769, doi: 10.1038/nclimate1972.

Golaz, J.-C., J.-C. Golaz, and H. Levy, 2013: Cloud tuning in a coupled climate model: Impact on 20th century warming. Geophysical Research Letters, 40, pp. 2246–2251, doi: 10.1002/grl.50232.

Grenier, Patrick, Diane Chaumont and Ramón de Elía, 2015: Statistical adjustment of simulated inter-annual variability in an investigation of short-term temperature trend distributions over Canada. EGU general meeting, Vienna, Austria.

Hourdin, Frederic, Thorsten Mauritsen, Andrew Gettelman, Jean-Christophe Golaz, Venkatramani Balaji, Qingyun Duan, Doris Folini, Duoying Ji, Daniel Klocke, Yun Qian, Florian Rauser, Cathrine Rio, Lorenzo Tomassini, Masahiro Watanabe, and Daniel Williamson, 2016: The art and science of climate model tuning. Bulletin of the American Meteorological Society, published online, doi: 10.1175/BAMS-D-15-00135.1.

Kiehl, J.T., 2007: Twentieth century climate model response and climate sensitivity. Geophysical Research Letters, 34, L22710, doi: 10.1029/2007GL031383.

Knutti, R., 2008: Why are climate models reproducing the observed global surface warming so well? Geophysical Research Letters, 35, L18704, doi: 10.1029/2008GL034932.

Murphy, J.M., D.M.H. Sexton, D.N. Barnett, G.S. Jones, M.J. Webb, M. Collins and D.A. Stainforth, 2004: Quantification of modelling uncertainties in a large ensemble of climate change simulations. Nature, 430, pp. 768–772, doi: 10.1038/nature02771.

Ribes, A., 2016: Multi-model detection and attribution without linear regression. 13th International Meeting on Statistical Climatology, Canmore, Canada. Abstract below.

Rowlands, Daniel J., David J. Frame, Duncan Ackerley, Tolu Aina, Ben B. B. Booth, Carl Christensen, Matthew Collins, Nicholas Faull, Chris E. Forest, Benjamin S. Grandey, Edward Gryspeerdt, Eleanor J. Highwood, William J. Ingram, Sylvia Knight, Ana Lopez, Neil Massey, Frances McNamara, Nicolai Meinshausen, Claudio Piani, Suzanne M. Rosier, Benjamin M. Sanderson, Leonard A. Smith, Dáithí A. Stone, Milo Thurston, Kuniko Yamazaki, Y. Hiro Yamazaki & Myles R. Allen, 2012: Broad range of 2050 warming from an observationally constrained large climate model ensemble. Nature Geoscience, 5, pp. 256–260, doi: 10.1038/ngeo1430 (manuscript).


MULTI-MODEL DETECTION AND ATTRIBUTION WITHOUT LINEAR REGRESSION
Aurélien Ribes
Abstract. Conventional D&A statistical methods involve linear regression models where the observations are regressed onto expected response patterns to different external forcings. These methods do not use physical information provided by climate models regarding the expected response magnitudes to constrain the estimated responses to the forcings. As an alternative to this approach, we propose a new statistical model for detection and attribution based only on the additivity assumption. We introduce estimation and testing procedures based on likelihood maximization. As the possibility of misrepresented response magnitudes is removed in this revised statistical framework, it is important to take the climate modelling uncertainty into account. In this way, modelling uncertainty in the response magnitude and the response pattern is treated consistently. We show that climate modelling uncertainty can be accounted for easily in our approach. We then provide some discussion on how to practically estimate this source of uncertainty, and on the future challenges related to multi-model D&A in the framework of CMIP6/DAMIP.


* Because this is the internet, let me say that "The new hiatus is already 4 month old." is a joke.

** The BAMS article calls any way to estimate a parameter "tuning". I would personally only call it tuning if you optimize for emerging properties of the climate model. If you estimate a parameter based on observations or a specialized model, I would not call this tuning, but simply parameter estimation or parameterization development. Radiative transfer schemes use the assumption that adjacent layers of clouds are maximally overlapped and that if there is a clear layer between two cloud layers that they are random overlapped. You could introduce two parameters that vary between maximum and random for these two cases, but that is not done. You could call that an implicit parameter, which shows that distinguishing between parameter estimation and parameterization development is hard.

*** Photo at the top: Grumpy Tortoise Face by Eric Kilby, used under a Attribution-ShareAlike 2.0 Generic (CC BY-SA 2.0) license.

Climate model ensembles of opportunity and tuning



Listen to grumpy old men.

As a young cloud researcher at a large conference, enthusiastic about almost any topic, I went to a town-hall meeting on using a large number of climate model runs to study how well we know what we know. Or as scientists call this: using a climate model ensemble to study confidence/uncertainty intervals.

Using ensembles was still quite new. Climate Prediction dot Net had just started asking citizens to run climate models on their Personal Computers (old big iPads) to get the computer power to create large ensembles. Studies using just one climate model run were still very common. The weather predictions on the evening television news were still based on one weather prediction model run; they still showed highs, lows and fronts on static "weather maps".

During the questions, a grumpy old men spoke up. He was far from enthusiastic about his new stuff. I see a Statler or Waldorf angrily swing his wooden walking stick in the air. He urged everyone, everyone to be very careful and not to equate the ensemble with a sample from a probability distribution. The experts dutifully swore they were fully aware of this.

They likely were and still are. But now everyone uses ensembles. Often using them as if they sample the probability distribution.

Before I wrote about the problems confusing model spread and uncertainty made in the now mostly dead "hiatus" debate. That debate remains important: after the hiatus debate is before the hiatus debate. The new hiatus is already 4 month old.* And there are so many datasets to select a "hiatus" from.


Fyfe et al. (2013) compared the temperature trend from the CMIP ensemble (grey histogram) to observations (red something) implicitly assuming that the model spread is the uncertainty. While the estimated trend is near the model spread, it is well within the uncertainty. The right panel is for a 20 year period: 1993–2012. The left panel starts in the cherry picked large El Nino year: 1998–2012.

This time I would like to explain better why the ensemble model spread is typically smaller than the confidence interval. These reasons suggest other questions where we need to pay attention: It is also important for comparing long-term historical model runs with observations and could affect some climate change impact studies. For long-term projections and decadal climate prediction it is likely less relevant.

Reasons why model spread is not uncertainty

One climate model run is just one realisation. Reality has the same problem. But you can run a model multiple times. If you change the model fields you begin with just a little bit, due to the chaotic nature of atmospheric and oceanic flows a second run will show a different realisation. The highs, lows and fronts will move differently, the ocean surface is consequently warmed and cooled at different times and places, internal modes such as El Nino will appear at different times. This chaotic behaviour is mainly found at the short time scales and is one reason for the spread of an ensemble. And it is one reason to expect that model spread is not uncertainty because models focus on getting the long term trend right and differ strongly when it comes to the internal variability.

But that is just reason one. The modules of a climate model that simulate specific physical processes have parameters that are based on measurements or more detailed models. We only know these parameters within some confidence interval. A normal climate model takes the best estimate of these parameters, but they could be anywhere within the confidence interval. To study how important these parameters are special "perturbed physics" ensembles are created where every model run has parameters that vary within the confidence interval.

Creating a such an ensemble is difficult. Depending on the reason for the uncertainty in the parameter, it could make sense to keep its value constant or to continually change it within its confidence interval and anything in between. It could make sense to keep the value constant over the entire Earth or to change it spatially and again anything in between. The parameter or how much it can fluctuate may dependent on the local weather or climate. It could be that parameter X is high also parameter Y is high (or low); these dependencies should also be taken into account. Finally, also the distributions of the parameters needs to be realistic. Doing all of this for the large number of parameters in a climate model is a lot of work, typically only the most important ones are perturbed.

You can generate an ensemble that has too much spread by perturbing the parameters too strongly (and by making the perturbations too persistent). If you do it optimally, the ensemble would still show too little spread because not all physical processes are modelled because they are thought not to be important enough to justify the work and the computational resources. Part of this spread can be studied by making ensembles using many different models (multi-model ensemble), which are developed by different groups with different research questions and different ideas what is important.

That is where the title comes in: ensembles of opportunity. These are ensembles of existing model runs that were not created to be an ensemble. The most important example is the ensemble of the Coupled Models Intercomparison Project (CMIP). This group coordinates the creating of a set of climate model runs for similar scenarios, so that the results of these models can be compared with each other. This ensemble will automatically sample the chaotic flows and it is a multi-model ensemble, but it is not a perturbed physics ensemble; these model runs are model aiming at the best possible reproduction of what happened. For this reason alone the spread of the CMIP ensemble is expected to be too low.

The term "ensembles of opportunity" is another example the tendency of natural scientists to select neutral or generous terms to describe the work of colleagues. The term "makeshift ensemble" may be clearer.

Climate model tuning

The CMIP ensemble also has too little spread when it comes to the global mean temperature because the model are partially tuned to it. There is just an interesting readable article out on climate model tuning in BAMS**, which is intended for a general audience. Tuning has a large number of objectives, from getting the mean temperature right to the relationship between humidity and precipitation. There is also a section on tuning to the magnitude of warming the last century. It states about the historical runs:
The amplitude of the 20th century warming depends primarily on the magnitude of the radiative forcing, the climate sensitivity, as well as the efficiency of ocean heat uptake. ...

Some modeling groups claim not to tune their models against 20th century warming, however, even for model developers it is difficult to ensure that this is absolutely true in practice because of the complexity and historical dimension of model development. ...

There is a broad spectrum of methods to improve model match to 20th century warming, ranging from simply choosing to no longer modify the value of a sensitive parameter when a match is already good for a given model, or selecting physical parameterizations that improve the match, to explicitly tuning either forcing or feedback both of which are uncertain and depend critically on tunable parameters (Murphy et al. 2004; Golaz et al. 2013). Model selection could, for instance, consist of choosing to include or leave out new processes, such as aerosol cloud interactions, to help the model better match the historical warming, or choosing to work on or replace a parameterization that is suspected of causing a perceived unrealistically low or high forcing or climate sensitivity.
Due to tuning models that have a low climate sensitivity tend to have stronger forcings over the last century and model with a high climate sensitivity a weaker forcing. The forcing due to greenhouse gasses does not vary much, that part is easy. The forcings due to small particles in the air (aerosols) that like CO2 stem from the burning of fossil fuels and are quite uncertain and Kiehl (2007) showed that high sensitivity models tend to have more cooling due to aerosols. For a more nuanced updated story see Knutti et al. (2008) and Forster et al. (2013).


Kiehl (2007) found an inverse correlation between forcing and climate sensitivity. The main reason for the differences in forcing was the cooling by aerosols.
This "tuning" initially was not an explicit tuning of model parameters, but mostly because modellers keep working until the results look good. Look good compared to observations. Bjorn Stevens talks about this in an otherwise also recommendable Forecast episode.

Nowadays the tuning is often performed more formally and an important part of studying the climate models and understanding their uncertainties. The BAMS article proposes to collect information on tuning for the upcoming CMIP. In principle a good idea, but I do not think that that is enough. In a simple example of climate sensitivity and aerosol forcing, the groups with low sensitivity and forcing and the ones with high sensitivity and forcing are happy with their temperature trend and will report not to have tuned. But that choice also leads to too little ensemble spread, just like the groups that did need to tune. Tuning makes it complicated to interpret the ensemble, it is no problem for a specific model run.

Given that we know the temperature increase, it is impossible not to get a tuned result. Furthermore, I mention several additional reasons why the model spread is not the uncertainty above that complicate the interpretation of the ensemble in the same way. A solution could be to follow the work in ensemble weather prediction with perturbed-physics ensembles and to tune all models, but to tune them to cover the full range of uncertainties that we estimate from the observations. This should at least cover the the climate sensitivity and ocean heat uptake, but preferably also other climate characteristics that are important for climate impact and climate variability studies. Large modelling centres may be able to create such large ensembles by themselves, the others could coordinate their work in CMIP to make sure the full uncertainty range is covered.

Historical climate runs

Because the physics is not perturbed and especially due to the tuning, you would expect that the CMIP ensemble spread is too low for global mean temperature increase. That the CMIP ensemble average fits well to the observed temperature increase shows that with reasonable physical choices we can understand why the temperature increased. It shows that known processes are sufficient to explain it. That is fits so accurately, does not say much. I liked the title of an article from Reto Knutti (2008): "Why are climate models reproducing the observed global surface warming so well?" Which implies it all.

Much more interesting to study how good models are, are spatial patterns and other observations. New datasets are greeted with much enthusiasm by modellers because they allow for the best comparison and are more likely to show new problems that need fixing and lead to a better understanding. Also model results for the deep past are important tests, which models are not tuned for.


That the CMIP ensemble mean fits to the observations is no reason to expect that the observations are reliable


When the observations peak out of this too narrow CMIP ensemble spread that is to be expected. If you want to make a case that our understanding does not fit to the observations, you have to take the uncertainties into account, not the spread.

Similarly, that the CMIP ensemble mean fits to the observations is no reason to expect that the observations are reliable. Because of the overconfidence in the data quality also many scientists took the recent minimal deviations from the trend line too seriously. This finally stimulated more research into the accuracy of temperature trends, into inhomogeneities in the ERSST sea surface temperatures, into the effect of coverage and how we blend sea, land and ice temperatures together. There are some more improvements under way.

Compared to the global warming of about 1°C up to now, these recent and upcoming corrections are large. Many of the problem could have been found long ago. It is 2016. It is about time to study this. If funding is an issue we could maybe sacrifice some climate change impact studies for wine. Or for truffles. Or caviar. The quality of our data is the foundation of our science.

That the comparison of the CMIP ensemble average with the instrumental observation is so central to the public climate "debate" is rather ironic. Please take a walk in the forest. Look at all the different changes. The ones that go slower as well as the many that go faster than expected.

Maybe it is good to emphasise that for the attribution of climate change to human activities, the size of the historical temperature increase is not used. The attribution is made via correlations with the 3-dimensional spatial patterns between observations and models. By using the correlations (rather than root mean square errors), the magnitude of the change in either the models or the observations is no longer important. Ribes (2016) is working on using the magnitude of the changes as well. This is difficult because of inevitable tuning, which makes specifying the uncertainties very difficult.

Climate change impact studies

Studying the impacts of climate change is hard. Whether dikes break depends not only on sea level rise, but also on the changes in storms. The maintenance of the dikes and the tides are important. It matters whether you have a functioning government that also takes care of problems that only become apparent when the catastrophe happens. I would not sleep well if I lived in an area where civil servants are not allowed to talk about climate change. Because of the additional unnecessary climate dangers, but especially because that is a clear sign of a dysfunctional government that does not prioritise protecting its people.

The too narrow CMIP ensemble spread can lead to underestimates of the climate change impacts because typically the higher damages from stronger than expected changes are larger than the reduced damages from smaller changes. The uncertainty monster is not our friend. Admittedly, the effect of the uncertainties is rather modest. This this is only important for those impacts we understand reasonably well already. The lack of variability can be partially solved in the statistical post-processing (bias correction and downscaling). This is not common yet, but Grenier et al. (2015) proposed a statistical method to make the natural variability more realistic.

This problem will hopefully soon be solved when the research programs on decadal climate prediction mature. The changes over a decade due to greenhouse warming are modest, for decadal prediction we thus especially need to accurately predict the natural variability of the climate system. An important part of these studies is assessing whether and which changes can be predicted. As a consequence there is a strong focus on situation specific uncertainties and statistical post-processing to correct biases of the model ensemble in the means and in the uncertainties.

In the tropics decadal climate prediction works reasonably well and helps farmers and governments in their planning.



In the mid-latitudes, where most of the researchers live, it is frustratingly difficult to make decadal predictions. Still even in that case, we would still have an ensemble where the ensemble can be used as a sample of the probability distribution. That is important progress.

When a lack of ensemble spread is a problem for historical runs, you might expect it to be a problem for projecting for the rest of the century. This is probably not the case. The problem of tuning would be much reduced because the influence of aerosols will be much smaller as the signal of greenhouse gasses becomes much more dominant. For long term projections the main factor is that the climate sensitivity of the models needs to fit to our understanding of the climate sensitivity from all studies. This fit is reasonable for the best estimate of the climate sensitivity, which we expect to be 3°C for a doubling of the CO2 concentration. I do not know how well the fit is for the spread in the climate sensitivity.

However, for long-term projections even the climate sensitivity is not that important. For the magnitude of the climatic changes in 2100 and for the impact of climate change in 2100, the main source of uncertainty is what we will do. As you can see in the figure below the difference between a business as usual scenario and strong climate policies is 3 °C (6 °F). The uncertainties within these scenario's is relatively small. Thus the main question is whether and how aggressively we will act to combat climate change.





Related information

New article (september 2017): Gavin Schmidt et al.: Practice and philosophy of climate model tuning across six US modeling centers.

Discussion paper suggesting a path to solving the difference between model spread and uncertainty by James Annan and Julia Hargreaves: On the meaning of independence in climate science.

Is it time to freak out about the climate sensitivity estimates from energy budget models?

Fans of Judith Curry: the uncertainty monster is not your friend.

Are climate models running hot or observations running cold?

Forecast: Gavin Schmidt on the evolution, testing and discussion of climate models.

Forecast: Bjorn Stevens on the philosophy of climate modeling.

The Guardian: In a blind test, economists reject the notion of a global warming pause.

Reference

Forster, P.M., T. Andrews, P. Good, J.M. Gregory, L.S. Jackson, and M. Zelinka, 2013: Evaluating adjusted forcing and model spread for historical and future scenarios in the CMIP5 generation of climate models. Journal of Geophysical Research, 118, 1139–1150, doi: 10.1002/jgrd.50174.

Fyfe, John C., Nathan P. Gillett and Francis W. Zwiers, 2013: Overestimated global warming over the past 20 years. Nature Climate Change, 3, pp. 767–769, doi: 10.1038/nclimate1972.

Golaz, J.-C., J.-C. Golaz, and H. Levy, 2013: Cloud tuning in a coupled climate model: Impact on 20th century warming. Geophysical Research Letters, 40, pp. 2246–2251, doi: 10.1002/grl.50232.

Grenier, Patrick, Diane Chaumont and Ramón de Elía, 2015: Statistical adjustment of simulated inter-annual variability in an investigation of short-term temperature trend distributions over Canada. EGU general meeting, Vienna, Austria.

Hourdin, Frederic, Thorsten Mauritsen, Andrew Gettelman, Jean-Christophe Golaz, Venkatramani Balaji, Qingyun Duan, Doris Folini, Duoying Ji, Daniel Klocke, Yun Qian, Florian Rauser, Cathrine Rio, Lorenzo Tomassini, Masahiro Watanabe, and Daniel Williamson, 2016: The art and science of climate model tuning. Bulletin of the American Meteorological Society, published online, doi: 10.1175/BAMS-D-15-00135.1.

Kiehl, J.T., 2007: Twentieth century climate model response and climate sensitivity. Geophysical Research Letters, 34, L22710, doi: 10.1029/2007GL031383.

Knutti, R., 2008: Why are climate models reproducing the observed global surface warming so well? Geophysical Research Letters, 35, L18704, doi: 10.1029/2008GL034932.

Murphy, J.M., D.M.H. Sexton, D.N. Barnett, G.S. Jones, M.J. Webb, M. Collins and D.A. Stainforth, 2004: Quantification of modelling uncertainties in a large ensemble of climate change simulations. Nature, 430, pp. 768–772, doi: 10.1038/nature02771.

Ribes, A., 2016: Multi-model detection and attribution without linear regression. 13th International Meeting on Statistical Climatology, Canmore, Canada. Abstract below.

Rowlands, Daniel J., David J. Frame, Duncan Ackerley, Tolu Aina, Ben B. B. Booth, Carl Christensen, Matthew Collins, Nicholas Faull, Chris E. Forest, Benjamin S. Grandey, Edward Gryspeerdt, Eleanor J. Highwood, William J. Ingram, Sylvia Knight, Ana Lopez, Neil Massey, Frances McNamara, Nicolai Meinshausen, Claudio Piani, Suzanne M. Rosier, Benjamin M. Sanderson, Leonard A. Smith, Dáithí A. Stone, Milo Thurston, Kuniko Yamazaki, Y. Hiro Yamazaki & Myles R. Allen, 2012: Broad range of 2050 warming from an observationally constrained large climate model ensemble. Nature Geoscience, 5, pp. 256–260, doi: 10.1038/ngeo1430 (manuscript).


MULTI-MODEL DETECTION AND ATTRIBUTION WITHOUT LINEAR REGRESSION
Aurélien Ribes
Abstract. Conventional D&A statistical methods involve linear regression models where the observations are regressed onto expected response patterns to different external forcings. These methods do not use physical information provided by climate models regarding the expected response magnitudes to constrain the estimated responses to the forcings. As an alternative to this approach, we propose a new statistical model for detection and attribution based only on the additivity assumption. We introduce estimation and testing procedures based on likelihood maximization. As the possibility of misrepresented response magnitudes is removed in this revised statistical framework, it is important to take the climate modelling uncertainty into account. In this way, modelling uncertainty in the response magnitude and the response pattern is treated consistently. We show that climate modelling uncertainty can be accounted for easily in our approach. We then provide some discussion on how to practically estimate this source of uncertainty, and on the future challenges related to multi-model D&A in the framework of CMIP6/DAMIP.


* Because this is the internet, let me say that "The new hiatus is already 4 month old." is a joke.

** The BAMS article calls any way to estimate a parameter "tuning". I would personally only call it tuning if you optimize for emerging properties of the climate model. If you estimate a parameter based on observations or a specialized model, I would not call this tuning, but simply parameter estimation or parameterization development. Radiative transfer schemes use the assumption that adjacent layers of clouds are maximally overlapped and that if there is a clear layer between two cloud layers that they are random overlapped. You could introduce two parameters that vary between maximum and random for these two cases, but that is not done. You could call that an implicit parameter, which shows that distinguishing between parameter estimation and parameterization development is hard.

*** Photo at the top: Grumpy Tortoise Face by Eric Kilby, used under a Attribution-ShareAlike 2.0 Generic (CC BY-SA 2.0) license.