Understanding CIBSE Weather Files for Energy and Overheating Analysis
How CIBSE builds TRY and DSY weather files, what morphing to UKCP18 does to them, and how to pick the right file for an energy or overheating assessment.
We work with building performance simulation almost every day in mechanical engineering. One of the most important inputs we use is weather data. Without good weather files, our energy calculations and overheating assessments simply cannot be trusted. In the UK, the Chartered Institution of Building Services Engineers (CIBSE) provides the standard set of hourly weather files that we rely on. These files come mainly in EPW format and are designed specifically for energy and overheating studies.
In this article, we will walk through the 2025 CIBSE Weather Data release, which is based on UKCP18 climate projections. We will explain what the different file types mean, how they are created, and how you should choose the right one for your project. The goal is to help you understand the files clearly so you can use them with confidence.
The Two Main Families of Weather Files
CIBSE gives us two main families of weather files. Each one serves a different purpose.
Test Reference Year (TRY) We use TRY files when we want to understand the typical or average performance of a building over a full year. These are the files we turn to for annual energy assessments and for calculations under Part L of the Building Regulations.
Design Summer Year (DSY) We use DSY files when we need to check how a building will behave during hot weather. These files are essential for overheating risk assessments under TM59, TM52 and Part O.
The 2025 release covers 28 climate zones across the UK. That gives us much better local detail than earlier sets.
How a Test Reference Year (TRY) Is Built
A TRY is not a real continuous year of weather. It is a carefully constructed composite year that represents typical conditions.
Here is the process CIBSE follows:
- They start with 30 years of real observational data (for the 2025 release this is 1994 to 2023).
- They look at every January from those 30 years, every February, and so on for each calendar month.
- For each month they calculate daily averages of the key weather variables such as dry-bulb temperature, solar radiation, humidity and wind.
- They build a long-term Cumulative Distribution Function that shows what “average” looks like for that month.
- They score every individual month using the Finkelstein-Schafer statistic. This statistic simply measures how closely a particular month matches the long-term average pattern.
- They pick the single most typical January, the most typical February, and so on for all twelve months.
- They stitch those twelve selected months together and smooth the joins so the weather does not jump unrealistically from one month to the next.
- Any missing data is carefully interpolated.
The result is a year that looks statistically normal. Extremes are deliberately kept under control. That is why a TRY is excellent for energy calculations but not suitable for overheating studies.
The build is easier to follow when you watch it happen. The model below runs the whole process on a synthetic 30 year record, one calendar month at a time.
Building a Test Reference Year out of twelve real months
Each column scores all 30 Januaries, all 30 Februaries, and so on. The most typical month in each column wins, and the twelve winners are stitched into one year. The record here is synthetic, not CIBSE data.
Turn the smoothing off and the largest step across a month boundary grows. That step is a day where the weather jumps from one year straight into another, which is exactly what the smoothing removes.
What the Finkelstein-Schafer statistic actually measures
The name sounds harder than the idea. We are comparing two curves.
The first curve is the long-term one. We take every day of every January in the 30 year record, sort the values, and plot the fraction of days at or below each temperature. That is the Cumulative Distribution Function for January.
The second curve is the same thing for one candidate January on its own.
The statistic is the average vertical distance between the two curves:
Here are the daily values of the candidate month, is the long-term cumulative distribution for that calendar month, and is the candidate month’s own. CIBSE scores several variables this way, not temperature alone, and combines the scores.
A candidate that sits on top of the long-term curve has small gaps and scores near zero. A month that ran warm or cold pulls its curve sideways, the gaps open up, and the score rises. Only the lowest score is kept.
The Finkelstein-Schafer statistic: how typical is this month?
The grey curve is the cumulative distribution of every day in this calendar month across all 30 years. The coloured curve is one candidate month. The statistic is the average vertical gap between them, so the smallest score wins.
A candidate that hugs the grey curve has small gaps and a low score. A month that was unusually warm or unusually cold pulls its curve sideways, the gaps open up, and the score rises. Only the lowest score is kept, which is why a TRY contains no extremes.
How a Design Summer Year (DSY) Is Selected
A DSY is different. It is a single continuous real year that contains significant periods of hot weather. We need these files because overheating risk depends on actual sequences of warm days, not on average conditions.
CIBSE produces three versions of DSY:
- DSY1 is a moderate hot year. It roughly matches a 7-year return period heat event. This is the file we normally use first for standard overheating assessments.
- DSY2 contains the most intense short heat events. We use it when we want to stress-test the building against sharp, severe heatwaves.
- DSY3 contains the longest continuous warm periods. We use it when we want to check how the building handles prolonged heat.
The selection process works like this:
- They take the same 30-year observational baseline.
- For every year they calculate heat metrics over the April to September period (usually forms of Weighted Cooling Degree Hours).
- They rank the years according to overall heat severity, intensity and duration.
- They choose:
- DSY1 as the year closest to a 7-year return period,
- DSY2 as the year with the highest intensity heat events,
- DSY3 as the year with the longest-duration heat events.
- They keep the entire continuous 8760-hour year.
Because a DSY is a real year, the hour-to-hour sequence of weather remains completely realistic. That realism is vital for accurate overheating analysis.
What a heat metric counts
A weighted cooling degree hour metric works on the hourly record. For each hour we take the amount by which the temperature sits above a base, raise it to a power, and add the result up across April to September:
The exponent is the whole argument. At every degree above the base counts the same, so a long mild summer can outscore a short fierce one. Raise and the hottest hours start to dominate the total. That single choice is what turns one record into both a severity ranking and an intensity ranking.
How a heat metric adds a summer up
Every hour above the base contributes, and the weighting decides how much a very hot hour counts for. Sweep the season and watch the running total build. The record here is synthetic, not CIBSE data.
At an exponent of 1 every degree above the base counts the same, so a long mild summer can outscore a short fierce one. Raise the exponent and the hottest hours start to dominate the total. That single choice is what separates a severity ranking from an intensity ranking.
Why the moderate year is not the hottest year
DSY1 carries a 7 year return period, and that is deliberately not the top of the ranking. If we rank summers from hottest down, the summer in position has a return period of about
For a 30 year record and a 7 year return period that gives . So DSY1 lands around the fourth warmest summer in the record, not the first. The three hotter summers are rarer than a designer should have to treat as routine. CIBSE works on the return period of the heat events themselves rather than on whole summer ranks, but the arithmetic has the same shape.
Change the metric below and the winner changes with it. That is the whole reason there are three DSY files and not one.
One record, three Design Summer Years
The same 30 synthetic summers, ranked three different ways. Switch the metric and the winner changes, which is the whole reason CIBSE publishes DSY1, DSY2 and DSY3 rather than one hot year.
Rank the summers by total weighted heat, then step down to the year with a one in seven likelihood. That is deliberately not the hottest year on record.
Turning Baseline Files into Future Climate Files
Both TRY and DSY files start from observed weather. To create future versions, CIBSE uses a process called morphing.
Morphing works in clear steps:
- Start with a high-quality observational baseline year (either a TRY or a DSY).
- Extract climate change signals (also called change factors) from the UKCP18 probabilistic projections.
- Apply a bounded weighted stretch morphing algorithm. This algorithm adjusts the means, the variances and the extremes of the weather variables while still keeping realistic hour-to-hour sequences and the natural relationships between temperature, solar radiation, humidity and wind.
- The final file therefore carries the statistical signature of a future climate but still feels like real weather.
The 2025 release also benefits from improved solar radiation data taken from satellite-based CAMS and ERA5 sources. This gives us more accurate solar gains in our simulations.
What shift and stretch do to a number
Morphing algorithms all descend from the same two operations. We shift the month’s mean, then stretch the departures from that mean:
Here is the baseline hour, is the change in that month’s mean, is the baseline monthly mean, and is the stretch factor. CIBSE’s 2025 files use the revised bounded algorithm of Eames and colleagues, which limits how far the stretch can run, but this is the form it grows from.
Take a July with a baseline mean of 17.0 °C. Apply a shift of 2.0 K and a stretch of 0.15. A warm hour at 26.0 °C becomes
and a cool hour at 12.0 °C becomes
The mean rose by 2.0 K. The warm hour rose by 3.35 K and the cool hour by only 1.25 K. The July range widened from 14.0 K to 16.1 K, which is the original range multiplied by 1.15. That is the stretch factor at work, and it is why a morphed file gives far more overheating hours than a uniform warming of the same size would.
Morphing: shift the mean, stretch the spread
The baseline year keeps its hour-to-hour shape. The change factors move its mean and widen its swings, so the file gains a future climate's statistics without becoming an invented sequence of weather. The change factors here are illustrative, not CIBSE or UKCP18 values.
Set the stretch to zero and the whole year rises by the same amount, so the peak gains exactly what the mean gains. Add stretch and the hot hours climb faster than the mean does. The hours above 26 °C then rise far more steeply than the shift alone would suggest, which is why the stretch term matters so much for overheating work.
What the new baseline does to the numbers
The change between releases is not small. CIBSE ran a semi-detached house through both sets. Heating demand fell everywhere, because the 1994 to 2023 baseline is warmer than the 1984 to 2013 one it replaces. The largest fall was between the old Cardiff file and the new Zone 5, at 1,335 kWh or 53 per cent. The smallest was between Edinburgh and Zone 15, at 628 kWh or 20 per cent. Every other comparison landed between 20 and 38 per cent.
Time Periods, Emission Scenarios and Percentiles
When we choose a future weather file we also decide three things: the time period, the emission scenario and the probability percentile.
Time periods available
- 2030s (covering 2019 to 2039). Many of us now treat this as the new “current” climate.
- 2050s (2039 to 2059)
- 2080s (2069 to 2089)
Emission scenarios (RCPs)
RCPs stand for Representative Concentration Pathways. The number tells us the approximate extra energy trapped in the climate system by the year 2100, measured in watts per square metre relative to pre-industrial levels.
- RCP 2.6 (labelled Low) assumes strong global mitigation. Emissions peak early and fall rapidly.
- RCP 4.5 (labelled Medium) assumes moderate mitigation. Emissions peak around mid-century and then decline.
- RCP 8.5 (labelled High) assumes continued high emissions through most of the century.
CIBSE mainly provides:
- Low (RCP 2.6) mostly for the 2080s
- Medium (RCP 4.5) for the 2050s and 2080s
- High (RCP 8.5) for the 2030s, 2050s and 2080s
Probability percentiles
Because UKCP18 is probabilistic, each projection comes with a range of possible outcomes. CIBSE gives us three percentiles:
- 50th percentile: the central or median estimate. This is the value we use for typical assessments.
- 90th percentile: a plausible upper bound. Only about 10 % of the climate models project more extreme conditions than this. We use it for stress-testing and resilience studies.
- 10th percentile: the lower bound of the projection range.
It is important to remember that these percentiles describe uncertainty in the climate models under a given emissions pathway. They do not tell us the probability that a particular weather year will actually occur.
What is actually published
Not every combination exists. CIBSE generated files for all four UKCP18 pathways, then released only the ones that were different enough from each other to be worth having. The earlier the period, the fewer the choices.
| Time period | Emission scenarios | Percentiles |
|---|---|---|
| 2030s (2019 to 2039) | High (RCP 8.5) | 50th |
| 2050s (2039 to 2059) | Medium (RCP 4.5), High (RCP 8.5) | 10th, 50th, 90th |
| 2080s (2069 to 2089) | Low (RCP 2.6), Medium (RCP 4.5), High (RCP 8.5) | 10th, 50th, 90th |
This shapes the advice below. There is no 90th percentile file for the 2030s, so a stress test has to move to a later period to find one.
Scenario, period, percentile: three separate dials
Each published combination gets its own band. Nothing joins the periods, because the 90th percentile of one period and the 90th of the next are separate points in a spread of model outputs, not one climate run followed through the century. The warming values here are illustrative, not UKCP18 outputs.
Count the bands. The 2030s has one, and it is a median with no spread at all, because only the 50th percentile is published for that period. The 2050s has two, the 2080s has three. The choice widens as the pathways pull apart, and before mid-century they are too alike to be worth publishing separately.
Practical Advice on Choosing the Right File
Here is a simple guide we can follow:
- For typical annual energy calculations, start with a TRY for the 2030s or 2050s at the 50th percentile.
- For a standard overheating check, use DSY1 for the 2030s or 2050s under the High scenario at the 50th percentile.
- For an intensity stress test, move to DSY2 and consider a higher percentile or a later time period.
- For a duration stress test, use DSY3 with a higher percentile or later period.
- For long-term resilience studies, look at DSY1, DSY2 or DSY3 for the 2080s under the High scenario at the 90th percentile.
Understanding the File Names
A typical 2025 file name looks like this:
Z1_DSY1_2030s_HIGH50_CIBSE_v1.1.epw
We can read it as:
- Z1 = climate zone
- DSY1 = file type
- 2030s = time period
- HIGH50 = High emissions scenario plus 50th percentile
- CIBSE_v1.1 = provider and version
Once you recognise the pattern, selecting the correct file becomes straightforward.
CIBSE also gives a longer reference form for quoting the file in a report, which spells the scenario and percentile out in full:
Z1_DSY1_2030s_HIGH_50th_CIBSE_2025_V1.1
Put that in the modelling parameters section of the report. A reviewer can then see exactly which file produced the results, without opening anything.
The chooser below works the other way round. Start from the question you are trying to answer, and it assembles the file name for you.
From the question to the file name
Pick what you are trying to find out and the file name assembles itself. Options CIBSE does not publish for a period are shown dashed and cannot be chosen.
A real year holding a moderate heat event with a one in seven year likelihood. This is the normal starting point for TM59, TM52 and Part O work.
Z1_DSY1_2030s_HIGH50_CIBSE_v1.1.epwZ1_DSY1_2030s_HIGH_50th_CIBSE_2025_V1.1Find the zone number for a real project with the CIBSE Weather Data Selection Tool, which takes a postcode or a latitude and longitude. The zone slider here just shows how the first field changes.
Quick Summary of the Key Ideas
- A TRY is a composite typical year. We use it for energy analysis.
- DSY1, DSY2 and DSY3 are continuous real hot years. We use them for overheating studies (moderate, intense and long-duration heat respectively).
- Morphing applies UKCP18 change factors to the baseline years so we can study future climates.
- RCPs define the emissions pathway (2.6, 4.5 or 8.5).
- Percentiles show the range of climate model uncertainty (50th is central, 90th is upper bound).
- The 2025 release improves on the 2016 set with a newer observational baseline, 28 climate zones, UKCP18 projections and better solar data.
When we understand these building blocks, we can choose weather files that match the exact question we are asking about a building. That leads to more reliable energy predictions and more robust overheating assessments.
Sources
- CIBSE Weather Data, including the Weather Data Selection Tool.
- Technical Briefing: CIBSE Weather Data, 2025 Release, v1.1.
- Eames ME, Xie H, Mylona A, Shilston R, Hacker J. A revised morphing algorithm for creating future weather for building performance evaluation. Building Services Engineering Research and Technology, 2023.
- Xie H, Eames M, Mylona A, Davies H, Challenor P. Creating granular climate zones for future-proof building design in the UK. Applied Energy, 2024.
- CIBSE TM59 (2017), Design methodology for the assessment of overheating risk in homes.