Why is mathematics important when a city feels hotter in one street than another? A thermometer gives a value only at its own place and time. A useful temperature map must connect many scattered observations, account for distance and landscape, represent uncertainty and avoid implying detail that was never measured. Spatial mathematics makes that possible.
Urban heat is also a modelling lesson in humility. A smooth colour surface is not a photograph of heat. It is an estimate built from sensors, assumptions and a chosen interpolation method. This guide explains coordinates, weighted averages, gradients, validation and careful interpretation, with Singapore examples from official sources.
- Separate heat concepts
- Build a spatial dataset
- Interpolate between stations
- Quantify uncertainty
- Work through examples
- Learn through investigations
Air temperature, surface temperature and heat stress
Air temperature describes the air measured under specified conditions. Land-surface temperature may be estimated from thermal remote sensing and can differ sharply from the air a person experiences. Heat stress depends on more than air temperature: humidity, radiant heat, wind, clothing, activity and exposure matter.
Singapore’s Meteorological Service heat-stress page states that Wet-Bulb Globe Temperature, or WBGT, is a composite measure involving air temperature, humidity, wind and solar radiation, and is different from air temperature. Its public readings are averaged over a recent interval. A map must label the variable rather than simply say “heat”.
Urban heat island is a comparison
An urban heat island refers to an urban area being warmer than a less urban reference under comparable conditions. The contrast can vary by hour, weather, season, location and the reference selected. It is not one permanent number attached to a city.
The Urban Redevelopment Authority explains that built environments can absorb and retain heat and describes planning strategies involving airflow, materials, greenery, spacing and building form. These are plausible physical mechanisms and design responses; a map is needed to examine where, when and by how much conditions differ.
Did You Know? A colour can hide several decisions
Before one map cell becomes red, someone has chosen a variable, time window, units, colour scale, interpolation method and classification boundary. Changing any one can change the picture without changing the underlying measurements.
That does not make maps deceptive by nature. It makes map literacy mathematical. Readers should inspect the legend, source, spatial resolution, observation time and uncertainty before drawing conclusions.
Every reading needs location, time and context
A temperature record might contain latitude, longitude, elevation, timestamp, measured value, instrument identity and quality flags. Urban features near the sensor—shade, vegetation, road surface, walls and ventilation—can affect representativeness. Two sensors at different heights or shielding conditions are not automatically comparable.
Coordinates locate readings in space. For neighbourhood-scale distance calculations, a suitable projected coordinate system may be more convenient than treating latitude and longitude as flat x-y units. Over large areas, Earth’s curvature matters. The choice of coordinate system belongs in the method.
Sampling design shapes the map
If most sensors cluster in parks and few cover dense commercial areas, the map knows parks better. If stations follow convenient rooftops, street-level exposure may be underrepresented. Interpolation cannot invent the missing diversity.
A good sampling plan considers the phenomenon’s spatial scale, important land-cover types, prevailing winds, elevation and practical maintenance. Random, systematic and stratified designs answer different needs. In an urban study, stratifying by local climate zone or surface type may improve coverage of meaningful contrasts.
Time alignment matters
Combining one station’s 2 p.m. reading with another station’s 4 p.m. reading can turn temporal change into apparent spatial difference. Observations should be synchronised or adjusted through a justified model.
The issue becomes stronger during clouds, rain or sea-breeze transitions. A map labelled with a date but not a time window may encourage a false static interpretation. Heat mapping is spatiotemporal mathematics, even when the final graphic is a single frame.
Quality control before interpolation
Check impossible values, stuck sensors, abrupt jumps, missing timestamps, coordinate errors and duplicated records. Compare nearby stations and instrument metadata. A suspicious outlier can be a real microclimate or a faulty sensor; automatic deletion is not a neutral act.
Keep raw data, quality flags and corrected data separately. Reproducible analysis should show which values were excluded and why. The colourful map comes after data discipline.
Spatial interpolation estimates unmeasured locations
Interpolation uses values at observed locations to estimate values between them. The simplest methods assume nearby places are more similar than distant places. That tendency is common in environmental data, but it is not a law: a coastline, hill, park or building canyon can create sharp boundaries.
We write observations as zi at locations si. The estimate at target s0 is often a weighted sum: z-hat(s0)=Σwi zi, with weights summing to one. Different methods define weights differently and therefore express different assumptions.
Nearest neighbour
Assign the nearest station’s value to each target. The result has abrupt region boundaries. It preserves observed values and is easy to explain, but ignores all other stations and creates blocky transitions.
Nearest neighbour can be appropriate for categorical labels or operational zones. For a smoothly varying field, it is usually a baseline rather than a final model.
Inverse distance weighting
Inverse distance weighting, or IDW, gives closer observations larger weights. A common form is wi = di^−p / Σdj^−p, where di is distance from target to station i and p is a positive power. When p increases, the nearest stations dominate more strongly.
IDW is deterministic and intuitive. It does not learn a spatial correlation model, and it can create circular “bullseyes” around isolated stations. Barriers and directional winds are absent unless added deliberately.
Triangulated interpolation
Stations can form a triangulation. Within each triangle, a plane fitted through the three vertex values gives a linear estimate. Barycentric coordinates provide the weights. The surface is continuous across shared edges when constructed consistently, though its gradient can change abruptly.
This creates a direct connection to Triangle Rasterisation, Barycentric Coordinates and Pixel Coverage once that article is published. In both cases, a point inside a triangle receives weights tied to geometry.
Kriging
Kriging uses a model of spatial dependence, commonly described through covariance or a variogram, to form a best linear unbiased estimate under stated assumptions. It can produce prediction uncertainty as well as a surface.
Kriging is not automatically superior. The variogram needs enough informative data; trend, anisotropy and non-stationarity must be considered. A sophisticated model poorly matched to the process can be less trustworthy than a transparent baseline.
Regression and hybrid models
Temperature may relate to elevation, vegetation, impervious surface, distance to coast, building density or sky-view factor. A regression can model these covariates, and spatial interpolation can model residuals. This is sometimes called regression kriging or a related hybrid approach.
Covariates can sharpen estimates where physical relationships are stable, but they also introduce measurement error and causal temptation. A predictive association between greenery and temperature does not by itself prove the temperature change produced by a specific intervention.
Distance is not the only kind of closeness
Two locations 300 metres apart across an open field may exchange air more freely than two equally close points separated by a dense block. A station upwind may be more informative than an equally distant one downwind. Elevation and shade can matter.
Anisotropy means spatial dependence differs by direction. A variogram may rise faster north-south than east-west, or a distance metric may be stretched along prevailing flow. Barriers can prevent weighting across water or terrain in some applications.
Scale changes relationships
At a 10-metre scale, tree shade and facade orientation can dominate. At a kilometre scale, land-cover pattern and coastal influence may be more visible. Aggregating cells can smooth extremes. A model calibrated for one resolution should not be casually interpreted at another.
This is the modifiable areal unit problem in a broader family of scale effects: patterns can change when spatial units or boundaries change. A district average cannot identify the hottest corner, and a single hot sensor does not define an entire district.
Singapore’s Ministry of Sustainability and the Environment stated in a 9 September 2026 parliamentary reply that observed temperature patterns did not follow residential district boundaries and that associating differences with heat inequality by those districts may not be meaningful. That is a timely example of why administrative boundaries should not be mistaken for physical heat regions.
Worked example: inverse distance weighting
Three stations record 31.0°C, 32.0°C and 33.5°C. A target lies 1 km, 2 km and 4 km away. With power p=2, unnormalised weights are 1, 1/4 and 1/16. Their sum is 1.3125.
Normalised weights are about 0.7619, 0.1905 and 0.0476. The estimate is 0.7619(31.0)+0.1905(32.0)+0.0476(33.5) ≈ 31.29°C. The nearest station dominates because squared inverse distance decreases quickly.
Check the result. Weights sum to one, so the estimate is a weighted average inside the observed range. Units remain degrees Celsius because dimensionless weights multiply temperatures.
Compare power choices
With p=1, unnormalised weights are 1, 1/2 and 1/4, giving normalised weights about 0.5714, 0.2857 and 0.1429. The estimate is about 31.64°C. The farther hot station has more influence than under p=2.
Neither power is universally correct. Cross-validation, process knowledge and sensitivity analysis can guide the choice. Publishing only one smooth surface without showing sensitivity can hide modelling uncertainty.
Exact station location
If the target coincides with a station, di=0 makes di^−p undefined. Implementations typically return that station value directly or handle collocated readings explicitly. This boundary case should be designed, not left to accidental division by zero.
Worked example: a temperature gradient
Suppose temperature along a 500-m east-west path rises from 30.8°C to 32.3°C. A simple average gradient is (32.3−30.8)/500 = 0.003°C per metre, or 0.3°C per 100 metres.
This slope summarises net change; it does not claim every intermediate step is linear. Shade and materials can produce local rises and falls. A profile plot of many points can reveal those variations.
If measurement uncertainty is ±0.2°C at each endpoint, the 1.5°C difference should not be reported with false precision. Error propagation and repeated observations help judge whether the spatial contrast is stable.
Gradients are vectors on a surface
On a two-dimensional temperature surface T(x,y), the gradient combines change in x and y. It points in the direction of steepest local increase, and its magnitude describes rate of change per unit distance under the model.
Urban planners may care about gradients because abrupt transitions can identify boundaries between materials or land covers. Yet a gradient computed from an interpolated surface inherits the interpolation assumptions and can amplify noise.
Worked example: comparing periods fairly
A neighbourhood has a mean night temperature of 28.4°C during one five-night campaign and 28.9°C during another. The 0.5°C difference is not automatically an effect of a new pavement treatment. Weather patterns, cloud, wind, rainfall, sensor placement and seasonal timing may differ.
A stronger comparison uses matched control locations, before-and-after measurements, repeated nights and a design that separates treatment from general change. A difference-in-differences calculation compares the treated site’s change with the control site’s change.
For example, treated rises 0.5°C while control rises 0.4°C; the relative change is 0.1°C. That still needs uncertainty analysis and design scrutiny. Mathematics helps resist a dramatic but unsupported causal story.
Validation asks how well the map predicts
Leave-one-out cross-validation removes one station, predicts its value from the others and records the error. Repeat for every station. Errors can be summarised with mean error, mean absolute error and root mean square error.
Mean error reveals systematic bias but can hide cancellation. Mean absolute error describes typical magnitude in original units. Root mean square error penalises larger errors more heavily. Report more than one when practical.
Spatial cross-validation is stricter
Randomly splitting dense neighbouring observations can overstate performance because a test point may sit beside a training point. Blocking data into spatial groups tests prediction across gaps more honestly.
The validation design should resemble the intended use. Predicting within a well-instrumented district is different from extrapolating to an unsampled coastal or industrial area.
Residual maps reveal structure
A single error score cannot show where a model fails. Map residuals—observed minus predicted—and look for clusters. Consistently positive residuals near a land-cover type may signal a missing covariate or non-stationary relationship.
Residual structure means the model has not captured all predictable spatial pattern. It can guide improvement, but repeatedly tuning to the same validation data risks overfitting. Reserve an independent test when decisions are important.
Prediction is not measurement
At station locations, a model may reproduce values exactly or smooth them, depending on method and measurement-error assumptions. Between stations, map values are estimates. Outside the convex hull of observations, extrapolation is often less reliable.
An uncertainty layer, station overlay and mask for unsupported areas can communicate this. A polished uninterrupted surface should not imply equal confidence everywhere.
Colour scales and classification
A sequential colour scale suits low-to-high temperature. A diverging scale may suit anomaly around a meaningful zero. Rainbow scales can introduce false visual boundaries and unequal perceptual steps.
Choose limits honestly. If each daily map rescales to its own minimum and maximum, small differences may look as dramatic as extreme ones, making days hard to compare. Fixed limits improve comparability; adaptive limits can reveal local structure. State the choice.
Bins change the story
Classifying 30.0–30.9°C as one colour and 31.0–31.9°C as another creates an apparent boundary between 30.99 and 31.00. Equal intervals, quantiles and policy thresholds answer different questions.
Show continuous legends when appropriate and do not let visual classes substitute for uncertainty. A threshold map may be useful operationally, but the source estimate near the boundary may be uncertain.
Maps need accessible design
Use readable legends, units, timestamps and descriptions that do not depend solely on colour. Test for common colour-vision differences. Provide data tables or textual summaries for key locations where possible.
Good visualisation is part of mathematical communication. The reader should be able to tell what was observed, what was estimated and what comparison the colours support.
Common misconceptions
“The red zone was measured at every red pixel”
Usually not. It may be an interpolated region supported by a limited set of stations or remotely sensed pixels. Overlay observations and read the method.
“The hottest surface is automatically the greatest human heat stress”
Surface temperature, air temperature and WBGT differ. Shade, humidity, wind and activity affect exposure.
“A smoother map is a more accurate map”
Smoothness can come from a stronger modelling assumption. Accuracy must be tested against held-out observations.
“Nearby always means similar”
Barriers, elevation, coastline, vegetation and building form can interrupt distance-based similarity.
“Correlation proves a cooling intervention worked”
Cooler places may differ in many ways. Causal evaluation needs design, controls and attention to confounding.
“A district average describes every resident”
Aggregation hides microclimates and time spent in different environments. Exposure is not identical to residential average.
A practical learning path
Stage 1: map points, not a surface
Collect or use a small ethical dataset with coordinates, times and temperatures. Plot coloured station points. Notice gaps before estimating them.
Stage 2: compute distances
Use a scale drawing or projected coordinates. Calculate distances from one target to several stations. Check units and identify the nearest observation.
Stage 3: build an IDW estimate
Try powers 1 and 2. Show raw weights, normalised weights and weighted temperature. Verify that weights sum to one.
Stage 4: make a grid
Estimate values at regularly spaced targets. Draw contours or cells with one fixed legend. Keep station markers visible.
Stage 5: validate
Hide each station in turn and predict it. Calculate mean error, MAE and RMSE. Map the residuals.
Stage 6: add context carefully
Compare residuals with greenery or surface categories. Phrase the result as association unless the design supports causality. Identify alternative explanations.
Stage 7: communicate uncertainty
Make a second layer showing distance to nearest sensor, cross-validation error or kriging standard error. Write a caption naming unsupported areas.
Investigation portfolio
Investigation 1: sensor density
Create a map with four corner stations, then add a central station with an unexpectedly high value. Compare surfaces. Explain how one new observation can change local structure.
Investigation 2: time mismatch
Give half the stations noon readings and half 3 p.m. readings. Make the misleading map, then align times. Describe which apparent spatial pattern was really temporal.
Investigation 3: choice of p
For one target, calculate IDW with p=0.5, 1, 2 and 4. Plot estimate against p. Explain why larger p approaches nearest-neighbour behaviour.
Investigation 4: leave-one-out error
Remove the hottest station and predict it. If the model strongly underestimates, ask whether it is an error, a local hotspot or an unsupported outlier.
Investigation 5: colour-scale ethics
Draw the same values with a narrow and wide legend range. Ask classmates which looks more alarming. Decide which comparison each scale supports.
Investigation 6: aggregation
Average a 4×4 grid into four large cells. Compare maximum, range and gradient before and after. Identify what disappeared.
Investigation 7: barriers
Place stations on both sides of an imaginary park or coastline. Compare ordinary distance with a barrier-aware rule. Discuss whether the barrier is physically justified.
Investigation 8: covariate model
Fit a simple relation between temperature and percentage vegetation in a toy dataset. Plot residuals. Explain why prediction does not prove the effect of planting a specific number of trees.
Guidance for students
Begin with the question, not the software. Are you mapping air temperature at one time, a daily maximum, a surface temperature or a heat-stress index? Write the quantity and unit at the top.
Keep observations visible. A station-point map prevents the smooth surface from becoming mentally detached from evidence. Mark missing and excluded values.
Compare methods rather than searching for one magic algorithm. Nearest neighbour, IDW and a triangle plane make different assumptions. Validation shows consequences.
Use honest language: “estimated”, “associated”, “under these conditions” and “within the sampled area”. These phrases make a result stronger because they match the evidence.
Guidance for parents and teachers
A safe classroom activity can use published data or simulated values rather than having students work in dangerous midday heat. The mathematics does not require exposure.
Ask students to critique a map legend and sampling pattern before calculating. Spatial judgement begins with the evidence layout.
Connect mapping to familiar weighted averages, coordinates and graphs. Then add complexity only when it answers a real limitation. A transparent IDW map with validation teaches more than an unexplained advanced model.
Discuss civic relevance without turning a school exercise into a policy verdict. Official planning involves richer data, physical modelling, design constraints and community needs.
Frequently asked questions
What is an urban heat island?
It is a temperature contrast in which urban conditions are warmer than a suitable less urban reference, varying with place, time and weather.
Is IDW machine learning?
It is usually described as a deterministic interpolation method. Broader labels matter less than understanding its distance-based weights and limitations.
Does kriging always give the best map?
No. It depends on data, variogram modelling and assumptions. Compare validated performance and interpretability.
Why can WBGT differ from air temperature?
WBGT incorporates humidity, wind and radiant conditions as well as temperature-related measurements, so it addresses heat stress rather than air temperature alone.
How many sensors are enough?
There is no universal number. It depends on area, spatial variability, desired resolution, design and acceptable uncertainty.
Can a map identify the cause of a hotspot?
It can suggest hypotheses. Establishing cause usually needs additional variables, physical reasoning and a study design that addresses confounding.
Useful next reading
A miniature heat-map studio
Create a fictional 1-km square neighbourhood with sensors at (0,0), (0,1), (1,0), (1,1) and (0.5,0.5), where coordinates are kilometres. Assign simultaneous temperatures 30.4, 31.0, 31.6, 32.0 and 30.8°C. The centre value deliberately breaks a simple southwest-to-northeast gradient.
Estimate temperature at (0.5,0.25) by nearest neighbour and by IDW with p=1 and p=2. Because the target is close to the central sensor, larger p pulls the estimate toward 30.8°C. Record each estimate, then ask which observations dominate and why. The method comparison is more instructive than choosing a winner immediately.
Remove the centre station and repeat. The corner-only surface will look smoother and may miss the local cool feature. Put the station back but hide its value, predict it from the corners and calculate the residual. That residual quantifies what the sparse network failed to capture at one location.
Next change only the northeast reading from 32.0 to 33.0°C. Observe how far its influence spreads under each power. A low p distributes influence broadly; a high p localises it. This is sensitivity analysis: alter one input or assumption and examine the output.
Finally, make two legends—one from 30 to 34°C and one from 30.5 to 32°C. The same estimates will look calmer in the wide scale and more dramatic in the narrow scale. Neither display is automatically dishonest, but the first supports comparison with hotter conditions while the second reveals local contrast. Captions should state the range and purpose.
Uncertainty through repeated measurement
Suppose every station records five consecutive 15-minute averages. Compute each station’s mean and range. A station with values 31.0, 31.1, 31.0, 31.2 and 31.1 is temporally stable over that window; one with 30.5, 31.2, 32.0, 31.4 and 30.8 is not. Mapping only the mean hides that difference.
Create a second map of temporal standard deviation or range. A highly variable location may need cautious comparison, especially if stations were not perfectly synchronised. The map of uncertainty does not cancel the map of estimates; it tells readers how much weight to give local contrasts.
Measurement uncertainty and interpolation uncertainty are distinct. A precise sensor can still sit in a large unsampled gap, making the spatial estimate uncertain. Conversely, a dense network of poorly calibrated instruments can produce detailed but biased patterns. A sound study addresses both.
From association to a testable urban question
Imagine cooler estimates appear near parks. A responsible next question is not “How many degrees do parks always cool?” It is “Under comparable times and weather, how does temperature vary with distance from selected green spaces, after accounting for other measured features?” That question names a population, comparison and covariates.
A design might pair park-edge and built-up locations, record them simultaneously over many days, and model weather conditions. Even then, park type, canopy, irrigation, wind and surrounding buildings can differ. The aim is not to make causal study impossible; it is to prevent an attractive map from skipping the work.
A data-ethics checkpoint
Environmental sensors usually seem impersonal, but fine-scale location data can reveal activity around homes, schools or workplaces. Projects should minimise personal data, obtain permissions for private sites and publish spatial detail proportionate to purpose. A precise coordinate is not always necessary for a public classroom map.
Equity claims need equal care. Temperature exposure, housing, health, outdoor work and access to cooling interact. A district-colour map alone cannot represent every resident’s experience. Combine quantitative evidence with appropriate community and policy expertise.
What a trustworthy caption includes
A concise caption can carry surprising analytical value: variable and unit; date and time window; number and type of observations; interpolation method and parameters; spatial resolution; validation summary; and a warning about extrapolation. For example, “Air temperature estimates, 2:00–2:15 p.m.; IDW p=2 from 18 quality-checked stations; cells beyond 1 km from a station hatched; leave-one-out MAE 0.4°C.”
That sentence lets a reader distinguish evidence from appearance. Writing it also forces the analyst to confront gaps before publication.
Comparing interpolation methods fairly
Use identical training observations, targets and validation folds for every method. Tune IDW power only within the training portion; otherwise the held-out data quietly influence the model and the reported error becomes optimistic. For kriging, estimate trend and variogram without peeking at test values.
Create a score table with bias, MAE, RMSE, maximum absolute error and computation time. Add a qualitative column for interpretability and a map of residuals. The method with lowest RMSE may have a serious local bias or require assumptions that are hard to defend.
Compare a simple baseline. Predicting every location with the network mean will be poor spatially, but it establishes whether interpolation adds real skill. Nearest neighbour is another baseline. An advanced model that barely improves either may not justify its complexity.
Cross-sections make surfaces easier to inspect
Draw a line across the map and plot predicted temperature against distance for each method. Nearest neighbour produces steps. IDW produces curves influenced by station positions. Triangular linear interpolation produces straight segments within triangles. A kriging profile follows its covariance and trend model.
Profiles expose oscillations or bullseyes hidden by colour. Add station values near the line and an uncertainty band. The one-dimensional graph complements the two-dimensional map and invites familiar slope reasoning.
Resolution is not accuracy
A model can output a value every metre even when stations are kilometres apart. The grid spacing is computational resolution, not guaranteed spatial detail. Upsampling creates more cells, not more observations.
Choose a display resolution consistent with intended use and evidence. If the surface varies little at the sensor spacing, a fine grid may be harmless for visual smoothness, but its caption should not imply metre-scale validation. Where decisions operate by neighbourhood, a coarser honest summary can be more useful.
Heat maps across a day
Repeat mapping hourly and treat each cell as a time series. Calculate daily maximum, time of maximum, night-time minimum and hours above a defined threshold. Two locations can share the same daily mean while one has a sharper afternoon peak and the other stays warm overnight.
Animations reveal movement but can overwhelm comparison if legends change. Small multiples with a fixed scale often support careful reading. Temporal interpolation introduces another assumption: a smooth line between hourly observations may miss short cloud or rain events.
Spatiotemporal validation can hold out whole stations and whole time blocks. This tests whether the model transfers to both new places and new periods. Random individual readings are a much easier test because neighbouring times and locations leak information.
A final reader checklist
Before using a heat map, ask six questions. What variable is mapped? When was it observed? Where are the actual sensors? How were gaps estimated? How accurate were held-out predictions? Which regions have weak support? If the map cannot answer them, treat its colours as an invitation to investigate rather than a precise ranking.
For student reports, add one sentence on what would most improve the study. It might be another station in a large gap, simultaneous measurements, better shade metadata or repeated nights. Naming the next useful evidence keeps uncertainty constructive. Mathematics does not merely criticise a map; it helps design the next measurement.
Read Air Quality, PSI and Particulate Matter for another environmental example where indices, sensors and communication matter. Explore Rainwater Harvesting, Runoff Coefficients and Tank Sizing for spatially variable weather and design assumptions. For networks rather than surfaces, see Search Engines, PageRank and Network Centrality.
Urban heat mapping shows why mathematics matters in daily life and public decisions. It connects a finite set of measurements to a continuous-looking city while keeping the limits visible. The best map is not simply the prettiest one. It is the one whose variable, weights, scale, validation and uncertainty are clear enough for a reader to make a proportionate judgement.
