VIEW THIS AS

Auto mode follows the Route Engine until you choose a viewpoint.

YOU ARE HERE

ROUTE CHECK

CONNECTED TO

WHAT NEXT

Use the canonical route for this room, or HELP if you are unsure.

Why Mathematics? | Computer Vision, Convolution and Edge Detection

Three learners review open books together at a classroom table, with stacks of textbooks, stationery and a whiteboard in the bright room.

Why is mathematics important when a computer looks at a photograph? The machine does not begin with “cat”, “road”, “tumour” or “face”. It begins with an array of sampled values. Computer vision uses mathematics to turn those values into measurements of brightness, colour, change, shape, texture, motion and probability.

One of the clearest mechanisms is edge detection. An edge is a place where image intensity changes strongly across nearby pixels. A small matrix called a kernel can move across the image, multiply neighbouring values and add the results. Different kernels smooth noise, estimate horizontal and vertical derivatives, sharpen detail or highlight particular patterns. This local calculation is convolution or, in some software, a closely related correlation operation.

Modern computer vision includes much more than hand-designed filters. Convolutional neural networks learn many filters from data, while transformers and hybrid systems use other structures. Yet the foundations remain deeply mathematical: vectors, matrices, coordinates, derivatives, probability, optimisation and evaluation.

The worked examples here are deliberately small so a student can calculate them by hand. Real systems must also address data quality, privacy, bias, lighting, viewpoint, safety and human oversight. A high model score does not automatically make a deployment appropriate.


Choose the computer-vision question you want to answer

  • To understand the raw material, begin with pixels as numbers in an array.
  • To see how a filter works, multiply a neighbourhood by a kernel and sum.
  • To detect change, study finite differences and image gradients.
  • To trace edges, combine gradient strength, direction and thresholds.
  • To connect with AI, see how learned convolutional filters become feature maps.
  • To judge performance, use confusion matrices and task-appropriate metrics.
  • To understand limits, test lighting, scale, blur, occlusion and distribution shift.
  • To build skill, use small open images and document every transformation.

The route begins with a surprisingly simple idea: before a computer can interpret an image, it needs a numerical representation.


A greyscale image is a matrix

A small greyscale image can be represented as a matrix whose entries describe intensity. With an 8-bit convention, values commonly run from 0 to 255. Zero may represent black and 255 white, with intermediate values representing shades of grey.

Consider a 3×3 patch:

2020200
2020200
2020200

The left two columns are dark and the right column is bright. A human immediately sees a vertical boundary. The computer can detect it by calculating how quickly values change from left to right.

The matrix is not the object itself. It is a sampled measurement produced by a camera pipeline. Exposure, sensor noise, compression and colour conversion already shape the numbers.


Colour adds channels

A colour image often uses three channels such as red, green and blue. Each pixel is then a vector, for example (R,G,B)=(120,80,40). The whole image is a three-dimensional array: height × width × channels.

Different colour spaces rearrange the representation. One may separate brightness from colour components; another may use perceptual coordinates. The best choice depends on the task.

Channel order matters. If software expects RGB but receives BGR, colours change even though the array dimensions look correct. This is a broader lesson in mathematical computing: shape alone does not guarantee semantic compatibility.


Coordinates give every pixel an address

An image needs a coordinate convention. Rows commonly increase downward and columns increase to the right, while Cartesian graphs usually increase upward in y. Libraries may write coordinates as row-column or x-y.

An algorithm that swaps conventions can rotate a gradient direction, misplace a bounding box or mirror a transformation. The arithmetic may be correct while the mapping is wrong.

Good code documents image origin, axis direction, pixel-centre convention and units. If the image is tied to the physical world, calibration may convert pixels into millimetres, degrees or another meaningful measure.


A kernel is a local weighted rule

A kernel is a small matrix of weights. To apply it at a location, align it with a neighbourhood, multiply corresponding entries and add. A 3×3 averaging kernel is:

1/91/91/9
1/91/91/9
1/91/91/9

Applied to nine pixel values, it returns their mean. Moving the kernel produces a new filtered image.

This local rule is powerful because the same calculation is reused across the image. Translation of the kernel lets one mathematical pattern search many locations.


A worked smoothing example

Suppose a neighbourhood is:

101211
95010
111012

The centre value 50 is much larger than its neighbours. The mean is (10+12+11+9+50+10+11+10+12)/9=135/9=15. The averaging filter replaces the centre output with 15 at that kernel position.

Smoothing reduces isolated variation, but it also blurs genuine edges. The operation therefore trades noise suppression against localisation. A wider kernel usually smooths more strongly and removes more fine detail.

This is not a flaw unique to vision. Measurements often require a choice between sensitivity to detail and resistance to noise.


Weighted smoothing can preserve locality better

A Gaussian-like kernel gives more weight near the centre:

1/162/161/16
2/164/162/16
1/162/161/16

The weights sum to one, so a constant region remains unchanged. Nearby pixels contribute more than diagonal ones. This approximates smoothing by a Gaussian function.

OpenCV’s Canny tutorial begins with noise reduction because edge detection is sensitive to noise. The filter is not an optional decoration; it changes which intensity variations survive into the gradient calculation.

Students should check the weight sum. If smoothing weights sum to more or less than one, overall brightness can change.


Convolution and correlation differ by a flip

Mathematical convolution flips the kernel before sliding it. Many image-processing functions compute correlation, which slides the kernel without flipping, while still using “convolution” as a broad engineering term.

For symmetric smoothing kernels, flipping makes no difference. For directional derivative kernels, the sign or orientation can change.

This is a useful reminder to read documentation rather than trusting a familiar word. When reproducing results, state the library operation and kernel orientation. A method can be mathematically related yet not byte-for-byte identical.


Boundaries force a modelling decision

Near an image edge, part of the kernel falls outside the available array. Software must decide what values to use. Options include zero padding, reflecting the image, repeating the nearest border or returning output only where the kernel fully fits.

Each choice changes results. Zero padding can create an artificial dark boundary. Reflection may reduce that discontinuity but introduces a different assumption about the unseen region.

Padding also affects output size. For stride one, an input width W, kernel width K and padding P produce output width W−K+2P+1 under the simplest discrete formula. Keeping dimensions unchanged with an odd K often uses P=(K−1)/2.


Stride changes sampling density

Stride is how far the kernel moves between outputs. Stride one evaluates every adjacent position. Stride two skips alternate positions and produces a smaller feature map.

For input width 7, kernel width 3, no padding and stride 2, output width is floor((7−3)/2)+1=3. The floor appears because a final partial placement is not included.

Larger stride reduces computation and spatial resolution. It can also miss small details. Downsampling should consider aliasing: high-frequency variation can masquerade as lower-frequency patterns when sampled too coarsely.


Image gradients estimate change

In calculus, a derivative measures rate of change. A digital image is discrete, so derivatives are approximated by finite differences.

A simple horizontal difference [−1,0,1] compares values on either side. In two dimensions, Sobel kernels combine differencing with smoothing. OpenCV documents the horizontal Sobel kernel:

-101
-202
-101

Multiplying this kernel by a patch estimates change in the x direction. A transposed version estimates y-direction change.


A worked Sobel calculation

Use the earlier vertical boundary patch:

2020200
2020200
2020200

Multiply by the horizontal Sobel kernel and add:

Gx=(−1×20+0×20+1×200)+(−2×20+0×20+2×200)+(−1×20+0×20+1×200)=720.

The large positive result indicates a strong increase from left to right. If the bright and dark sides were reversed, the sign would be negative. Magnitude measures strength; sign helps indicate direction.


Gradient magnitude combines two directions

After estimating Gx and Gy, gradient magnitude can be G=√(Gx²+Gy²). A quicker approximation is |Gx|+|Gy|. Gradient direction is θ=atan2(Gy,Gx).

If Gx=30 and Gy=40, magnitude is 50 and direction is about 53.1 degrees from the positive x direction under the chosen coordinate convention.

The two-argument arctangent preserves the quadrant. Plain arctangent of Gy/Gx loses information when signs differ or Gx is zero.


Edge direction and gradient direction are perpendicular

The gradient points toward the fastest increase in intensity. The visible edge runs approximately perpendicular to that direction.

If intensity changes left to right, the gradient is horizontal and the boundary is vertical. This can feel backwards until the student separates the direction of change from the direction of the line.

The distinction matters in non-maximum suppression, which checks whether a gradient response is locally strongest along the gradient direction, not along the edge itself.


Canny edge detection is a sequence, not one threshold

The Canny method is commonly explained through several stages: smooth noise, calculate gradients, thin responses with non-maximum suppression, apply two thresholds and connect weak edges to strong ones through hysteresis.

OpenCV’s current tutorial describes Gaussian smoothing, Sobel derivatives, gradient magnitude and direction, non-maximum suppression, and hysteresis thresholding. The stages solve different problems.

A single threshold on raw gradients often produces thick, broken or noisy edges. Canny’s sequence shows how a useful algorithm can be built from simple mathematical operations whose roles are explicit.


Non-maximum suppression thins the edge

A real intensity transition can produce several neighbouring pixels with high gradient magnitude. Non-maximum suppression keeps a candidate only if it is locally strongest in the gradient direction.

Imagine magnitudes 20, 75 and 40 along the relevant direction. The centre 75 survives; its neighbours are suppressed. Repeating this rule produces a thinner contour.

Direction is usually quantised into a few bins for efficient neighbour selection. Quantisation introduces approximation, another reminder that digital algorithms balance accuracy and computation.


Two thresholds reduce fragile decisions

With hysteresis, values above a high threshold are strong edge candidates. Values below a low threshold are rejected. Values between them survive only when connected to strong edges according to the algorithm’s neighbourhood rule.

Suppose low=40 and high=100. A response of 130 is strong; 20 is removed; 70 is conditional. This helps preserve a faint continuation of a real edge without accepting every medium response as a separate edge.

Thresholds are hyperparameters, not truths. Their suitable values depend on image scale, noise and intensity range.


Threshold scale must match data scale

An image stored from 0 to 255 and an image normalised from 0 to 1 use different numerical scales. Reusing thresholds 40 and 100 on the normalised image would reject everything.

Similarly, gradients change with kernel choice, input bit depth and preprocessing. Threshold values are meaningful only with the full pipeline.

This is why reproducible computer vision reports configuration, not just an algorithm name. “Used Canny” is incomplete without preprocessing, thresholds, aperture and data conventions.


Sharpening is controlled contrast amplification

A sharpening kernel might increase the centre and subtract neighbours:

0-10
-15-1
0-10

In a constant region, the weights sum to one, so brightness is retained. Around a boundary, the subtraction emphasises local contrast.

Sharpening can make noise and compression artefacts more visible. Stronger-looking detail is not automatically more truthful detail. In scientific or medical contexts, transformations must be documented and validated for the intended interpretation.


The Laplacian responds to curvature

The Laplacian combines second derivatives. A discrete kernel such as centre 4 with four neighbouring −1 values responds where intensity bends away from a locally linear pattern.

Second derivatives can be very sensitive to noise, so smoothing is often paired with them. The Laplacian of Gaussian concept first smooths and then finds rapid changes in slope.

Students can compare first and second derivatives on a one-dimensional brightness profile. The first derivative peaks at a transition; the second derivative may change sign near its centre. This connects image processing to calculus.


Edges are not objects

An edge map marks local changes. It does not by itself know whether a contour belongs to a bicycle, a shadow, a leaf or text. Texture can create many edges; low-contrast objects can create few.

Object recognition requires grouping and context. Features may include corners, shapes, regions, motion or learned representations. A complete system may combine multiple stages.

This limit is educationally useful. It prevents the leap from “the algorithm highlighted boundaries” to “the computer understood the scene”.


Convolutional neural networks learn filters

In a convolutional neural network, kernels contain trainable weights. During learning, optimisation adjusts them to reduce a chosen loss on training data.

Early layers often respond to simple patterns such as oriented changes or textures, while deeper layers combine previous feature maps. That description is a useful intuition, not a guarantee that every channel has a simple human label.

The mathematical operation remains local weighted summation plus nonlinearity and other transformations. Learning changes where the weights come from: data and optimisation rather than manual design.


Channels mix through learned weights

A learned filter usually spans all input channels. For an RGB input, one output value can sum products from red, green and blue neighbourhoods, then add a bias.

If a layer has Cin input channels, Cout output channels and K×K kernels, its basic weight count is Cout×Cin×K×K, plus biases when used. With 16 input channels, 32 output channels and 3×3 kernels, that is 32×16×9=4,608 weights before biases.

Counting parameters helps students understand memory and model capacity. More parameters increase expressive power but can also raise computation and overfitting risk.


Nonlinear activation changes what layers can represent

Stacking purely linear convolutions without nonlinearities collapses into another linear transformation. Activation functions such as ReLU introduce nonlinearity, allowing networks to represent more complex relationships.

ReLU applies max(0,x). It keeps positive responses and sets negative ones to zero. This simple piecewise function changes the geometry of the network’s mapping.

The lesson is broader: complex models often grow from repeated combinations of linear algebra and simple nonlinear functions.


Pooling and downsampling trade detail for invariance

Max pooling keeps the largest value in a window; average pooling keeps the mean. Both reduce spatial resolution. A strong feature can remain visible even if it shifts slightly within the window.

That tolerance may help classification, but precise tasks such as segmentation need spatial detail. Modern architectures use multiple strategies, including strided convolution, skip connections and upsampling.

There is no universally best resolution. The correct representation depends on whether the task asks “what is present?” or “where exactly is every boundary?”


Classification, detection and segmentation are different jobs

Classification assigns a label to an image. Object detection predicts labels and bounding boxes. Segmentation assigns labels at pixel or region level.

Metrics must match the job. Classification accuracy says little about whether a box overlaps the object correctly. Pixel accuracy can hide poor performance on rare classes when background dominates.

Before choosing a model, define the output mathematically. Ambiguous tasks produce ambiguous evaluation.


A confusion matrix makes errors visible

For a binary detector, outcomes are true positive, false positive, true negative and false negative. Precision is TP/(TP+FP). Recall is TP/(TP+FN).

Suppose a model finds 80 true objects, misses 20 and raises 40 false alarms. Precision is 80/120=66.7 per cent. Recall is 80/100=80 per cent.

Accuracy alone may look high if most images contain no object. Precision and recall reveal different costs. The acceptable balance depends on the application and consequences.


Threshold changes precision and recall

A model often outputs a score rather than a final yes/no answer. Raising the decision threshold usually reduces positive predictions. False positives may fall, but false negatives may rise.

Plotting precision against recall across thresholds shows the trade-off. The threshold should be selected on appropriate validation data and linked to real costs, not chosen because 0.5 looks natural.

A safety-critical medical aid and a photo-sorting tool need different decision policies even if they share the same underlying model.


Intersection over union evaluates overlap

For bounding boxes or masks, intersection over union is IoU=area of overlap ÷ area of union. If a predicted box overlaps 60 square units with a reference box and their union is 100, IoU is 0.60.

A prediction can have the correct class but poor localisation. IoU makes location quality explicit.

The reference annotation is itself a human or procedural judgement. Ambiguous boundaries and annotator disagreement mean the “ground truth” may contain uncertainty.


Image augmentation encodes assumptions

Training data may be augmented by flips, crops, brightness changes or rotations. Each transformation asserts that the label should remain valid under that change.

A horizontal flip may be suitable for many animals but not for written text or traffic signs with direction. Rotating a medical scan arbitrarily may violate acquisition or anatomy conventions.

Augmentation is therefore mathematical and semantic. The transformation must preserve the task’s meaning, not merely produce more arrays.


Normalisation makes optimisation easier but changes scale

Pixel values may be scaled to 0–1 or standardised by channel means and standard deviations. This can improve numerical conditioning and match pretrained model expectations.

If training uses one normalisation and deployment another, predictions can fail even though the image looks ordinary to a person. The model receives different numbers.

Preprocessing is part of the model contract. It must be versioned, tested and applied consistently.


Blur can be modelled as convolution

An out-of-focus or motion-blurred image can be represented approximately as a sharp image convolved with a point-spread function, plus noise. Deblurring tries to invert that process.

Inverse problems are difficult because information may be attenuated or lost. Dividing by tiny frequency responses amplifies noise. Regularisation adds assumptions that stabilise the solution.

This connects computer vision to the mathematics in MRI reconstruction and CT scans: measurements do not automatically reveal the desired image without a model.


Lighting creates a nuisance variable

Pixel intensity depends on illumination, surface reflectance, camera response and geometry. The same object can have very different values under sunlight, shade or indoor lighting.

Colour constancy, histogram methods and learned augmentation try to reduce sensitivity, but none guarantees perfect invariance. A model trained in one environment may fail in another.

Testing should therefore stratify performance by lighting and other relevant conditions rather than report one average.


Scale and viewpoint change the pattern

An object far away covers fewer pixels. Rotating it changes edges and occlusion. Perspective changes apparent lengths and angles.

Image pyramids, multi-scale features and geometric transformations help. Homogeneous coordinates and projective geometry explain how 3D points map to a 2D image.

Students can continue with Why Mathematics? | Perspective Drawing, Vanishing Points and Projective Geometry to see the geometry beneath camera images.


Bias can enter before training begins

If training images underrepresent relevant groups, devices, locations or conditions, performance may differ across them. Labels can also encode human disagreement or historical decisions.

Mathematics can measure subgroup error rates and calibration, but a metric cannot decide which inequalities are acceptable. Dataset documentation, domain expertise and stakeholder review are also required.

Responsible computer vision distinguishes technical accuracy from social permission. A system can detect something reliably and still be inappropriate to deploy.


Privacy is not solved by accuracy

Images can contain faces, locations, documents and bystanders. Collecting more data may improve a model while increasing privacy risk.

Data minimisation, purpose limitation, access control and retention decisions sit beside algorithm design. Blurring or embedding does not automatically remove identifiability.

Students should learn that computing capability is constrained by law, ethics and human rights. Mathematics helps quantify risk; governance decides what should be done.


Distribution shift tests transfer

A model may perform well on a held-out sample drawn from the same source yet fail on a new camera, season, hospital or road. The input distribution has shifted.

Robust evaluation uses external datasets, time-based splits or site-based splits when relevant. Monitoring compares deployment data and error patterns with validation assumptions.

This is why a benchmark result is not a permanent certificate. Evidence must match the place and time of use.


Separable kernels reduce computation

Some two-dimensional kernels can be written as the outer product of one column vector and one row vector. A 3×3 Gaussian-like kernel, for example, can be applied as one horizontal 1×3 pass followed by one vertical 3×1 pass.

A direct K×K filter needs roughly K² multiplications per output channel and location. A separable version needs about 2K. For K=7, that changes 49 multiplications into about 14, before implementation details.

The output is mathematically equivalent when the kernel is exactly separable and the same boundary rules are used. This is an elegant example of algebra improving efficiency without changing the intended transformation.


Frequency gives another view of image detail

Slowly varying brightness corresponds to low spatial frequencies; rapid alternation corresponds to high spatial frequencies. Smoothing suppresses high-frequency components, while derivative and sharpening filters emphasise certain high-frequency changes.

The convolution theorem states that convolution in the spatial domain corresponds to multiplication in the frequency domain. For large kernels, fast Fourier transform methods can sometimes be efficient, although small local kernels are often faster directly.

Frequency language does not mean an image is vibrating through time. Spatial frequency measures repetition across distance in the image. Keeping the domain clear prevents a useful analogy from becoming a category error.


Quantisation limits available information

Digital intensities use finite levels. An 8-bit channel has 256 possible integer values. Mapping a wide physical brightness range into those levels can lose subtle differences.

If two scene intensities map to the same stored value, later processing cannot recover their original distinction from that image alone. Increasing contrast can make the step between levels more visible but does not recreate missing measurements.

Bit depth, clipping and compression therefore matter before edge detection begins. A perfect gradient formula cannot repair information that the acquisition pipeline never retained.


Computational cost scales across dimensions

For an output of height H and width W, Cout filters, Cin input channels and K×K kernels, a rough multiplication count is H×W×Cout×Cin×K². Real libraries optimise this heavily, but the scaling reveals which choices are expensive.

Doubling image width and height multiplies pixel locations by four. Doubling both input and output channels multiplies channel combinations by four again. Resource limits shape model design.

Efficiency is not merely about speed. Smaller computations can reduce energy use, enable on-device privacy and make systems accessible where connectivity is limited.


Explainability needs task-specific evidence

Saliency maps and feature visualisations can show where a model response changes, but they do not necessarily provide a causal or human-complete explanation.

A colourful heatmap can look persuasive even when unstable. Explanations should be tested for faithfulness, not judged only by appearance.

In high-stakes settings, the system may need interpretable measurements, uncertainty, audit trails and human review in addition to predictive performance.


A safe student investigation

Choose a small public image with clear geometric shapes. Convert it to greyscale, inspect pixel values, then apply an averaging kernel and Sobel kernels using a spreadsheet or a short program.

Predict the result before running it. Compare zero, reflected and repeated border handling. Change one threshold at a time and record how the edge map changes.

Do not use private photographs or scrape people’s images. The learning goal is to understand transformations and evaluation, not to build surveillance.


Common misconceptions to repair

  • “A pixel is a tiny object.” It is a sampled value tied to a sensor and coordinate.
  • “Convolution understands the picture.” It performs a weighted local calculation.
  • “Every strong gradient is an object boundary.” Shadows, texture and noise also create gradients.
  • “Edge detection and object recognition are the same.” Edges are one kind of local evidence.
  • “A high accuracy score proves safety.” Dataset fit, subgroup errors and deployment context still matter.
  • “More training images always solve bias.” Coverage, labels, objectives and governance also matter.
  • “AI removes the need for mathematics.” AI systems are built and evaluated through mathematical models.

A practical learning path for students

Begin with arrays, coordinates and weighted sums. Calculate a 3×3 filter by hand. Then study finite differences, vector magnitude and arctangent. Learn output-size formulas for padding and stride.

Next, implement smoothing, Sobel and Canny on open images. Build confusion matrices and vary a threshold. After that, study linear algebra, probability, optimisation and neural networks.

Keep a notebook that records input range, preprocessing, kernel, border rule, threshold, dataset and metric. Reproducibility is a technical skill and a thinking habit.


What parents can encourage

Ask the student to explain what the computer actually receives. If the answer jumps straight to “it sees a dog”, return to pixels, arrays and measurements.

Encourage small worked examples before large libraries. A student who can multiply one neighbourhood by one kernel understands more than a student who only presses Run.

Also ask what could make the model fail and whose data were used. Technical curiosity and ethical curiosity belong together.


Did you know? A three-by-three matrix can find a boundary

The Sobel example used only nine weights, yet it detected a strong change across a tiny image patch. That compact operation scales because it can be repeated across millions of locations and many channels.

Modern models learn thousands or millions of such weights, but the core discipline remains: multiply, add, transform, compare and test.

This is one of the clearest benefits of learning mathematics. Simple structures can be composed into systems that work on problems far larger than the original example.


Frequently asked questions

Is convolution the same in mathematics and software?

Not always. Mathematical convolution flips the kernel; some software functions perform correlation without the flip while using convolution terminology. Symmetric kernels behave the same. Check the documentation.

Why smooth before detecting edges?

Derivative operations amplify rapid variation, including noise. Smoothing reduces small fluctuations so the later gradient is less likely to treat them as edges. Excessive smoothing can remove real detail.

What does an edge detector return?

It returns a numerical response or an edge map based on intensity change and selected rules. It does not automatically identify the object or explain the scene.

Why use two thresholds in Canny?

The high threshold identifies strong candidates. Medium responses survive only when linked to strong edges, helping keep faint continuations without accepting every moderate response.

Do neural networks still use kernels?

Many convolutional networks learn kernels. Other modern architectures may use attention or mixed designs. The broader mathematical foundations include linear algebra, optimisation and probability.

Which school mathematics matters most?

Matrices, coordinates, functions, trigonometry, vectors and statistics are useful. Calculus explains gradients and optimisation. Programming turns the ideas into repeatable experiments.

Does computer vision guarantee a career in AI?

No. It is one foundation among many. Relevant pathways also need computing, data practice, domain knowledge, communication, responsible design and continued learning.


Useful next reading

Start with Why Mathematics? | Digital Images, Aspect Ratios and Pixel Counts for image dimensions, then continue to Why Mathematics? | Machine Learning, Loss Functions and Gradient Descent for learning and optimisation. Why Mathematics? | Bézier Curves, Animation and Digital Design offers a complementary geometry route.

For official implementation references, see OpenCV’s Sobel Derivatives tutorial, Canny Edge Detector tutorial, and image-filtering documentation.

The central idea is joyful and precise: a computer image begins as numbers, but those numbers are not meaningless. Mathematics gives us disciplined ways to reveal change, shape and structure while keeping limitations visible.

Discover more from eduKate Singapore

Subscribe now to keep reading and get access to the full archive.

Continue reading