Why is mathematics important in JPEG images? Because a camera or website cannot always store every pixel value directly. JPEG compression changes a picture into colour components, divides it into blocks, expresses each block as weighted spatial patterns, rounds those weights and encodes the resulting symbols efficiently.
The method is powerful because natural photographs usually contain smooth regions and correlated neighbouring pixels. A compact set of low-frequency patterns can describe much of that structure. Fine variations often need less precision, although the choice can create visible blocking, ringing and colour artefacts.
JPEG is lossy in its common photographic mode: decompression does not generally recreate every original sample. Mathematics makes the loss controlled and measurable. It also explains why saving the same edited JPEG repeatedly can degrade it and why text, diagrams and sharp interface graphics may need another format.
The Joint Photographic Experts Group maintains the official JPEG standards site, and the W3C's historical JPEG JFIF overview distinguishes the compression method from a common file interchange format. This article focuses on the classic block-based process associated with the original JPEG standard, not every later format created by the committee.
> Did You Know? JPEG does not simply throw away random pixels. It changes coordinates: from individual sample values to coefficients that say how much of each cosine pattern appears in an eight-by-eight block.
Quick Navigation
- The compression problem
- Pixels, channels and colour conversion
- Why JPEG uses blocks
- The discrete cosine transform
- A one-dimensional worked transform
- Quantisation creates the main loss
- Zigzag order, runs and entropy coding
- Decompression and visible artefacts
- Rate-distortion trade-offs
- A safe student simulation
- Guidance and pathways
- Frequently asked questions
The Compression Problem
Imagine an image with 4000 by 3000 pixels and three eight-bit channels. A simple uncompressed count is
\[ 4000\times3000\times3=36{,}000{,}000\text{ bytes}, \]
about 36 MB before headers and metadata. A photograph may contain substantial redundancy: neighbouring sky pixels are similar, broad walls change slowly, and colour detail can be less visually sensitive than brightness detail.
Compression seeks a shorter representation. Lossless compression must recover every input value exactly. Lossy compression permits differences chosen to reduce size while retaining useful visual quality.
Redundancy is opportunity
If a row contains 500 identical values, recording the value once plus a run length can be shorter than writing it 500 times. Photographs are not made only of identical runs, but local correlation provides a related opportunity.
Transform coding goes further. Instead of asking whether adjacent values repeat exactly, it asks whether a block can be described by a small number of smooth patterns.
File size is not only pixel count
Two images with identical dimensions and quality settings can produce different JPEG sizes. A smooth, slightly blurred sky tends to need fewer nonzero high-frequency coefficients than grass, hair or sensor noise. Entropy coding also depends on how predictable the symbols are.
Dimensions set the number of blocks, but image statistics determine how efficiently those blocks encode.
Compression ratio
If an uncompressed representation is 36 MB and the JPEG is 4.5 MB, the simple compression ratio is
\[ \frac{36}{4.5}=8:1. \]
This ratio depends on what is counted as the original and whether units use decimal or binary prefixes. A ratio does not by itself state visual quality.
Pixels, Channels and Colour Conversion
A digital colour pixel is often represented initially by red, green and blue components. Classic JPEG commonly transforms them into a luma-like component (Y) and two chroma-like difference components.
A simplified linear transform has the form
\[ \begin{bmatrix}Y\\C_b\\C_r\end{bmatrix} =M\begin{bmatrix}R\\G\\B\end{bmatrix}+b, \]
where matrix (M) contains channel weights and (b) supplies offsets for the chosen numeric representation. Exact coefficients and ranges depend on the coding convention.
Why separate brightness and colour?
Human vision is often more sensitive to fine spatial detail in luma than in chroma. By separating them, an encoder can preserve luma resolution while sampling colour-difference channels more sparsely.
This is not a claim that colour never matters. Thin coloured text, scientific heat maps and saturated edges can be damaged by chroma subsampling.
Chroma subsampling
Notation such as 4:4:4 and 4:2:0 describes relative chroma sampling arrangements. In a common 4:2:0 arrangement, chroma has lower horizontal and vertical sampling density than luma.
If a 4000 by 3000 image has 12 million luma samples, full-resolution two chroma channels add 24 million samples. At one quarter spatial sample count per chroma channel, they add 6 million instead. Before transform coding, total component samples fall from 36 million to 18 million.
The exact interpretation includes sampling geometry, but the arithmetic reveals why subsampling saves data.
Conversion is reversible only under conditions
A floating-point colour transform can be inverted mathematically if its matrix is nonsingular. In practice, limited ranges, rounding and chroma subsampling lose information. The later inverse estimates RGB values; it does not reopen a hidden exact original.
Why JPEG Uses Blocks
In the common sequential DCT process, each component is divided into eight-by-eight sample blocks. A 4032 by 3024 luma plane divides exactly into 504 by 378 blocks. Other dimensions may require edge extension or padding so complete blocks can be processed.
Locality
Within a small region, brightness often changes smoothly enough for a few cosine patterns to describe it. Applying a transform to the entire image would create different computational and error-localisation properties. Blocks make the process practical and independently decodable.
The price of boundaries
Each block is transformed and quantised separately. At strong compression, neighbouring blocks may reconstruct with slightly different averages or edge behaviour, revealing an eight-pixel grid.
The block structure is therefore both an engineering advantage and a source of characteristic artefacts.
Level shifting
Eight-bit samples commonly range from 0 to 255. Before the DCT, an encoder level-shifts them by subtracting 128, giving values centred around zero.
Centred values make the transform's constant component represent deviation around a neutral midpoint. The inverse process adds 128 and clips or rounds to the output sample range.
An eight-by-eight block as a vector
Sixty-four samples can be arranged as a vector (x\in\mathbb R^{64}). The transform is a matrix multiplication
\[ c=Tx, \]
where (T) is an orthogonal transform matrix under a suitable normalisation and (c) contains 64 coefficients. The two-dimensional separable transform can be computed by rows and columns rather than multiplying one enormous matrix directly.
The Discrete Cosine Transform
The two-dimensional DCT represents an eight-by-eight block as a weighted sum of cosine basis patterns. One standard normalised form is
\[ F(u,v)=\frac14 C(u)C(v)\sum_{x=0}^{7}\sum_{y=0}^{7} f(x,y)\cos\frac{(2x+1)u\pi}{16}\cos\frac{(2y+1)v\pi}{16}, \]
where (C(0)=1/\sqrt2) and (C(k)=1) for nonzero (k).
DC and AC coefficients
(F(0,0)) is the DC coefficient. Both cosine factors are constant, so it is proportional to the block's average level after shifting.
The other 63 are AC coefficients. Small indices represent slow horizontal or vertical changes. Large indices represent rapid alternation.
Orthogonality
The basis patterns are orthogonal. Their inner products are zero when they are different, under the transform's discrete weighting. Orthogonality allows each coefficient to measure one independent basis direction.
If (T) is orthonormal,
\[ T^{-1}=T^T. \]
Before quantisation and rounding, the inverse transform can reproduce the input block apart from finite numerical precision.
Energy compaction
Natural image blocks often place much of their squared coefficient magnitude in low frequencies. This is called energy compaction. It does not mean every image behaves that way. Checkerboards, fine fabric and noise can create large high-frequency coefficients.
The transform itself does not compress the number of values: 64 samples become 64 coefficients. Compression becomes possible because the new values have a distribution that can be quantised and encoded more economically.
Separable computation
A two-dimensional DCT can be applied across every row, then down every column:
\[ F=AfA^T. \]
This reduces computation and makes the structure easier to teach. It also shows how one-dimensional transformations combine to describe two-dimensional frequency.
A One-Dimensional Worked Transform
Use a four-sample example to make the idea visible. Let shifted samples be
\[ x=[10,10,10,10]. \]
This sequence is constant. Its transform has one nonzero DC coefficient and zero ideal AC coefficients. No oscillating pattern is needed to describe it.
Now consider
\[ y=[-10,-10,10,10]. \]
The average is zero, so its DC term is zero. Low-frequency cosine terms describe the broad transition from negative to positive. A rapidly alternating sequence
\[ z=[-10,10,-10,10] \]
places more magnitude at high frequency.
Projection onto a basis
If (b_k) is a unit-length basis vector, coefficient (c_k) is a dot product:
\[ c_k=x\cdot b_k. \]
The dot product is large when the sample pattern resembles the basis. Reconstruction adds all weighted bases:
\[ x=\sum_k c_kb_k. \]
JPEG's two-dimensional version uses 64 basis images instead of four basis vectors.
A constant eight-by-eight block
Suppose every shifted sample is 20. The DC coefficient under the displayed normalisation is
\[ F(0,0)=\frac14\frac1{\sqrt2}\frac1{\sqrt2}(64)(20)=160. \]
All ideal AC terms vanish because positive and negative cosine portions cancel. One meaningful number describes the entire block before coding overhead.
Why edges need more coefficients
A sharp edge has an abrupt transition. Reconstructing it with only smooth low-frequency cosines produces overshoot or ringing near the boundary. Adding higher frequencies sharpens the transition. Quantising those components heavily therefore affects edges more visibly.
Quantisation Creates the Main Loss
After transformation, each coefficient is divided by a corresponding quantisation step and rounded:
\[ Q(u,v)=\operatorname{round}\left(\frac{F(u,v)}{S(u,v)}\right). \]
The decoder multiplies back:
\[ \hat F(u,v)=Q(u,v)S(u,v). \]
Unless (F) was an exact multiple of (S), \(\hat F\ne F\). This is the principal lossy step.
Worked coefficient
If a coefficient is 37 and its step is 8,
\[ Q=\operatorname{round}(37/8)=5, \]
and reconstruction gives (5\times8=40). The coefficient error is +3.
If another coefficient is 3 with step 12, it rounds to zero. That spatial-frequency pattern disappears from the block.
The quantisation matrix
JPEG stores or references tables of step sizes, commonly using smaller steps for visually important low frequencies and larger steps for high frequencies. Luma and chroma can use different tables.
A software “quality” slider is not a standard universal percentage. Applications map slider values to tables in their own ways. Quality 80 in one program need not match quality 80 in another.
Error in coefficient and pixel space
With an orthonormal transform, total squared error is preserved between coefficient space and sample space before clipping and other operations. If coefficient error vector is (e_c), pixel error is (e_x=T^Te_c), and
\[ \|e_x\|_2^2=\|e_c\|_2^2. \]
The transform redistributes error spatially; it does not make error vanish. Perceptual design chooses which coefficient errors are less noticeable.
Dead zones and zeros
Rounding creates an interval around zero that maps to zero. Larger high-frequency steps therefore produce long runs of zero coefficients after ordering. Those zeros are excellent for later compression but remove fine detail.
Zigzag Order, Runs and Entropy Coding
After quantisation, coefficients are arranged in a zigzag sequence that starts at DC and generally moves from low to high spatial frequency. Because high-frequency values are often zero, the sequence tends to end with a long zero run.
DC differences
Neighbouring block averages are often similar. Rather than encode each DC coefficient independently, the encoder can code the difference from the previous block's DC value.
If successive DC values are 120, 124, 123 and 125, differences after the first are +4, -1 and +2. Small differences can be encoded efficiently.
Run-length descriptions
An AC sequence might begin
\[ [7,-2,0,0,0,1,0,0,\ldots]. \]
Instead of recording every zero, symbols describe the number of preceding zeros and the size category of the next nonzero value. A special end-of-block symbol marks that all remaining coefficients are zero.
Entropy coding
Frequent symbols receive shorter codes and rare symbols longer codes, subject to the coding method. Classic JPEG commonly uses Huffman coding, while arithmetic coding is also defined in the broader standard family.
If symbol probabilities are (p_i), entropy is
\[ H=-\sum_i p_i\log_2p_i. \]
It gives a theoretical average-information measure. A practical prefix code has constraints and headers, so its average length is not automatically exactly the entropy.
A small Huffman idea
Suppose symbol A occurs with probability 0.5, B with 0.25, C and D with 0.125 each. Codes A=`0`, B=`10`, C=`110`, D=`111` have average length
\[ 0.5(1)+0.25(2)+0.125(3)+0.125(3)=1.75\text{ bits}. \]
Using two bits for every symbol would average 2 bits. Probability has reduced the expected length without losing symbol information.
The second lossless stage
Run-length and entropy coding are lossless with respect to the quantised coefficients. They shorten the representation but can be reversed exactly. The loss already entered through chroma subsampling, rounding and quantisation.
Decompression and Visible Artefacts
A decoder entropy-decodes symbols, restores runs, de-zigzags coefficients, multiplies by quantisation steps, applies the inverse DCT, adds the level shift, upsamples chroma and converts colour for display.
Blocking
Heavy quantisation can change the reconstructed average or low-frequency slope differently in neighbouring blocks. Their boundaries become visible as an eight-by-eight grid.
Ringing
Removing high-frequency coefficients from a sharp edge leaves a limited sum of smooth cosine waves. Oscillations can appear near the edge as halos or ripples.
Mosquito noise
Small high-contrast details may acquire flickering-looking speckles or ripples around them. This is especially distracting around text and line art.
Colour bleeding
Chroma subsampling and quantisation can spread or shift colour boundaries. Red text on a dark background is a classic stress case because chroma detail matters strongly.
Banding
Smooth gradients can show steps after strong quantisation, colour conversion or limited precision. Dithering can trade structured bands for less noticeable noise, but JPEG does not guarantee a band-free gradient.
Repeated saving
Opening and viewing a JPEG does not inherently degrade it. Decoding, editing and re-encoding can. A second encoder transforms pixels already containing first-generation error, and block alignment or settings may differ.
For work in progress, retain an original or lossless master. Export JPEG for an intended delivery size and quality rather than treating it as an endlessly editable source.
Rate-Distortion Trade-offs
Compression balances rate, meaning bits or file size, against distortion, meaning reconstruction error under a chosen measure.
Mean squared error
For original samples (x_i) and reconstruction \(\hat x_i\),
\[ MSE=\frac1n\sum_{i=1}^{n}(x_i-\hat x_i)^2. \]
MSE penalises large numeric errors strongly. It does not know whether an error falls on a face, flat sky or invisible texture.
Peak signal-to-noise ratio
For maximum sample value (MAX),
\[ PSNR=10\log_{10}\left(\frac{MAX^2}{MSE}\right). \]
With eight-bit values, (MAX=255). If MSE is 25,
\[ PSNR=10\log_{10}(255^2/25)\approx34.15\text{ dB}. \]
A higher PSNR means lower MSE, but it does not guarantee a more pleasing image. Two images with equal MSE can distribute error differently.
Bits per pixel
If a 12-megapixel JPEG occupies 3 MB using decimal units, approximate bits per pixel are
\[ \frac{3{,}000{,}000\times8}{12{,}000{,}000}=2.0. \]
Metadata and headers are included in file size, so the actual compressed image payload rate differs slightly.
Optimising a table
A conceptual objective is
\[ J=D+\lambda R, \]
where (D) is distortion, (R) is rate and \(\lambda\) states how much rate matters. A larger \(\lambda\) favours smaller files; a smaller value favours fidelity.
Real encoders use heuristics, perceptual models and search strategies. The equation clarifies that “best” has no meaning until the trade-off is specified.
Content changes the curve
Smooth portraits, noisy night scenes, screenshots and scanned text have different rate-distortion behaviour. A setting that works for one image is not a universal prescription.
Progressive JPEG and Resilience
A baseline sequential file sends the main information in one scan order. A progressive JPEG can send coefficient information in multiple scans so a rough image appears before finer detail arrives.
Spectral selection and refinement
One progressive strategy transmits low-frequency coefficients first and higher frequencies later. Another refines coefficient precision over successive scans. The final decoded samples can represent the same quantised coefficients as a sequential encoding while the transmission experience differs.
Restart markers
Encoded data can include restart intervals that help a decoder resynchronise after certain errors and permit bounded processing segments. They add overhead but improve operational robustness.
Metadata is separate
A JPEG-containing file may store camera settings, colour profiles, thumbnails, location and copyright information. Removing metadata can shrink the file and protect privacy, but does not change the core DCT coefficients unless the image is re-encoded.
Students should distinguish image content, compression syntax and surrounding metadata rather than calling all of them “the JPEG algorithm.”
Integer Arithmetic and Implementation Accuracy
The transform formula is written with cosines and real numbers, but production codecs often use carefully scaled integer arithmetic for speed and reproducibility. Multiplication by a real constant can be approximated by multiplying an integer and shifting a fixed number of binary places.
Fixed-point representation
Suppose a coefficient 0.7071 is represented with scale (2^{12}=4096). Its stored integer approximation is
\[ \operatorname{round}(0.7071\times4096)=2896. \]
Multiplying value (x) by 2896 and then dividing by 4096 approximates (0.7071x). The approximation error must remain within the codec's conformance expectations.
Scaling can be folded into quantisation
Fast DCT algorithms factor the transform into fewer additions and multiplications. Some postpone normalisation factors and combine them with quantisation steps. Algebraically equivalent factorizations can therefore have different intermediate values while producing acceptably matched decoded results.
Overflow and range analysis
An eight-bit input becomes a signed level-shifted value. Sums of 64 values and scaled multiplications require wider intermediate types. Engineers bound the largest possible magnitude before choosing integer widths.
If every shifted sample has magnitude at most 128, the magnitude of their unweighted sum is at most
\[ 64(128)=8192. \]
Cosine scaling and factorisation add further requirements. A program that ignores these bounds can overflow and create dramatic artefacts even when the mathematics on paper is correct.
Decoder differences
Standards specify decoding requirements and acceptable precision, but two implementations can differ slightly because of rounding and colour conversion. Testing uses known vectors and tolerance rules rather than assuming every intermediate bit must match one reference program.
This distinction is useful beyond JPEG: a mathematical function, a numeric approximation and a software implementation are three layers that must be validated separately.
Markers, Tables and the Byte Stream
The compressed coefficients need a structured byte stream so a decoder knows image dimensions, component sampling, quantisation tables, code tables and scan boundaries.
Marker segments
JPEG streams use markers to announce structural information such as start of image, frame parameters, tables, scans and end of image. A parser must respect segment lengths and validate inputs rather than searching casually for a byte pattern.
Tables are part of the meaning
A quantised coefficient value of 5 does not reconstruct to one universal number. The decoder needs the corresponding quantisation step. Likewise, a sequence of entropy bits needs the code table that maps prefixes to symbols.
If a file loses or corrupts required tables, the remaining coefficients cannot be interpreted reliably. Data and metadata are mathematically coupled.
Byte stuffing
Because marker bytes have special meaning, entropy-coded data uses an escaping convention when a marker-like byte value occurs as data. This is an example of syntax design: reserve patterns for control, then provide a reversible way to represent the same pattern as ordinary payload.
Robust parsing is a safety issue
Image decoders process files from cameras, websites and messages. Dimensions, table identifiers, lengths and arithmetic must be checked before allocating memory or indexing arrays. A mathematically valid transform does not make an unchecked parser safe.
Students interested in cybersecurity can see the connection: format mathematics defines valid relationships, while secure engineering rejects values that would violate memory, range or resource limits.
Choosing an Image Format for the Job
JPEG's strengths should be evaluated against the content and workflow.
Photographic sharing
Continuous-tone photographs often benefit from JPEG's energy compaction and wide compatibility. Resize to the required output dimensions before export, because storing thousands of unseen pixels wastes bandwidth even if compression is efficient.
Text and diagrams
Screenshots, line drawings and text have sharp edges and repeated flat colours. Lossless formats can preserve exact boundaries and often compress those patterns effectively. JPEG may create halos and coloured fringes.
Transparency
Classic JPEG does not provide a general alpha transparency channel. A logo needing soft transparency requires another representation or a chosen background before JPEG export.
Archival masters
An archive should preserve enough information for future uses. A delivery JPEG may be perfect for a current webpage yet unsuitable as the only surviving master. Original capture data, colour profile, editing history and lossless derivatives may have long-term value.
Web performance
Smaller images can reduce transfer time, but page performance also depends on dimensions, responsive selection, caching, connection and decode cost. A 200 KB image displayed at 300 pixels wide can still be wasteful if its encoded dimensions are 5000 pixels.
The optimisation objective is the user experience at adequate quality, not winning a file-size contest in isolation.
A Second Worked Example: Quantisation and MSE
Consider four transform coefficients
\[ F=[160,37,-11,3] \]
with steps
\[ S=[8,8,10,12]. \]
Quantised values are
\[ Q=[20,5,-1,0]. \]
Dequantised coefficients are
\[ \hat F=[160,40,-10,0]. \]
Coefficient errors are ([0,3,1,-3]). Their squared-error sum is
\[ 0^2+3^2+1^2+(-3)^2=19. \]
Under an orthonormal four-element transform, the sample-domain squared-error sum is also 19. Sample MSE is (19/4=4.75).
If every step doubles, more coefficients may round to zero and rate may fall, but distortion generally rises. If steps halve, distortion generally falls but more magnitude and nonzero symbols must be encoded.
This tiny example omits chroma and entropy overhead, yet it demonstrates rate-distortion causality. The quantisation table is not decoration; it is the main precision budget.
Ethical Use and Image Evidence
Compression artefacts can resemble edges, texture or small objects that were not present in the source. Enlarging a low-quality JPEG does not restore the discarded coefficients; an enhancement model may invent a plausible estimate.
In journalism, science and investigations, keep provenance and original files. Do not interpret a block boundary or ringing halo as physical evidence without checking the source and compression history.
Metadata can be edited or removed, so it is not proof by itself. Conversely, absent metadata does not prove manipulation. Evidence requires a documented chain, corroboration and methods appropriate to the claim.
This is a vital mathematical habit: distinguish observed samples, encoded representation, reconstructed pixels and human interpretation. Each transformation changes what can be concluded.
A Safe Student Simulation
Students can model JPEG ideas with an eight-by-eight greyscale spreadsheet or a short local program. Use invented values, not private photographs.
Step one: make three blocks
Create a constant block, a smooth gradient and a checkerboard. Subtract 128 if values are in the eight-bit range. Predict which will need the most high-frequency coefficients.
Step two: apply a DCT
Use a trusted numerical library or teacher-provided cosine matrix. Compute (F=AfA^T). Display coefficient magnitudes as a heat map.
The constant block should concentrate at DC. The gradient should favour low frequency. The checkerboard should show high-frequency energy.
Step three: quantise
Divide coefficients by a table and round. Count nonzero coefficients. Multiply back and apply the inverse DCT.
Repeat with steps scaled by 0.5, 1 and 2. Record file-size proxies such as number of nonzeros and distortion such as MSE.
Step four: inspect boundaries
Place four blocks side by side and quantise independently. If neighbouring gradients reconstruct differently, block edges may appear. Try shifting the image by one pixel before blocking and compare artefacts.
Step five: entropy exercise
Write a zigzag sequence, convert zero runs to symbols, count frequencies and build a small Huffman tree. Compare fixed-length and average code length.
Report limitations
A classroom simulation may omit exact integer DCT scaling, marker syntax, byte stuffing, sampling alignment, colour profiles and encoder optimisation. It should demonstrate the mechanism without claiming bit-for-bit compatibility.
Common Misconceptions
“JPEG deletes every nth pixel”
Common JPEG transform coding converts blocks to coefficients, quantises them and encodes the results. Chroma may be subsampled, but the core luma process is not periodic pixel deletion.
“The DCT is the lossy step”
The ideal transform is invertible. Quantisation and subsampling introduce the principal designed losses.
“Quality 90 means 90 per cent of data remains”
The slider is an application-specific mapping to encoder choices. It is not a universal retained-information percentage.
“A smaller JPEG always looks worse”
At equal dimensions and competent encoding, stronger compression often increases distortion, but image content, encoder and metadata also affect size. A noisy large file can look worse than a clean smaller one.
“Opening a JPEG damages it”
Decoding for viewing does not change the stored file. Editing and re-encoding can add loss.
“JPEG is best for every image”
Photographs often compress well. Text, logos, diagrams, transparency and repeated editing may favour lossless or other modern formats.
“No visible difference means no mathematical difference”
Pixel values can differ below a particular viewing threshold. Visibility depends on display, size, distance, content and observer.
Guidance and Pathways
JPEG connects school mathematics to matrices, trigonometry, vectors, rounding, probability, logarithms and optimisation. It is an unusually complete example of a pipeline in which each mathematical choice serves an engineering constraint.
Students should practise asking:
- Which representation are these numbers in?
- Which step is reversible?
- Where does rounding enter?
- What error measure matches the purpose?
- Does the image contain smooth detail, edges or noise?
- Are file size and visual quality being compared at equal dimensions?
- Is the master file protected before export?
For parents and teachers, the best demonstration uses tiny invented blocks. A learner who can explain why a constant block needs one coefficient and a checkerboard needs high frequencies has understood the central mechanism.
This mathematics leads toward image processing, computer vision, multimedia engineering, web performance, medical imaging, remote sensing and digital preservation. No one codec guarantees a career. The Mathematics Pathways guide helps students connect foundations to later choices.
Useful Next Reading
Computer vision, convolution and edge detection explains how algorithms extract structure after an image is decoded. MRI, k-space and Fourier image reconstruction uses frequency-domain reasoning for a very different measurement problem.
Ray tracing, vectors and light reflection shows how synthetic images are generated before compression. Students can use how to check your work to verify dimensions, units and reconstruction calculations.
The eduKate Sengkang Mathematics Hub provides broader navigation, while the Mathematics Year 0 to Adulthood Engineer Series places coding mathematics in a longer learning pathway.
Frequently Asked Questions
What does JPEG stand for?
It names the Joint Photographic Experts Group, the standards committee associated with the image coding family.
Is JPEG a file format or a compression method?
JPEG defines coding processes; formats such as JFIF and Exif organise JPEG-compressed images and metadata for exchange. Everyday language often calls the whole file a JPEG.
Why are blocks eight by eight?
The classic design balances local energy compaction, computation, memory and coding overhead. The choice is part of the standard process, not a law of nature.
What is the DC coefficient?
It is the zero-frequency transform coefficient, proportional to the block's average after level shifting.
Why are there 63 AC coefficients?
An eight-by-eight block has 64 basis patterns. One is constant DC; the remaining 63 describe spatial variation.
What does quantisation do?
It divides coefficients by step sizes and rounds them, reducing precision and creating many zeros. This is the main lossy step.
Why does JPEG look blocky at low quality?
Blocks are processed separately. Strong quantisation makes differences at their boundaries visible.
Why is text often poor in JPEG?
Sharp high-contrast edges require high-frequency information and colour subsampling can damage coloured edges. Lossless formats often preserve such graphics better.
Can the original be recovered from a JPEG?
Not generally after lossy quantisation and subsampling. Decompression reconstructs an approximation consistent with stored coefficients.
Does a higher PSNR guarantee a better-looking image?
No. It means lower mean squared error under a specific numeric comparison. Perceptual importance and artefact location are not fully captured.
Is progressive JPEG higher quality?
Progressive coding changes how information arrives in scans. With equivalent quantised coefficients, its final reconstruction need not be inherently higher quality.
Should students repeatedly save experiments as JPEG?
Keep an original or lossless working file and export fresh JPEG copies. Repeated decoding, editing and re-encoding can accumulate loss.
Final Perspective: Change Coordinates, Then Spend Precision Wisely
JPEG compression is a chain of mathematical ideas. Colour conversion separates kinds of information. Subsampling reduces selected spatial detail. The DCT changes from pixels to cosine coefficients. Quantisation spends precision unevenly. Run-length and entropy coding exploit the resulting probabilities.
Every stage has a purpose and a limit. Smooth photographs often compress efficiently because low frequencies dominate. Sharp text and repeated patterns reveal the assumptions. Strong rounding creates small files and larger artefacts.
That is why mathematics matters in image compression. It turns “make this picture smaller” into a transparent set of choices about representation, error, probability and human perception. The best file is not simply the smallest. It is the smallest representation that remains fit for its intended use while preserving an appropriate master for the future.
