VIEW THIS AS

Auto mode follows the Route Engine until you choose a viewpoint.

YOU ARE HERE

ROUTE CHECK

CONNECTED TO

WHAT NEXT

Use the canonical route for this room, or HELP if you are unsure.

Why Mathematics? | Hard Disk Drives, Areal Density and Rotational Latency

Three learners review open books together at a classroom table, with stacks of textbooks, stationery and a whiteboard in the bright room.

Why is mathematics important in a hard disk drive? A file feels like one object, yet the drive stores patterns on rotating surfaces arranged into tracks and sectors. To retrieve a block, an actuator moves a head to the right radius, the platter rotates the desired sector underneath it, and electronics transfer and check the data. Geometry, angular motion, probability and scheduling all shape the result.

The mathematics explains two different questions. Capacity depends on usable area, bit density, track spacing, formatting and redundancy. Access time depends on seek motion, rotational position, transfer rate and waiting behind other requests. A larger drive is not automatically faster, and a high rotational speed does not remove every delay.

This article builds an educational model of magnetic hard disk drives. Modern products use proprietary servo systems, coding, caching and recording technologies, so the simplified calculations are not specifications or purchasing advice. They are a way to understand why data storage needs much more than counting bytes.


From Platter Area to Addressable Blocks

Concentric tracks

A hard disk platter is a circular surface coated with magnetic material. The recording region is an annulus rather than a full disc: it has an inner usable radius Ri and outer usable radius Ro. The head follows approximately circular tracks at different radii.

Area of one recording surface is A = pi(Ro squared – Ri squared). If Ro = 45 millimetres and Ri = 15 millimetres, usable geometric area is pi(2025 – 225) = 1800 pi, or about 5,655 square millimetres. A drive with four recording surfaces would have about 22,619 square millimetres before excluding servo, defects, formatting and reserved space.

This annulus formula matters because subtracting radii and then squaring is wrong. Pi(Ro – Ri) squared would describe a different circle, not the ring between two circles. Drawing both boundaries makes the subtraction of areas obvious.

Tracks, sectors and blocks

Each track is divided into sectors. Modern interfaces usually present logical block addresses, or LBAs, rather than asking software to specify a physical cylinder, head and sector. The drive’s controller maps logical requests to physical locations and manages defects and internal organisation.

A sector is not merely “one angle.” Outer tracks have greater circumference than inner tracks. If every track stored the same number of sectors, the longer outer circumference would waste potential recording length. Zone bit recording groups tracks into zones, with outer zones generally holding more sectors per revolution than inner zones.

Therefore capacity cannot be calculated accurately by multiplying one track count by one universal sectors-per-track value unless that value is explicitly an average or the model assumes constant geometry. The zone table is part of the physical format.

Circumference changes with radius

Track circumference is C = 2 pi r. At radius 40 millimetres, circumference is about 251.3 millimetres. At 20 millimetres, it is about 125.7 millimetres. The outer track is twice as long.

If linear bit spacing along the track were constant, the outer track could hold about twice as many bits. Real recording density, servo information and zone boundaries complicate that picture, but the proportional reasoning remains useful.

Logical capacity and physical capacity

Operating systems report logical capacity available through the interface. Physical recording contains more than user payload: sector headers, synchronisation fields, servo information, error-correcting redundancy, spare sectors and controller-managed metadata.

Manufacturers may state capacity with decimal prefixes, where one terabyte is 10 to the power 12 bytes. Many software displays use binary units or historically label them ambiguously. One tebibyte is 2 to the power 40 bytes, about 1.0995 trillion bytes. A nominal 1 TB drive therefore appears as about 0.91 TiB before partitions and filesystems.

The difference is a unit convention, not lost material. Filesystem overhead and reserved space are separate effects. Good explanations name both the numerical value and the unit definition.

Did You Know? The outside of a platter travels farther per revolution than the inside, even though every radius completes the same revolution at the same time. This is why outer zones can often deliver a higher sequential transfer rate.


Areal Density and Capacity

Bits per area

Areal density describes stored bits per unit surface area. A simplified recording model multiplies linear density along a track by track density across the radius:

areal density approximately equals bits per unit track length times tracks per unit radial length.

The first quantity is often discussed as bits per inch, BPI. The second is tracks per inch, TPI. Their product has units bits per square inch. This product is a useful conceptual model, although modern recording and coding make the engineering definition more nuanced.

If linear density were 1.2 million bits per inch and track density 300,000 tracks per inch, the product would be 3.6 times 10 to the power 11 bits per square inch, or 360 gigabits per square inch. Dividing by eight gives 45 gigabytes of raw bits per square inch before overhead.

Capacity from annular area

In a uniform teaching model, raw bit capacity equals usable area times areal density times number of recording surfaces. Units must agree. If radii are in millimetres but density is bits per square inch, convert area using 25.4 millimetres per inch.

For Ro = 45 mm and Ri = 15 mm, one-surface area is 5,654.9 mm squared. Divide by 25.4 squared, or 645.16, to obtain about 8.765 square inches. At 360 gigabits per square inch, raw capacity is about 3.155 terabits per surface, or 394 gigabytes per surface.

Four surfaces would give about 1.58 terabytes raw in this illustrative calculation. User capacity would be lower after format and reserved overhead. The chosen densities are illustrative and should not be attributed to an actual product.

Nonuniform reality

The uniform-area method hides zones, guard bands, servo wedges, head-dependent performance and unusable regions. It is valuable for scaling: doubling areal density approximately doubles raw capacity if usable area and surfaces stay fixed. But it cannot reconstruct a specific drive’s format.

A zone model is more detailed. For zone j, capacity equals number of tracks in the zone times sectors per track times bytes per sector. Total logical capacity is the sum across zones and surfaces, after the controller’s presentation rules.

Why shrinking dimensions is difficult

Higher linear density places transitions closer together. Higher track density narrows track pitch. The read head must distinguish smaller signals while servo control keeps it accurately positioned. Thermal stability, neighbouring-track interference and signal-to-noise ratio become challenging.

Mathematics supports trade-offs among coding rate, error probability, track pitch, head positioning and manufacturing variation. A headline areal-density number does not reveal all these constraints.

Percentage growth

If areal density rises from 800 to 1,000 gigabits per square inch, percentage increase is (1,000 – 800)/800 times 100 = 25 percent. The new value is 1.25 times the old.

If usable area falls by 4 percent while density rises 25 percent, combined raw capacity factor is 1.25 times 0.96 = 1.20, or 20 percent growth. Percentage changes multiply; they should not always be added.

Decimal and binary conversion

A sector with 4,096 payload bytes contains 32,768 payload bits. One million such sectors contains 4.096 billion bytes, or 4.096 GB in decimal units. In binary, divide by 2 to the power 30 to obtain about 3.815 GiB.

The same data can have different numerical sizes when units differ. That is not measurement uncertainty. It is a definition choice that should be stated.


RPM, Angular Speed and Rotational Latency

Revolutions per minute

RPM states how many full revolutions occur per minute. Rotational frequency f in revolutions per second is RPM/60. Angular speed omega = 2 pi f radians per second.

For 7,200 RPM, f = 120 revolutions per second. One revolution takes 1/120 second = 8.333 milliseconds. Angular speed is 240 pi, about 754 radians per second.

For 5,400 RPM, one revolution takes 60/5,400 second = 11.111 milliseconds. For 10,000 RPM, it is 6 milliseconds. These are revolution times, not complete access times.

Rotational latency

After the head reaches the right track, the desired sector may not yet be under it. Rotational latency is the wait for the platter to rotate to the sector’s angular position.

If request arrival is independent of platter angle and target positions are uniformly distributed, latency ranges from nearly zero to one revolution. The expected latency is half a revolution:

average rotational latency = 30/RPM seconds.

At 7,200 RPM, this is 30/7,200 second = 0.004167 second, or 4.167 milliseconds. At 5,400 RPM it is 5.556 ms; at 10,000 RPM it is 3.000 ms.

This is an average under a randomness assumption. One request might arrive just before its sector and wait almost nothing. Another might just miss it and wait nearly a full revolution.

Latency distribution

In the uniform model, rotational latency L is uniformly distributed from 0 to T, where T is revolution time. Mean is T/2 and standard deviation T divided by square root 12.

For a 7,200 RPM drive, T = 8.333 ms and standard deviation is about 2.406 ms. A deterministic statement such as “rotational latency is exactly 4.17 ms” confuses an expected value with every observation.

The 95th percentile of a uniform latency is 0.95T, about 7.917 ms. Tail performance is different from average performance, and interactive systems may care about both.

Linear velocity

Tangential surface speed v = omega r = 2 pi r times RPM/60. At r = 40 mm = 0.040 m and 7,200 RPM, v is about 30.16 metres per second. At r = 20 mm, it is about 15.08 m/s.

If linear bit density and electronics permit, outer tracks pass more recorded length under the head per second. This supports higher sequential transfer rates in outer zones.

Angular sector width

If a zone has S sectors per track, each sector occupies an average angular span 360/S degrees, ignoring gaps and overhead. At 1,200 sectors per revolution, average is 0.3 degree per sector.

Time to rotate through one sector is revolution time/S. At 7,200 RPM and 1,200 sectors, it is 8.333 ms/1,200 = about 6.94 microseconds. Reading a multi-sector block requires that time per sector plus encoding and electronics effects.


Seek, Transfer and Queueing

Seek time

Seek time is the interval for the actuator and head-positioning system to reach the desired track and settle sufficiently to read or write. It depends on distance, direction, current motion and control strategy.

A short adjacent-track seek can be much faster than a long stroke across the platter. Therefore one “average seek time” summarises a workload distribution and test method. It is not a constant for every request.

A simple teaching model might use t_seek(d) = a + b square root d for track-distance d, reflecting acceleration and settling behaviour. Actual drive models are proprietary and may have more complex piecewise dynamics.

Transfer time

After positioning, transfer time for B bytes at sustained rate R bytes per second is B/R. Reading 4 MiB at 160 MiB/s takes 0.025 second, or 25 ms, if the transfer is sequential and sustained at that rate.

For a small random 4 KiB block, pure media transfer time may be tiny compared with seek and rotational latency. This is why sequential throughput and random input/output operations per second describe different capabilities.

Average service time

A simplified service time is:

t_service = t_seek + t_rotation + t_transfer + t_controller.

If average seek is 8 ms, average rotation 4.17 ms, controller overhead 0.3 ms and transfer for the request 0.1 ms, mean service time is about 12.57 ms. A rough single-request ceiling is 1/0.01257 = 79.6 operations per second, before queueing and workload optimisation.

This reciprocal is not a guaranteed benchmark. Requests vary, caching occurs, and drive scheduling can reorder work.

Queueing delay

If requests arrive faster than they can be served, they wait. Let arrival rate be lambda requests per second and mean service rate mu. Utilisation rho = lambda/mu in a simple queue. As rho approaches one, expected waiting can rise sharply.

The M/M/1 formula W = 1/(mu – lambda) is only a teaching reference because hard-disk service times are not exponential and scheduling is not necessarily first-come-first-served. Still, it shows why operating near full utilisation can create long latency.

Why Mathematics? | Queueing Theory, Arrival Rates and Waiting Times explains those assumptions in depth. A disk adds physical location, so service order can change seek distance and rotation.

Request scheduling

Serving requests strictly in arrival order may make the head travel back and forth. A scheduler can reorder queued requests to reduce total seek distance or exploit sectors about to rotate under the head.

The elevator or SCAN idea sweeps across cylinder locations in one direction before reversing, analogous to an elevator serving floors. Shortest-seek-first can reduce movement but may delay far-away requests. Fairness, deadlines and write durability complicate optimisation.

Native command queueing lets a drive see multiple outstanding requests and choose an efficient order within protocol constraints. The benefit depends on workload and queue depth. Reordering cannot eliminate physical movement or every latency tail.

Caching

A drive cache may satisfy repeated reads without media access, combine writes or buffer transfers. The operating system also caches data. A measured access time can therefore describe memory rather than platter mechanics unless tests control caching.

Write completion semantics matter. Acknowledging data in volatile cache is different from confirmed durable media storage. Power-loss protection and flush commands are system concerns beyond one timing formula.


Worked Example: A 7,200 RPM Drive

Consider a simplified drive with 7,200 RPM, average seek 8.0 ms, controller overhead 0.25 ms and sequential rate 180 MB/s in decimal units. We compare a 4 KiB random request and an 8 MB sequential request. Assume a cache miss and no queue.

Revolution and average rotation

Revolutions per second = 7,200/60 = 120. Revolution time = 1/120 = 0.008333 s = 8.333 ms. Average rotational latency = 4.167 ms.

Small random transfer

4 KiB = 4,096 bytes. At 180,000,000 bytes per second, ideal transfer time = 4096/180,000,000 = 22.8 microseconds = 0.0228 ms.

Average service time = 8.0 + 4.167 + 0.25 + 0.0228 = 12.440 ms. Reciprocal is about 80.4 requests per second if every request behaved like this and no scheduling or caching changed it.

Positioning accounts for over 98 percent of this simple small-read service time. Doubling sequential transfer rate would barely improve the result because 0.023 ms is already tiny.

Larger sequential transfer

8 MB in decimal units takes 8,000,000/180,000,000 = 0.04444 s = 44.44 ms. If positioning is required first, total is 8.0 + 4.167 + 0.25 + 44.44 = 56.86 ms.

For a long already-contiguous stream, seek and rotation occur once, then many bytes transfer. Throughput approaches the media rate. This explains why workload shape matters more than a single label such as “fast drive.”

Outer-zone comparison

Suppose an outer zone transfers at 200 MB/s and inner zone at 140 MB/s. The same 8 MB media transfer takes 40.0 ms outer and 57.14 ms inner. With the same positioning assumption, totals are about 52.42 and 69.56 ms.

Real seek locations and caching may differ, but circumference and zoning make location-dependent sequential performance plausible.

Queue example

Suppose requests arrive at 60 per second and mean service capacity is 80 per second. In an M/M/1 teaching model, mean total time W = 1/(80 – 60) = 0.05 s, or 50 ms. The service time alone is 1/80 = 12.5 ms; the remaining 37.5 ms is average waiting.

At 40 requests per second, W = 1/(80 – 40) = 25 ms. Halving arrival rate halves utilisation but also reduces total mean time by half in this particular example. Near saturation, small load changes can have large effects.

Do not use this exponential queue formula as a hard-drive benchmark. It shows the nonlinearity of waiting, while an empirical trace or more suitable service model is needed for prediction.

Expected latency versus percentile

With uniform rotation, median and mean rotational latency are both 4.167 ms, but 95th percentile is 7.917 ms. Add fixed 8.25 ms seek-plus-controller and 0.023 ms transfer: median service is about 12.44 ms, while rotation-only 95th service is about 16.19 ms.

If seek time also varies, total 95th percentile cannot be found by simply adding individual 95th percentiles unless dependence and distribution assumptions justify it. Percentiles of sums require convolution, simulation or data.

Results table

QuantityValueMeaning
Revolution time8.333 msOne full turn at 7,200 RPM
Average rotational latency4.167 msHalf-revolution expectation
4 KiB transfer at 180 MB/s0.0228 msMedia time after positioning
Random 4 KiB service estimate12.44 msSeek + rotation + overhead + transfer
Rough reciprocal rate80.4/sNot a guaranteed benchmark
8 MB transfer at 180 MB/s44.44 msSequential payload time

Reliability, Formatting and Error Correction

Location mathematics is not integrity mathematics

A drive must locate data and recover it accurately. Magnetic signals are noisy and media have defects. Error-detecting and error-correcting codes add redundancy so many imperfect readings can still reconstruct intended bits.

This article owns geometry and access-time intent. The coding mathematics belongs to Why Mathematics? | Error-Correcting Codes, Parity and Reliable Data Transfer. Linking the topics is useful; merging them would make each search intent less clear.

Raw and unrecoverable errors

The read channel may experience many raw errors that internal coding corrects invisibly. An unrecoverable read error is what remains after retries and correction limits. A product specification’s rate depends on exact definitions and test conditions.

Do not turn a probability such as one error per 10 to the power n bits into a deterministic promise that the nth bit must fail. It is a long-run rate model. Errors may be correlated, and drive failures are not captured by one bit-error number.

IBM Research reported empirical measurements of disk failure and error rates from large data movement. The study is useful evidence about observed systems, but its hardware, era and workload should not be treated as every modern drive.

Redundancy costs capacity

Sector codes, servo information and spare areas consume physical space. More powerful coding can improve reliability but lowers code rate or requires more signal-processing complexity. The advertised user capacity is therefore not raw magnetic transitions divided by eight.

Backups are a system property

Internal correction does not protect against accidental deletion, malware, theft, fire or controller failure. Redundant drives are not the same as independent backups. Storage reliability is a layered probability and operations problem.


Measurement and Benchmarking

Define workload

A benchmark should state request size, read/write mix, random or sequential pattern, queue depth, test duration, drive fill, zone, cache treatment and temperature. Without these, “MB/s” or “IOPS” is hard to interpret.

Average can hide tails

Two drives can have the same average access time but different 99th-percentile latency. Interactive systems, databases and real-time logging may care about occasional long stalls.

Report a distribution or several percentiles. Also distinguish service time measured at the device from end-to-end application response time.

Warm-up and nonstationarity

Background maintenance, thermal recalibration, error recovery and write caching change performance over time. A short test may capture a burst from cache. A long test may include steady-state behaviour.

Units again

MB/s may mean 10 to the power 6 bytes per second, while MiB/s means 2 to the power 20. IOPS means operations per second, but operation size must be stated. Latency should name mean, median or percentile.

Comparisons with solid-state storage

Solid-state drives have no platter rotation or mechanical seek. They still have controller, flash-programming, garbage-collection, queueing and error-correction delays. The correct lesson is not “SSDs have zero latency,” but that their latency mechanisms differ.


What the Simple Model Leaves Out

Zoned and shingled recording

Zone bit recording varies sectors per track. Shingled magnetic recording overlaps tracks like roof shingles, increasing density but complicating random writes. Drive-managed and host-managed forms have different operational behaviour.

Servo control

Heads fly extremely close to the surface and follow servo information. Track misregistration, vibration and thermal expansion affect positioning. The track is not a thick painted ring.

Multiple heads and surfaces

Only selected heads read at a time in many designs, and switching heads has a cost. Cylinders historically grouped tracks at similar radii across surfaces, but modern logical mapping is more abstract.

Variable seek and rotation correlation

Schedulers may choose a request based on both seek distance and sector arrival. Rotation and seek are then not independent uniform variables. Adding their separate averages can misrepresent an optimised queue.

Failures and recovery

Bad sectors can trigger retries and remapping, creating long latency tails. Averages under healthy conditions do not describe degraded behaviour. Monitoring signals are not perfect predictors.

Filesystems and applications

Files may be fragmented or cached. Read-ahead can make sequential access faster. Database alignment, journalling and operating-system scheduling change request patterns. Drive mechanics are one layer.

Simplified-model table

AssumptionWhat it teachesWhat real analysis adds
Uniform areal densityCapacity scales with area and densityZones, servo and format overhead
Uniform random platter angleMean rotation is half a revolutionScheduling and workload correlation
One average seek timePositioning mattersSeek-distance distribution and control
Constant transfer ratePayload time is size/rateZone, cache and sustained behaviour
No queueDevice service timeArrival process and request reordering
Error-free transferGeometry and timingCoding, retries and failure modes

Misconceptions Worth Correcting

“RPM is the same as transfer rate”

RPM sets rotational period and influences media speed. Transfer rate also depends on linear bit density, zone, electronics and formatting. Two drives with the same RPM can have different throughput.

“Average rotational latency is one revolution”

Under a uniform arrival-angle model, it is half a revolution. One revolution is the approximate worst case after just missing the target.

“A terabyte is always 1,024 gigabytes”

SI terabyte is 1,000,000,000,000 bytes. Binary tebibyte is 1,099,511,627,776 bytes. Labels should be explicit.

“More capacity means faster random access”

Capacity and access time depend on different combinations of density, mechanics, cache and workload. Higher density may help sequential transfer, but random access remains positioning-dominated.

“A block is stored at the LBA number’s obvious physical position”

Logical block addresses are mapped internally. Defect remapping, zones and firmware make physical location opaque.

“Error correction makes backups unnecessary”

No. It addresses certain data errors within limits. Backups address wider loss scenarios and should be independent and tested.


How Students Can Learn This Mathematics

Draw the annulus

Choose inner and outer radii, calculate usable area and compare with the full disc. Change inner radius by ten percent and observe the nonlinear area effect.

Build a paper zone model

Divide a disc into three radial zones. Assign more sectors per track to outer zones. Sum capacity by zone and compare with a constant-sectors model.

Convert RPM three ways

Calculate revolutions per second, seconds per revolution and radians per second. Use units at each step. Verify that 7,200 RPM gives 120 Hz and 8.333 ms per revolution.

Simulate rotational latency

Generate random angles from 0 to 360 degrees and convert each to time. Compare sample mean with half a revolution and plot the uniform distribution.

Compare request sizes

Calculate transfer time for 4 KiB, 1 MiB and 100 MiB at one rate. Add the same positioning time. Identify when transfer begins to dominate.

Model scheduling

Place requested track numbers on a line. Compare first-come-first-served head travel with SCAN and shortest-seek-first. Discuss efficiency and fairness.

Measure carefully

If students use public benchmark data, record workload and units. Avoid dismantling powered drives or handling damaged media. Use safe simulations or retired sealed components for visual inspection.

Disk timing connects circles, rates, distributions, queueing and optimisation. Data integrity connects coding theory. Satellite communications show a different route from physical signal to reliable bits in Why Mathematics? | Satellite Communications, Link Budgets and Signal-to-Noise Ratios.

Keep pathways open

Interested students may explore computing, electronics, data engineering, information security, mechatronics or applied physics. Mathematics supports these pathways but guarantees neither admission nor employment. Check current official programme requirements.

For Singapore cohorts, verify current MOE and SEAB terminology for Posting Groups, Full Subject-Based Banding, G1/G2/G3 subjects and the Singapore-Cambridge Secondary Education Certificate. When relevant, use SEC Additional Mathematics Examination for G2/G3 rather than assuming an older examination label.


Guidance for Parents and Teachers

Start with a spinning circle and a moving arm. Students often understand the physical wait before they understand why average latency is half a turn.

Ask for definitions with every number. Is 8 ms a mean seek, a median access time or a percentile? Is MB decimal or binary? Is capacity raw or user-addressable?

Separate layers. Geometry explains where data can fit. Timing explains when it arrives. Coding explains whether it can be trusted. Filesystems explain how software organises it.

Use estimation. A 7,200 RPM drive cannot have a 20 ms revolution time. A 4 KiB payload at hundreds of MB/s should not take milliseconds to transfer once positioned. Order-of-magnitude checks catch unit errors.

Encourage model comparisons. One average service time is useful, but a latency distribution tells a richer story. One areal-density number is useful, but zones explain outer/inner differences.

Do not turn hardware metrics into consumer guarantees. Product performance and reliability depend on the actual model, workload, firmware, environment and system configuration.


Frequently Asked Questions

What is areal density?

It is stored bits per unit recording surface area. A simple model multiplies linear bit density along tracks by track density across radius.

Why are hard-disk tracks circular?

The platter rotates, so a fixed-radius head path traces a circle. Servo control follows the intended track while the actuator changes radius between tracks.

Why do outer tracks store more sectors?

They have larger circumference. Zone bit recording uses this extra length by assigning more sectors per revolution to outer zones.

What is average rotational latency at 7,200 RPM?

About 4.17 milliseconds under a uniform random-angle assumption, half the 8.33 ms revolution time.

Is seek time the same as access time?

No. Access time can include seek, rotational latency, transfer, controller overhead and queueing. Definitions in specifications or benchmarks should be checked.

Why are small random reads slow?

The payload transfer is tiny, but mechanical positioning still takes milliseconds. Seek and rotation dominate.

Why can sequential transfer be faster on outer tracks?

Outer tracks pass more circumference under the head per revolution and typically contain more sectors in a zone.

Does doubling RPM halve total access time?

It halves revolution time and average rotational latency, but seek, controller, transfer and queueing remain. Total improvement is therefore less than twofold unless rotation dominates.

Why does a drive show less space than its label?

Decimal versus binary units create a major numeric difference; partitions and filesystems add overhead. This is separate from hidden physical redundancy inside the drive.

Can these equations predict a specific drive?

Only when its format, zones, measured timing distributions and controller behaviour are known. The examples teach mechanisms, not product certification.


Useful Next Reading

A hard disk drive makes abstract mathematics wonderfully mechanical. Area determines how much surface is available. Angular speed turns RPM into waiting time. Probability describes where a sector is when a request arrives. Queueing and optimisation decide how requests share the moving head. Once these layers are separated, storage performance becomes a reasoned model rather than a mysterious number on a box.

Discover more from eduKate Singapore

Subscribe now to keep reading and get access to the full archive.

Continue reading