Why is mathematics important to data protection? Removing names is not always enough. A table can omit every name yet still reveal something about a person when its rows are rare, when several releases can be combined, or when an attacker already knows part of the story. Differential privacy approaches the problem differently: it asks how much a published result can change when one person’s data is added or removed, then limits that change with a mathematically defined random mechanism.
This is a powerful example of mathematics in technology, statistics and everyday civic life. Schools, hospitals, companies and public agencies often want to learn from groups without turning every participant into an open book. The challenge is not to choose between “publish everything” and “publish nothing”. It is to quantify a privacy–accuracy trade-off, state the assumptions and spend a limited privacy budget carefully.
The subject connects ratios, inequalities, probability distributions, functions, sensitivity, simulation and error analysis. It also teaches humility. Differential privacy can limit what a release reveals through a defined channel, but it does not repair insecure databases, poor consent, biased samples or careless interpretation. This article is educational, not legal or cybersecurity advice.
Choose the privacy question you want to answer
- Start with neighbouring datasets to understand whose presence the guarantee protects.
- Use sensitivity to measure how much one record can move an answer.
- Study random noise to see why identical queries need not return identical outputs.
- Explore epsilon to understand a formal privacy-loss parameter without treating it as a magic score.
- Follow composition to see why repeated releases consume a budget.
- Compare accuracy with privacy through expected error and confidence-like ranges.
- Examine limits to separate mathematical protection from governance, security and fairness.
- Build skill with synthetic data and simulation, never with private records you are not authorised to use.
The central habit is to define the data unit. Does “one person” mean one row, one household, one device, one trip or all records contributed by one user? The mathematics is only as honest as that boundary.
A privacy question begins with two almost-identical datasets
Imagine a school survey containing one row per student. Dataset A contains 500 rows. Dataset B contains the same rows except that one student’s row is absent. These are neighbouring datasets under an add-or-remove-one-person definition. A different application might define neighbours by replacing one row, and that choice can change sensitivity by a factor of two.
The goal is not to make outputs equal. Useful statistics should still respond to data. The goal is to make the distribution of possible outputs sufficiently similar that an observer cannot confidently infer whether the extra person participated solely from the released answer.
This shift—from protecting a stored value to comparing probability distributions—is the mathematical heart of differential privacy. It is also why one noisy answer should not be judged by asking whether it exactly matches the confidential total.
Differential privacy is a statement about mechanisms
A randomised mechanism M is often described as epsilon-differentially private if, for every neighbouring pair D and D′ and every possible set of outputs S, Pr[M(D)∈S]≤e^ε Pr[M(D′)∈S]. The definition compares probabilities, not the raw tables.
The exponential factor matters. If ε is small, e^ε is close to 1, so the two output distributions must be close. If ε is larger, the permitted ratio is larger. A mechanism may also use an approximate form with a second parameter, delta, allowing a small additive probability term. That extension needs careful interpretation; delta is not simply “chance that privacy fails”.
NIST’s SP 800-226, published in March 2025, describes differential privacy as a framework that quantifies privacy loss and warns that implementation details can undermine a formal-looking claim.
Sensitivity measures one person’s maximum influence
Before adding noise, ask how far a query can move between neighbouring datasets. For a numerical query f, global sensitivity under a chosen norm is the largest possible |f(D)−f(D′)| across all allowed neighbours.
For a simple count—“How many respondents selected option A?”—adding or removing one person changes the count by at most 1. Its sensitivity is 1. For a sum, sensitivity depends on the range of each person’s contribution. If a person may contribute any number without a stated bound, the worst-case change is unbounded, and a standard noise calculation has no finite scale.
This is why clipping is not an irritating preprocessing detail. It defines the maximum individual influence. A question about weekly study time might clip each response to 0–40 hours before calculating a sum.
A bounded sum has visible sensitivity
Suppose each participant contributes a number from 0 to 40 after clipping. Adding or removing one record can change the sum by at most 40, so the add/remove sensitivity is 40 hours. If the permitted interval were 0–100, sensitivity would be 100, requiring more noise for the same epsilon.
Clipping improves the privacy–accuracy trade-off but can introduce bias when genuine values exceed the cap. Choosing 40 because it is convenient is not enough. Analysts should use subject knowledge, public pre-specified limits or a separately protected procedure.
The example shows a recurring principle: privacy depends on controlling influence. Range restrictions, contribution limits and aggregation rules are part of the model. They must be disclosed because they affect both the guarantee and the meaning of the statistic.
The Laplace mechanism links sensitivity to noise
For certain numerical queries, a classic mechanism releases f(D)+Z, where Z follows a Laplace distribution centred at zero with scale b=Δf/ε. Here Δf is sensitivity. Larger sensitivity produces more noise; larger epsilon produces less noise.
If a count has sensitivity 1 and ε=0.5, the Laplace scale is 1/0.5=2. The noise can be positive or negative. Its expected value is zero, but the expected absolute error is b, so about 2 in this simplified example. A particular draw might be 0.4, −3.1 or another value.
The mechanism is not “subtract two people”. It samples from a continuous probability distribution whose tails allow larger errors with decreasing probability. Randomness prevents an observer from reading participation directly from a deterministic change.
A worked count should show the random step
Assume the confidential count is 237, sensitivity is 1 and ε=0.5. With Laplace scale 2, suppose one random draw is −1.6. The raw private output is 237−1.6=235.4. If the publication requires a whole count, a post-processing rule might round this to 235.
Another run could return 239. The privacy claim concerns the distribution over many possible runs, not closeness of every one. Publishing repeated independent runs to “average away the noise” is dangerous because each release composes and can reveal more.
For a classroom simulation, students can generate uniform random numbers and transform them into Laplace noise, then plot thousands of outcomes. They should use a made-up count. The exercise reveals both the central concentration and the nonzero chance of a larger deviation.
Did You Know? Randomisation can protect, not merely confuse
Students often meet randomness as unwanted error: dice wobble, sensors fluctuate, samples differ. Differential privacy uses calibrated randomness deliberately. The noise is not decorative blur. Its distribution is chosen so neighbouring datasets lead to sufficiently similar output probabilities.
Uncalibrated noise can be useless. Adding a random number between −1 and 1 to a count with sensitivity 1 creates outputs that may still identify a boundary event. Adding an unknown amount “because privacy” gives no auditable guarantee. The protection comes from the relationship among adjacency, sensitivity, the distribution and privacy parameters.
This is an optimistic lesson about probability. Randomness can be engineered as a resource. The same mathematics that quantifies uncertainty can limit inference, provided the mechanism is specified and implemented correctly.
Epsilon is a parameter, not a universal grade
It is tempting to say ε=0.1 is “excellent” and ε=10 is “bad”. Context matters. The neighbouring relation, protected unit, number of releases, delta, sampling method, threat model and consequences all influence interpretation. Comparing epsilon values across systems without these details is like comparing speeds without units or distances.
Because e^ε grows nonlinearly, differences are not merely cosmetic. e^0.1≈1.105, e^1≈2.718, and e^3≈20.086. Yet those ratios are components of a formal bound, not direct probabilities that a person is exposed.
NIST advises evaluating the full differential-privacy “pyramid”, including assumptions, implementation and governance. A responsible explanation names epsilon but does not reduce the entire privacy decision to one number.
Delta needs a disciplined explanation
Approximate differential privacy is commonly written (ε,δ)-DP. The probability comparison receives an additive δ term. Choosing delta casually can weaken the guarantee, especially at population scale.
A useful rule is that delta should be much smaller than an inverse-population scale, but the appropriate choice is system-specific and belongs in a documented risk decision. It is misleading to describe delta as a simple probability that “anything can happen” or as a percentage of people who lose privacy.
For students, the important mathematical idea is that inequalities can include both multiplicative and additive allowances. Multiplication controls relative change; addition allows a small absolute slack. Interpreting both requires looking at units, events and magnitudes rather than memorising symbols.
Privacy loss can be viewed as a log ratio
For a particular output y, a privacy-loss quantity can be written as L(y)=ln(Pr[M(D)=y]/Pr[M(D′)=y]) when probabilities or densities are defined appropriately. An epsilon bound limits this log ratio in the pure-DP setting.
Logarithms turn probability ratios into additive quantities. That is helpful for composition: multiplying likelihood ratios corresponds to adding their logarithms. This is one reason privacy accounting feels related to information theory and statistical evidence.
The expression should not be calculated from a single observed output without a specified mechanism. It is a property of distributions over possible outputs. The lesson transfers beyond privacy: whenever a ratio spans many scales, logarithms can turn multiplication into addition and reveal structure.
Composition explains why repeated queries cost privacy
If an ε1-private mechanism and an ε2-private mechanism are run on the same people, a basic sequential-composition bound is ε1+ε2. Ten releases each with ε=0.2 can therefore have a simple bound of 2.0, not 0.2.
More advanced accounting can sometimes give tighter results, especially under approximate definitions or specific mechanisms. But the basic sum teaches the operational point: privacy is a limited resource across releases.
An organisation cannot publish hundreds of independently noised cuts, then claim each table is safe in isolation. A privacy ledger should record who is protected, which data overlap, which mechanism ran, which parameters were used and how the cumulative budget was calculated. This resembles a financial budget because new choices depend on prior spending.
Parallel composition depends on disjoint people
When mechanisms run on disjoint groups and each person contributes to only one group, the privacy cost can sometimes be governed by the largest group-specific epsilon rather than the sum. This is parallel composition.
The disjointness condition is essential. If a student appears in both the sports and music groups, the releases are not parallel for that student. If people can contribute multiple devices or visits, row-level separation may not equal person-level separation.
This is where database design meets set theory. Analysts need to reason about intersections, identifiers and contribution rules. A Venn diagram can reveal why a claimed “separate dataset” is not actually separate at the protected unit. Mathematics prevents a convenient organisational label from being mistaken for a proof.
Post-processing does not spend extra privacy by itself
A fundamental property says that processing a differentially private output without returning to the confidential data cannot worsen its differential-privacy guarantee. Rounding, charting, combining categories or applying a deterministic formula to the released result is post-processing.
Suppose a mechanism releases a private count of 235.4. Rounding to 235, converting it to a percentage using a public denominator, or colouring a chart does not add a new query to the source dataset. However, poor post-processing can still harm usefulness, fairness or credibility. It can also create impossible-looking values.
The phrase “without returning to the confidential data” is crucial. Choosing how to edit an output after secretly checking the true count is data-dependent and may consume privacy. Workflow boundaries matter as much as algebra.
Negative noisy counts are mathematically possible
Laplace or Gaussian noise has unbounded support. A true count of 1 can become −2.7 after noise. A negative number of people is physically impossible, but the random mechanism is operating on a numerical query, not simulating literal people.
One option is to clamp published counts at zero as post-processing. This improves face validity but introduces bias near the boundary because negative errors are pushed upward while positive errors remain. Another option is a mechanism designed for discrete, nonnegative outputs or a consistency optimisation across a table.
The right choice depends on intended use. Students should learn that “fixing” impossible outputs changes statistical properties even when it does not worsen the DP guarantee. Accuracy has more than one dimension.
A histogram query has vector sensitivity
Consider a histogram where each person belongs to exactly one category. Adding one person increases one bin by 1, so the L1 sensitivity of the whole vector is 1 under add/remove adjacency. If replacement adjacency lets one person move from category A to B, one bin decreases and another increases, giving L1 change 2.
This example makes adjacency concrete. The same visible table can have different sensitivity depending on whether neighbours add/remove or substitute.
If people can select three categories, contribution limits change again. A person may affect up to three bins. The database must enforce the promised bound; writing “maximum three” in documentation while accepting unlimited selections breaks the reasoning.
Means require bounded values and a denominator decision
The sample mean is sum/n. Protecting it can be more subtle than protecting a count because both sum and n may be confidential and because an unbounded value can dominate the answer.
One approach clips values, privately releases a sum and privately releases a count, then divides through post-processing. Suppose hours are clipped to 0–40, the private sum is 4,920 and the private count is 198. The released mean is about 24.85 hours.
If the noisy count is small or near zero, the ratio can become unstable. Analysts may restrict publication to sufficiently large groups, use a public known denominator when justified, or choose a purpose-built mechanism. Algebra exposes a real design problem: division magnifies denominator error.
Utility is not just average error
An analyst might report mean absolute error, root-mean-square error, bias or a high-probability bound. Each highlights something different. For Laplace noise with scale b, expected absolute error is b and variance is 2b². The standard deviation is √2 b.
With b=2, expected absolute error is 2 and standard deviation is about 2.83. Those summaries do not promise that every release lies within 2 or 2.83 of truth. The distribution has tails.
For a public-health or school-resource decision, subgroup error may matter more than overall error. A mechanism can have acceptable average accuracy yet poor small-area results. Utility evaluation should mirror the decisions the statistics will support.
A simple tail calculation makes uncertainty visible
For Laplace noise with scale b, Pr(|Z|≥t)=e^(−t/b). If b=2, the probability that absolute noise reaches at least 6 is e^(−3)≈0.0498, about 5 per cent.
Equivalently, about 95 per cent of noise draws fall within roughly ±b ln(20)=±5.99. This is a property of the noise distribution, not automatically a confidence interval for every later derived statistic.
The calculation helps students connect exponentials, inverse functions and risk statements. It also demonstrates why a private output should come with uncertainty information when users need quantitative decisions. Hiding the noise model can make a rounded table look more precise than it is.
Gaussian noise supports approximate guarantees
The Gaussian mechanism adds normal noise, often calibrated to sensitivity, epsilon and delta. Its familiar bell curve makes some calculations convenient, and it appears in modern privacy accounting and machine-learning systems.
The details of calibration matter. One should not copy a standard deviation from a blog without matching the privacy definition and accountant. Different theorems use different constants and assumptions. Software libraries can also implement neighbouring relations or clipping differently.
For classroom understanding, compare 10,000 Laplace and Gaussian draws with the same standard deviation. The histograms show different peaks and tails. The exercise builds distribution literacy without implying that visual similarity proves equal privacy.
Privacy amplification through sampling has conditions
Randomly sampling participants before applying a private mechanism can sometimes strengthen the effective privacy guarantee. Intuitively, an outside observer is uncertain whether a person was sampled at all.
But amplification depends on the sampling scheme: sampling with or without replacement, Poisson sampling and fixed-size sampling have different analyses. If participation is public, or if the sampling step leaks, the intuition can fail. If a person contributes repeatedly, inclusion probability changes.
This topic is a useful warning against slogan-level mathematics. “We sampled, so privacy is amplified” is not a proof. A theorem needs premises, and a deployment needs evidence that the premises match the code and data pipeline.
Privacy budgets belong to questions, not just teams
Imagine a school allocates total ε=1.0 to a survey release. It might spend 0.3 on year-level counts, 0.4 on subject-interest histograms and 0.3 on a wellbeing aggregate. The numbers are illustrative; they do not define a recommended policy.
Changing the allocation changes accuracy. A high-priority statistic may receive more budget, but that leaves less for others. Analysts can optimise allocations against a stated loss function, such as weighted mean-square error.
This is mathematics serving governance. The objective weights are value choices; the resulting optimisation is technical. A transparent process should separate them so a formula does not disguise whose questions were prioritised.
Invariants can improve consistency but leak information
Statistical systems sometimes hold selected totals fixed—an invariant—while adding noise to lower-level cells. If the national population total is treated as public and exact, noisy subgroups can be adjusted to sum to it.
The invariant may improve utility, yet every exact data-dependent fact excludes possible neighbouring datasets and changes what is protected. An invariant cannot be added casually after the proof. It must be part of the mechanism and threat model.
The U.S. Census Bureau explains that disclosure avoidance balances statistical accuracy, data availability and confidentiality, and that it has used small random changes or “noise” since the 1990 Census. Its current disclosure-avoidance page also records ongoing research for the 2030 Census.
Differential privacy is not the same as anonymisation
Traditional anonymisation may remove names, addresses and obvious identifiers. Re-identification can still occur by linking rare combinations with external information. Differential privacy instead bounds how much a mechanism’s output distribution depends on one protected unit.
The approaches are not interchangeable. A private mechanism still needs access controls and careful handling of raw data. Conversely, a de-identified table is not automatically differentially private.
This distinction helps students resist a common misconception: privacy is not a visual property of a dataset. You cannot inspect a few rows and declare it safe. Protection depends on release rules, adversary knowledge, contribution limits, future composition and implementation.
Security and differential privacy solve different problems
Encryption can protect data in transit or at rest. Authentication can restrict who queries a system. Audit logs can record access. Differential privacy limits inference from authorised outputs under a defined model.
If an attacker steals the raw database, a private chart published last week does not make the stolen records safe. If the password is weak, epsilon cannot strengthen it. If staff export unprotected microdata, a privacy budget for the dashboard is irrelevant.
Good systems use layers. Mathematics helps draw the boundary: confidentiality controls access to raw data; differential privacy constrains a release; governance sets legitimate purposes; security engineering defends infrastructure; law and ethics define obligations.
Fairness is not guaranteed by privacy
A mechanism can satisfy differential privacy and still produce biased or harmful results. If the underlying sample excludes a community, calibrated noise does not repair representation. If a model allocates opportunities unfairly, hiding individual influence does not make the decision fair.
Small groups may experience larger relative error because fixed noise is large compared with their counts. Suppressing, combining or reallocating budget involves trade-offs that should be assessed with affected users.
Privacy and fairness can also interact. Publishing more detail may help detect inequity but increase disclosure risk. The responsible response is not to claim that one principle automatically wins; it is to state objectives, measure impacts and build oversight around the mathematical mechanism.
Correlated records complicate interpretation
Differential privacy protects participation under the chosen neighbouring relation even when records are statistically correlated, but it does not stop all inferences about a person that can be made from population patterns.
If twins share genetics or household members share an address, learning about one may rationally update beliefs about another. DP aims to limit the extra inferential effect of including the protected person’s record, not erase knowledge already implied by correlated public facts.
Group privacy bounds can grow with group size. This matters when one person contributes many rows or when the protected unit should be a household. Defining adjacency at row level because it gives better-looking numbers may protect the wrong thing.
Synthetic data is not automatically private
Synthetic data is generated rather than copied row for row. That does not prove privacy. A model can memorise rare training examples or reproduce them with high probability.
Synthetic data can inherit a differential-privacy guarantee if the training process is differentially private and the final data are post-processing of the private model. But the training algorithm, clipping, sampling, accountant and parameters all matter.
For student projects, the safest route is to start with fully invented data. Synthetic should mean “created without real people”, not “generated from confidential records using an unverified tool”. This distinction supports both ethical practice and clean experimentation.
Reproducibility requires controlled randomness
Analysts often set a pseudorandom seed so a simulation can be repeated. In a production privacy mechanism, reusing randomness incorrectly can be dangerous. Publishing the seed may let observers subtract the noise; using correlated noise across queries can expose differences.
Reproducible research can separate a public demonstration from the live release. The demonstration uses synthetic data and a fixed seed. The live mechanism uses a cryptographically appropriate random source and protected operational controls, while preserving auditable code and configuration.
The broader lesson is that “random” has engineering requirements. A distribution on paper is not enough. The software must sample it accurately, handle floating-point behaviour and keep secrets appropriately.
Careers use both mathematics and judgement
Differential privacy appears in official statistics, technology, health research, mobility analysis and machine learning. Relevant roles include statistician, privacy engineer, data scientist, software engineer, security researcher, policy analyst and data steward.
Mathematics supports probability, optimisation, algorithms and error measurement. Computing turns definitions into reliable systems. Domain knowledge chooses meaningful bounds and utility tests. Law and ethics shape legitimate use. Communication helps non-specialists interpret noisy results.
No single mathematics course guarantees entry to these careers. Students can keep options open by building algebra, probability, statistics and programming, then practising how to explain assumptions. A correct inequality that nobody can interpret is not a complete public-data product.
A safe classroom investigation
Create an invented list of 200 binary responses. Compute the true count. Then add Laplace noise at ε values 0.1, 0.5, 1 and 2 with sensitivity 1. Repeat each setting 1,000 times.
For each setting, calculate mean error, mean absolute error and the fraction of results within five of the truth. Plot the distributions. Observe that mean error approaches zero across many trials while individual errors remain. Compare empirical mean absolute error with the theoretical scale b.
Then simulate five releases and add their basic epsilon costs. The investigation demonstrates privacy–accuracy trade-offs, Monte Carlo reasoning and composition without touching private information.
A strong learning progression
At primary level, students can learn that grouped data should not expose classmates and that randomised answers can protect individuals. In lower secondary mathematics, they can use ratios, histograms and simple simulation. Upper-secondary students can add logarithms, probability distributions and bounded functions.
At pre-university or polytechnic level, students can study inequalities, calculus, statistics and programming. University study may add measure-theoretic probability, optimisation, machine learning, databases, cryptography, law or public policy.
The progression is not a race. The transferable habit is to ask what one input can change, what uncertainty a release contains and which assumptions the guarantee needs.
Common misconception: more noise always means better privacy
Noise must be calibrated to sensitivity and a proven mechanism. Huge arbitrary noise may destroy utility without establishing a formal guarantee. Tiny noise may be adequate for a low-sensitivity query under one budget and inadequate for another.
Likewise, privacy cannot be inferred from visual messiness. A chart can look noisy while repeated releases reveal the exact total. A clean-looking output can be private if it comes from a carefully designed mechanism and post-processing.
Students should replace “looks random” with four questions: What are neighbours? What is sensitivity? What distribution is used? What is the cumulative privacy accounting?
Common misconception: the private answer is the true answer with a harmless wobble
Noise can affect rankings, thresholds and small groups. If two categories differ by 1 and noise scale is 3, declaring a definitive winner is not justified. Downstream users need uncertainty and minimum-quality rules.
Private outputs may also undergo consistency adjustments that correlate errors. Treating every cell as independent Laplace noise can be wrong after post-processing.
The responsible stance is neither to dismiss private statistics nor to overtrust them. Use them at the resolution they support, combine them with domain evidence and avoid decisions that hinge on changes smaller than the expected uncertainty.
Questions students should ask before trusting a private release
- What person, household or device is the protected unit?
- Which neighbouring relation is used?
- What contribution bounds or clipping rules apply?
- What are epsilon and delta, and what is the cumulative budget?
- Which mechanism and privacy accountant were used?
- What uncertainty, bias or suppression affects the published statistic?
- Are repeated or overlapping releases accounted for?
- What security, consent and governance protections sit outside the mathematics?
These questions turn privacy from a label into an auditable claim.
A practical guide for parents and teachers
When a student says “the computer added noise”, ask them to explain why, how much and relative to what sensitivity. Encourage them to report ranges and limitations rather than chasing the exact confidential answer.
Parents can connect the topic to familiar school data: class averages, attendance summaries and survey charts. The lesson is not that institutions should hide everything. It is that useful group information can be designed with respect for individuals.
Teachers should use fictional datasets and avoid exercises that invite students to infer sensitive facts about classmates. The mathematics becomes more meaningful when ethical boundaries are part of the problem statement.
Frequently asked questions
Does differential privacy guarantee complete anonymity?
No. It gives a mathematical bound for a specified mechanism, neighbouring relation and parameter set. Raw-data security, consent, governance and other releases still matter. “Anonymous” is too broad to substitute for the actual guarantee.
Is a smaller epsilon always better?
A smaller epsilon usually represents a stronger bound within the same setup, but utility may fall and cross-system comparisons can mislead. Delta, composition, adjacency, contribution bounds and implementation must also be examined.
Can noise be removed by averaging many answers?
Repeated independent answers may average toward the truth, which is precisely why each release must be accounted for. A system should not answer unlimited repetitions as if they were free.
Why not publish only large groups?
Minimum group size can reduce some risks but is not a general proof against linkage or repeated-query attacks. It can be one governance rule inside a broader protection design.
Does differential privacy make data unbiased?
No. Some noise mechanisms have zero mean before post-processing, but sampling bias, clipping bias, nonresponse and model bias remain. Boundary corrections can also create bias.
Can students implement it themselves?
They can simulate basic mechanisms on invented data. Real deployments should use reviewed libraries, qualified experts, documented accounting and organisational controls. A classroom formula is not a production privacy programme.
Useful next reading
- Read the official NIST Guidelines for Evaluating Differential Privacy Guarantees for the current framework and implementation hazards.
- See how the U.S. Census Bureau describes disclosure avoidance and the accuracy–availability–confidentiality trade-off.
- Continue with Why Mathematics? | Survey Sampling, Margin of Error and Weighting to separate sampling uncertainty from privacy noise.
- Compare privacy randomness with information coding in Why Mathematics? | Data Compression, Entropy and Huffman Coding.
The hopeful conclusion
The importance of mathematics here is not that an equation can make every data release safe. It is that mathematics makes a promise testable. It identifies the protected unit, bounds individual influence, calibrates randomness, accumulates repeated costs and quantifies error.
That clarity helps students become better citizens and technologists. They learn to ask what a statistic reveals, what it hides, whose data shaped it and how much uncertainty remains. Differential privacy is one tool among many, but it demonstrates a beautiful principle: with careful definitions, probability can help society learn from groups while limiting what is learned about one person.
