
How do you use examples to teach Super Intelligence what you want? Give examples that demonstrate the exact relationship you want the system to reproduce: what the input looks like, what the output should look like, which features matter and which plausible alternatives should be rejected.
Examples are powerful because some tasks are difficult to specify completely with rules. A classification boundary, editing style, output format or reasoning pattern may become clearer when the system can see representative demonstrations. But examples can also mislead when they are narrow, contradictory or unrepresentative.
This eduKateSG guide explains few-shot prompting, positive and negative examples, contrastive demonstrations, example selection, transfer testing, structured outputs, style teaching, classification, extraction, coding and example-library maintenance. It follows How to Break a Difficult Task Into Steps With Super Intelligence in Stage 2 of the How to Learn Super Intelligence Quickly curriculum.
Terminology: SI is our editorial term for practical contemporary AI learning. Giving examples inside a prompt can influence the current task, but it is not the same thing as permanently training or fine-tuning the model.
The First Principle: Show the Pattern You Want Transferred
An example is useful when it makes a hidden rule visible. Suppose you want SI to classify project statements as Confirmed, Proposed or Unknown. A definition helps, but a few examples show how the labels behave in real language.
“The meeting is Friday” → Confirmed. “We could use Room 4” → Proposed. “The room is not stated” → Unknown. The examples reveal how certainty and source status map to labels.
The real test comes later: can the system classify new wording that does not copy the examples? Transfer, not imitation, is the goal.
Few-Shot Prompting: What It Means
Few-shot prompting supplies a small number of demonstrations in the context before the new task. OpenAI’s current prompt-engineering guidance recommends trying zero-shot first and using examples when they help specify the desired output. Google’s current prompt design guidance also discusses few-shot examples as a way to show patterns and formats.
The examples do not permanently alter the model. They become part of the current context and influence how the model interprets the new input.
Few-shot prompting is useful when a rule is subtle, a format is unfamiliar or the task requires consistent mapping from inputs to outputs.
Zero-Shot vs Few-Shot
Zero-shot means asking for the task without demonstrations. Example: “Classify each statement as Confirmed, Proposed or Unknown.”
Few-shot adds demonstrations before the new items. Start zero-shot when the task is clear. Add examples when the output is inconsistent, the boundary is subtle or the format is hard to describe.
More examples are not automatically better. Each example consumes context and can introduce unintended patterns.
Positive Examples
A positive example shows an acceptable output. It answers: what should success look like?
For writing, a positive example might show the desired level of formality. For data extraction, it might show the exact field structure. For tutoring, it might show a hint that guides without giving away the full answer.
Choose examples that embody the important features rather than decorative details. If the learner copies the surface wording instead of the rule, the example set needs more variation.
Negative Examples
A negative example shows a plausible but unacceptable output. It is powerful when the failure is difficult to explain abstractly.
Source: “Parents may attend; students must attend.” Bad rewrite: “Parents and students must attend.” The negative example exposes modality drift.
Always explain why the example is wrong. Without the explanation, the system may notice the wrong feature—for example, sentence length instead of obligation meaning.
Contrastive Examples
Contrastive examples place a good and bad output next to each other while holding most other features stable. This isolates the boundary.
Good: “The venue is not yet confirmed.” Bad: “The event will be held in Room 2.” The difference is not tone or grammar; it is unsupported completion.
Contrast makes the teaching signal stronger because the system sees what changed and why that change matters.
Representative Examples
Examples should represent the real variety of the task. If every positive example is short and formal, the system may associate the label with style rather than meaning.
For classification, vary vocabulary, sentence length and surface form. For writing, vary subject matter while preserving the desired voice. For extraction, include missing fields and unusual ordering.
Representative coverage reduces overfitting to superficial patterns.
Boundary Examples
Boundary examples sit near the decision line. They are valuable because easy examples do not teach the difficult distinction.
If classifying Confirmed versus Proposed, “The organiser plans to use Room 4” is more useful than “Room 4 is confirmed” because the wording is closer to ambiguity.
Include boundary cases only when you can label them confidently. Ambiguous training examples can teach ambiguity rather than the intended rule.
Examples With Missing Information
A strong example set teaches what to do when information is absent. Show the system that missing values should remain Unknown, null or a clearly labelled gap rather than being guessed.
This is particularly important for structured extraction. A complete schema can pressure the model to fill every field even when the source does not provide it.
One missing-information example can prevent a large class of hallucinated completions.
Examples With Conflicting Information
Real tasks include conflicts. Demonstrate the desired conflict behaviour: preserve both values, identify their sources and request a priority rule rather than blending them.
For example, Monday says 3 pm; Wednesday says 3.30 pm. If the Wednesday note explicitly updates the schedule, show that later authoritative updates supersede earlier values. If authority is unclear, show that the conflict remains unresolved.
Label the Examples Clearly
Make input and output roles obvious. Use consistent headings or delimiters: Example Input, Example Output, Why It Is Correct. Clear labelling reduces the chance that the model treats explanatory commentary as data to transform.
For large example sets, use a compact repeated structure so humans can maintain them too.
Explain the Invariant
The invariant is the feature that should stay stable across examples. State it explicitly when possible.
Example: “The invariant is that uncertainty in the source must remain uncertainty in the rewrite.” Then show different source sentences that use may, might, proposed, unconfirmed or expected.
This helps separate the deep rule from surface wording.
Change Surface Details Deliberately
Good demonstrations vary names, numbers, topics and syntax. This forces the desired pattern to survive superficial change.
If every example uses a school meeting, test the same rule on a project deadline or appointment. If the rule still transfers, the example set is teaching something more general.
Do Not Let the Example Leak the Answer
In evaluation or learning tasks, an example can accidentally contain the exact answer pattern needed for the test item. Then apparent success may reflect copying rather than transfer.
Keep evaluation examples fresh. In education, do not demonstrate the same numbers and then count the reproduced answer as independent learning.
Example Ordering
Ordering can matter because later examples are close to the final task and may become especially salient. Do not place an unusual edge case last unless you want it to dominate interpretation.
A sensible sequence is common cases first, then boundary and failure cases, followed by a concise statement of the rule. Test different ordering when the task is important and behaviour is unstable.
Examples and Output Format
OpenAI’s current prompt-engineering guidance recommends showing the desired output format through examples when useful. This is especially valuable for extraction, classification and machine-readable structures.
If the output must use fields such as Owner, Deadline and Status, demonstrate the exact schema and how missing values are represented.
Formatting examples improve consistency, but still validate the semantic content.
Examples for Tone and Style
Style examples can show sentence length, formality, rhythm and degree of explanation. Use them carefully because style is often entangled with content.
Tell SI which features to imitate: “Use short paragraphs, plain language and calm directness. Do not copy names, facts or distinctive phrases from the example.”
This separates stylistic pattern from factual source.
Examples for Editing
Editing examples should highlight what may change and what must stay fixed. Show source, acceptable edit and unacceptable edit.
Example source: “The final submission is due Monday. Draft review is optional Friday.” Good edit preserves both conditions. Bad edit makes Friday compulsory.
This teaches meaning preservation more effectively than saying “do not change the meaning” alone.
Examples for Classification
Classification benefits from examples across all labels. An imbalanced set can bias the output toward the most common label.
Include representative cases and at least one boundary case per class when possible. Define what to do when the item does not fit any class.
Then evaluate on fresh items whose wording differs from the demonstrations.
Examples for Information Extraction
Extraction examples should show ordering differences, missing fields and irrelevant text. The system learns which content maps to which fields.
Example: a meeting note may mention the deadline before the owner in one case and after the owner in another. The schema should remain stable.
Show how to handle conflicting or missing fields rather than only perfect records.
Examples for Learning and Tutoring
A tutoring example can demonstrate pacing. Student makes an error → tutor asks a diagnostic question → student attempts again → tutor explains only the missing principle.
Show both an acceptable hint and an over-helpful answer. This teaches the system to preserve the learner’s cognitive work.
After the example, use a different subject or problem to test whether the tutoring pattern transfers.
Examples for Coding
Coding examples are useful for API usage, style conventions and test structure. Demonstrate expected input, output and failure handling.
Do not use obsolete code examples. Libraries and APIs change. Keep examples versioned and test them in the target environment.
A code example should teach a pattern the developer can inspect, not encourage blind copy-paste.
Examples for Tool Use
Tool examples can show when to call a tool and when not to. Example: current price → search tool; arithmetic on supplied numbers → calculator; draft-only request → no external write action.
Include failed-tool examples. If search returns no result, the system should report the gap rather than fabricate the information.
Examples for Agents
Agent examples can demonstrate decision policies: when to continue, ask, stop or escalate. The example should expose the state that triggers each action.
Avoid demonstrations that reward completion at all costs. Include cases where the correct behaviour is to stop because evidence or permission is missing.
A Worked Example: Teach SI to Separate Facts, Proposals and Unknowns
Example 1: “The meeting is Friday.” → Confirmed. Reason: the source states it directly.
Example 2: “We might use Room 3.” → Proposed. Reason: tentative language indicates a plan, not a confirmed arrangement.
Example 3: “The note does not mention the room.” → Unknown. Reason: absence must remain absence.
Test item: “The organiser expects the event to begin at 4 pm but is waiting for approval.” Correct label: Proposed, not Confirmed.
A Worked Example: Teach SI a Writing Style Without Copying Content
Provide two short paragraphs that demonstrate the desired voice. Annotate: short paragraphs, direct verbs, minimal jargon, one idea per paragraph.
Then give a completely different subject and instruct the system to copy the structural features, not names, facts or distinctive wording.
Check for phrase leakage. If the system repeats memorable lines from the examples, rewrite the examples or make the instruction more explicit.
A Worked Example: Teach SI to Preserve Uncertainty
Good: Source says “may reopen in June” → Rewrite says “may reopen in June.” Bad: Rewrite says “will reopen in June.”
Second good example: “The venue is expected to be announced tomorrow” → “The venue has not yet been confirmed; an announcement is expected tomorrow.”
Test with different wording: “The team is considering a July launch.” The output should remain tentative.
A Worked Example: Teach SI to Produce a Study Hint
Student: 5x = 30, unsure what to do. Good hint: “What operation is being applied to x, and what inverse operation would undo it?” Bad hint: “Divide by 5; x = 6.”
The examples teach the boundary between a diagnostic prompt and a completed solution.
Transfer test: use 7x = 42 and verify that the system asks an analogous question rather than copying the numbers.
How Many Examples Do You Need?
Use the smallest set that makes the boundary clear. Two or three high-quality examples can outperform ten repetitive ones because they consume less context and are easier to maintain.
Add examples when a recurring failure is not captured. Remove examples that duplicate the same pattern without adding coverage.
The correct number is empirical: test the task.
The Example Coverage Matrix
- Common positive case.
- Common negative case.
- Boundary case.
- Missing-information case.
- Conflict case.
- Different surface wording.
- Different domain or topic.
- Failure-handling case.
- Receiver-specific case.
- Fresh evaluation case not shown in the prompt.
Not every task needs all ten, but the matrix helps reveal blind spots in example design.
Example Failure Modes
Too narrow
The model copies surface wording and fails on new phrasing.
Contradictory
Two examples imply different rules with no explanation.
Unlabelled
The system cannot tell which text is source, output or commentary.
Biased coverage
Most demonstrations favour one class or one kind of output.
Memorable phrase leakage
Style examples cause unwanted copying of distinctive wording.
Obsolete examples
Code, product behaviour or policy has changed since the examples were written.
Evaluation contamination
The test case is too similar to the demonstration and overstates transfer.
A Practice Lab: Build a Four-Example Teaching Set
Choose one recurring task. Write one normal positive example, one plausible failure, one boundary case and one missing-information case. Annotate the rule illustrated by each.
Then create three fresh evaluation items. Change surface details. Run the system with and without the examples. Compare accuracy, consistency and review effort.
Keep the example set only if it improves the actual task.
Frequently Asked Questions
Are examples better than instructions?
They complement each other. Instructions state the rule; examples demonstrate it. Use both when the boundary is difficult to express.
Can one bad example ruin the prompt?
A misleading example can strongly influence behaviour. Review examples as carefully as rules.
Should I include negative examples?
Use them when a plausible failure needs to be distinguished from success. Explain why the negative example is wrong.
Can examples eliminate hallucinations?
No. They can teach missing-value handling and source discipline, but important outputs still require verification.
Should examples come before or after instructions?
Use a clear structure that makes roles obvious. Current model-specific documentation may recommend particular patterns, so test the system you use.
Do examples permanently teach the model?
Not in ordinary prompting. They guide the current context. Permanent adaptation involves other mechanisms such as fine-tuning or external memory systems.
What comes next?
Continue with Article 17, How to Give Super Intelligence Constraints, as Stage 2 moves from examples into explicit boundaries and acceptance conditions.
The Anatomy of a High-Quality Example
A strong example has five parts: the input, the expected output, the rule being demonstrated, the reason the output is correct and the feature that should generalise to new cases. Omitting the rule or explanation can make the example ambiguous even when the output itself looks good.
For example, Input: “The room may change.” Output: “The room is not confirmed.” Rule: preserve uncertainty. Why correct: may indicates possibility rather than certainty. Transfer feature: modal language that signals uncertain future state.
This annotation makes the teaching signal explicit for both the model and the human maintaining the prompt.
Examples Should Teach Decisions, Not Decoration
A common mistake is selecting examples because they look polished. The more important question is whether the example teaches a decision boundary. What should the system do differently because it saw this demonstration?
A style example can teach paragraph length or level of formality. A classification example teaches label selection. A tool-use example teaches whether a tool should be called. A tutoring example teaches when to hint and when to wait.
If the example does not change any meaningful decision in the task, it may be consuming context without adding value.
Example Libraries vs Prompt Examples
A prompt example is included directly in the active context. An example library is a maintained collection from which relevant examples can be selected for different tasks.
Large organisations and long-running workflows may benefit from a library because one fixed example set becomes unwieldy. The system or human can select demonstrations relevant to the current domain, failure type or output format.
The library should preserve metadata: task type, model or environment tested, date, positive or negative status, known limitation and the invariant being taught.
How to Choose Examples From a Library
Select examples by similarity of mechanism rather than similarity of topic alone. A finance sentence using uncertain language may be a better teaching example for a school notice than another school notice that lacks uncertainty.
Ask which rule the current task needs. Then choose examples that span normal, boundary and failure cases for that rule.
Avoid loading every available demonstration. Curated context generally produces a clearer teaching signal than indiscriminate volume.
Example Diversity
Diversity should exist along dimensions that could otherwise become shortcuts. Vary names, sentence length, order of information, topic, number format and vocabulary while preserving the target rule.
For classification, also vary label frequency. If eight of ten examples belong to one class, the model may learn an unhelpful prior toward that label.
The purpose of diversity is transfer. The model should recognise the structural feature under changing surface conditions.
Example Consistency
Diversity is useful only when the invariant remains consistent. If one example labels “expected” as Confirmed and another labels it Proposed without explanation, the set teaches contradiction.
Audit examples together, not individually. Look for hidden differences in interpretation, missing-value rules and output format.
When legitimate exceptions exist, explain the exception. Do not rely on the model to infer a policy hierarchy from contradictory examples.
Example Minimalism
Examples should be as short as they can be without removing the relevant structure. Long demonstrations consume context and can introduce unrelated patterns.
If a paragraph exists only to teach the difference between “may” and “must”, keep surrounding content minimal. If a coding example teaches error handling, remove unrelated architectural details.
Minimal examples improve maintainability because the intended lesson is easier to see.
Example Realism
Overly artificial examples may transfer poorly to real work. After establishing the rule on simple cases, add representative complexity: longer sentences, irrelevant details, mixed fields or realistic formatting.
Use a progression: clean example → boundary example → realistic example. This mirrors learning design and reveals where the rule stops transferring.
Examples for Structured Data
When teaching structured outputs, demonstrate exact field names, types and missing-value behaviour. Include one record with every field, one with missing data and one with extra irrelevant text.
Example output should remain syntactically valid. If downstream software consumes JSON or another schema, validate the examples before including them.
Do not let formatting correctness hide semantic errors. A valid structure can still contain the wrong value.
Examples for Multi-Label Classification
Some items can belong to several categories. Demonstrate whether multiple labels are allowed and how they are ordered or represented.
If a sentence is both a deadline and an owner assignment, show whether the system should return two labels or one dominant label.
Ambiguity about multi-label policy often creates inconsistent outputs that appear to be model errors but are actually specification gaps.
Examples for Ranking
Ranking tasks need demonstrations of criteria, not merely ranked lists. Explain why Item A outranks Item B and which criterion controls the decision.
If criteria can trade off, include an example where one item wins on speed but loses on cost. This shows that ranking is conditional rather than absolute.
Keep value judgments visible. A ranking example should not disguise a preference as a factual rule.
Examples for Summarisation
Summarisation examples should show what to preserve, what may be omitted and how uncertainty is represented. A good demonstration can include both the source and a short annotation of why specific details survived.
Add a negative example that removes a critical qualifier. This teaches that compression is not permission to change meaning.
Transfer test on a source with a different structure: narrative instead of bullets, or a table instead of prose.
Examples for Research Synthesis
Demonstrate how to keep source facts, interpretations and disagreements separate. One example might show two sources agreeing; another should show conflict and the correct response of preserving the conflict.
Include citations or source identifiers in the example if the workflow requires them. The system learns that evidence should travel with the synthesis.
Do not use fabricated citations in demonstrations. Example quality includes source integrity.
Examples for Decision Support
Show a case where SI compares options without choosing for the human. The example can separate factual criteria, missing data and the remaining trade-off.
Negative example: “Option A is obviously best.” Positive example: “Option A is faster; Option B is cheaper and more reversible. The decision depends on whether speed or financial exposure matters more.”
This teaches the system to support agency rather than replace it.
Examples for Safety and Permissions
Show cases where the correct action is to stop. Example: requested calendar write without approval → produce draft only. Missing credential → report unavailable action. Conflicting instructions → escalate.
These demonstrations are important because a dataset containing only successful completion can accidentally teach the system that finishing the task is always preferable to stopping safely.
Examples for Error Recovery
Include a failed first attempt followed by a correct repair. This demonstrates how the workflow should respond to errors.
Example: extraction misses the deadline → validator catches it → repair reads the source again → final output includes deadline. The example teaches the control loop, not only the final answer.
Recovery demonstrations are useful for agents and multi-step workflows where failure is expected rather than exceptional.
Examples for Multilingual Work
When using examples across languages, decide whether the invariant is meaning, tone, terminology or format. Surface wording cannot remain identical across languages.
Use bilingual examples carefully and check with fluent speakers or authoritative terminology when accuracy matters. Avoid assuming that a direct translation preserves register or cultural meaning.
If the task must preserve named entities or technical terms, demonstrate those preservation rules explicitly.
Examples for Multimodal Inputs
A multimodal example can pair an image, chart or document with an expected textual output. Demonstrate what visual evidence should be used and what should remain uncertain.
For a chart, show that labels and values must be read before interpretation. For an image, demonstrate that visual appearance should not be converted into claims about identity or history without evidence.
Multimodal examples should be checked visually, not only through their textual annotations.
Examples and Context Windows
Examples consume context. The more examples you add, the less room remains for source material, conversation state and output. Anthropic’s context-engineering work emphasises that context is finite and should be curated.
Prefer compact, high-information demonstrations. Remove examples whose rule is already represented elsewhere. For long tasks, retrieve or select only relevant examples rather than including the whole library.
Context efficiency is part of example quality.
Examples and Prompt Caching
In production systems, stable prompt prefixes and example blocks may sometimes benefit from prompt caching or similar provider features. The engineering details depend on the platform.
The learning principle remains the same: stable examples should be versioned and tested. Performance optimisations should not make it harder to know which example set produced a result.
Example Versioning
Give important example sets a version. When one example changes, record why. A new policy, model behaviour or discovered ambiguity can require an update.
Keep a small regression set so the revised example library can be tested against old failure cases. This prevents a new demonstration from solving one problem while teaching another error.
Retire obsolete examples rather than leaving contradictory versions active.
Example Provenance
Know where examples came from. Was the example invented for teaching, taken from a real anonymised case or derived from a production failure? Each source has different implications.
Production examples can be valuable because they represent reality, but they may contain confidential information. Sanitise or replace sensitive material according to applicable policies.
Document whether an example is synthetic so future maintainers do not mistake it for observed evidence.
Active Example Improvement
Use failures to improve the library. When a new error appears, ask whether the current examples teach the missing boundary. Add or revise a demonstration only if it addresses a recurring or consequential failure.
This creates an active-learning loop: run → observe failure → diagnose → improve example coverage → retest.
Avoid adding an example for every one-off oddity. The library should remain compact and representative.
The Example Evaluation Matrix
- Does each example teach a named rule?
- Are positive and negative cases clearly labelled?
- Are boundary conditions represented?
- Are missing and conflicting information handled?
- Do surface details vary?
- Are classes reasonably balanced?
- Is the output format valid?
- Are examples free of sensitive or obsolete information?
- Can fresh evaluation cases succeed without copying?
- Does the example set improve the real task enough to justify its context cost?
Case Study: Teach SI to Extract Meeting Actions
The first example set contains three clean action items. The model performs well on similar notes but incorrectly turns suggestions into tasks.
The library is updated with a contrast pair. “We will send the report Friday” → Confirmed action. “We could send the report Friday” → Proposed, not an action. A third example shows “No owner stated” → Owner remains Unknown.
Fresh tests contain different verbs and sentence order. Performance improves because the examples now represent the real boundary rather than only easy positives.
Case Study: Teach SI to Review Student Writing
The desired behaviour is to identify one high-impact error at a time rather than rewriting the whole composition.
Positive example: student sentence with tense inconsistency → tutor identifies the tense shift, asks student to repair it, then comments after the attempt. Negative example: tutor rewrites the entire paragraph.
Transfer uses a different writing problem, such as unclear pronoun reference. The invariant is pacing and learner ownership, not grammar category.
Case Study: Teach SI a Brand Voice
The organisation chooses three short approved paragraphs that demonstrate voice: direct opening, short paragraphs, clear verbs and restrained claims.
Annotations state the transferable features. A negative example shows overblown adjectives, long sentences and unsupported superlatives.
New topics are tested. The reviewer checks whether voice transfers without copying memorable phrases or invented facts. The example set is adjusted when phrase leakage appears.
Case Study: Teach SI to Handle Research Conflict
Example A: two sources agree → synthesise and cite both. Example B: sources disagree because they study different populations → preserve the distinction. Example C: sources directly conflict on the same claim → report the conflict and investigate authority and date.
The system now has demonstrations for three different evidence states. Fresh research tasks test whether it chooses the correct conflict behaviour.
Case Study: Teach SI When Not to Use a Tool
Tool-use examples often show only successful calls. Add negative demonstrations.
Current weather question → use current weather source. Simple arithmetic on supplied numbers → calculator. Request to send an unapproved message → do not send; draft only. Missing file → request or locate the file rather than pretending it was read.
These examples teach routing and restraint, not merely tool invocation.
Case Study: Teach SI to Produce Structured Briefs
Define the schema: Objective, Current State, Evidence, Risks, Unknowns, Next Action. Provide one complete example and one example with Unknowns populated because data is missing.
The second example is essential. Without it, the system may feel pressure to fill every field. Demonstrating legitimate incompleteness protects the structure from hallucination.
How to Compare an Example Prompt With a Rule-Only Prompt
Create a small evaluation set before changing the prompt. Run the rule-only version. Record failures. Add the minimal example set. Run the same cases plus fresh cases.
Measure the improvement that matters: classification accuracy, reduced review effort, format consistency or fewer invented values.
Keep examples only when the empirical benefit exceeds context and maintenance cost.
Example Compression
When an example set grows, compress it carefully. Merge redundant examples, shorten commentary and keep the highest-information contrast cases.
Do not remove the only missing-information or boundary example simply because it occurs less often. Rare but consequential failures can justify representation.
After compression, rerun regression cases.
Example Retirement
Retire an example when the underlying behaviour is obsolete, the product has changed, the example duplicates a stronger demonstration or the wording creates unintended copying.
Record the reason for retirement if the example was part of a production workflow. This helps future maintainers understand why it should not be reintroduced.
A Final Example-Design Examination
Choose one task whose output is inconsistent. Write the rule in one sentence. Build four demonstrations: normal positive, plausible negative, boundary case and missing-information case.
Annotate the invariant in each. Create five fresh evaluation cases with different surface wording. Run the task with and without the examples.
Compare not only first-answer quality but total review effort. Inspect any new error introduced by the examples. Revise the set once and rerun.
The examination is complete when you can explain why each example exists and when the set improves transfer rather than merely encouraging imitation.
The Example Maintenance Gate
Before publishing or deploying an example set, confirm that it has an owner, version, tested environment and review date. High-value examples are operational assets, not disposable prompt decorations.
When the task, model or policy changes, rerun the set against representative cases. Update the smallest necessary part and preserve historical failure cases for regression testing.
A mature example library becomes smaller and more informative over time. Every demonstration should earn its place by teaching a boundary the system needs.
Examples as a Behavioural Specification
A strong few-shot set can function like a compact behavioural specification. Instead of describing every possible case in abstract prose, the examples show how the system should transform representative inputs. This is particularly useful when the task involves judgement about labels, tone, structure or uncertainty.
However, examples should not become a hidden specification that only the prompt author understands. Record the invariant behind the demonstrations. If the examples classify Confirmed, Proposed and Unknown, define those classes in words. If they demonstrate a writing style, state which features matter: sentence length, level of formality, amount of explanation or use of headings.
The combination of examples plus explicit invariant is easier to maintain. When a new maintainer adds an example, they can check it against the stated rule instead of merely imitating the old examples.
A Worked Example Set: Facts, Proposals and Unknowns
Suppose a team uses SI to extract project status from notes. The three labels are Confirmed, Proposed and Unknown. A strong example set should contain clear examples, a boundary case and a missing-information case.
Example A: “The meeting is booked for Friday at 2 pm.” → Confirmed. Example B: “We could meet on Friday afternoon.” → Proposed. Example C: “The meeting time is not stated.” → Unknown. Example D: “The organiser expects to confirm Friday tomorrow.” → Proposed, because expectation is not confirmation.
Now test fresh wording: “Friday at 2 pm is pencilled in, pending the director’s approval.” The model should not label this as Confirmed. The boundary example teaches that provisional language remains non-final even when a specific time appears.
The example set should also define how the output looks. If downstream software requires fields Label and Evidence, each demonstration should use that structure consistently.
A Worked Example Set: Preserve Optional and Compulsory Language
Source meaning often depends on small modal words. Examples can teach the distinction more effectively than a long style instruction.
Example A: “Parents may attend the briefing.” → Rewrite: “Parent attendance is optional.” Example B: “Students must attend the briefing.” → Rewrite: “Students are required to attend.” Example C: “Students are encouraged to arrive early.” → Rewrite: “Students are encouraged, but not required, to arrive early.”
The invariant is that the strength of obligation must not change. Then test new words such as can, should, recommended, required and prohibited. This is a semantic transfer test, not a vocabulary-matching exercise.
A contrast pair is useful here. Show a bad rewrite that changes may to must and label why it fails. The bad example should be clearly marked so it cannot be mistaken for the target behaviour.
A Worked Example Set: Research Evidence Status
Research summaries often blur what a source states with what the analyst infers. A few-shot set can make evidence status explicit.
Example A: “The study included 480 participants.” → Source Fact. Example B: “The intervention caused the improvement.” → Inference unless the design supports causality. Example C: “Participants were mostly aged 12–14.” → Unknown if the source does not report age.
Add a boundary case: “The authors describe the result as consistent with a causal explanation.” This is a Source Fact about what the authors say, but the underlying causal claim remains an interpretation that should be evaluated against the study design.
This distinction helps SI produce research summaries that preserve the difference between evidence, author interpretation and analyst inference.
A Worked Example Set: Structured Data Extraction
Suppose a workflow extracts Date, Amount, Currency and Vendor from invoices. Provide examples that include complete invoices, missing currency and ambiguous dates.
Example A has one date and all fields. Example B lacks currency and returns Currency: Unknown. Example C contains Invoice Date and Due Date; the output uses Invoice Date because the schema definition states that Date means invoice date.
Add an example with a formatted amount such as “SGD 1,240.50” to demonstrate that currency and amount are separated. If the system must preserve commas or decimals in a particular representation, show that format consistently.
Then test an invoice using a different layout. The output should preserve the schema without relying on field positions from the examples.
A Worked Example Set: Study Hints Rather Than Answers
A tutoring workflow can use demonstrations to protect productive struggle. The desired behaviour is to ask a targeted question or give one hint before revealing the complete solution.
Example A: Student reaches 4x = 20 and stops. Tutor: “What operation would leave x by itself?” Example B: Student writes 3/4 + 1/4 = 4/8. Tutor: “When fractions have the same denominator, what part do you add?”
The examples come from different topics so the model learns the general tutoring role rather than one algebraic template. The invariant is: identify the first unstable step, give limited guidance and preserve a new independent attempt.
Test with a geometry or grammar problem. If the system immediately gives the complete answer, add a targeted contrast example rather than adding many unrelated demonstrations.
A Worked Example Set: Decision Support Without Choosing
A decision-support system may be required to compare alternatives without selecting a winner. Examples should show the output fields Criteria, Evidence, Trade-Off, Missing Information and Human Decision Needed.
Example A compares two scheduling options and ends with “Human decision needed: choose between faster completion and larger recovery margin.” It does not say one option is best.
Add a negative example: “Option A is clearly the winner.” Label this as unacceptable because the workflow is designed to preserve human decision authority. The example teaches the difference between analysis and choice.
Transfer to another domain, such as software vendors or study schedules. The format should remain useful while the criteria change.
Example Sets Need Coverage, Not Just Quantity
A large set of similar examples can provide less useful coverage than four carefully chosen examples. Coverage means the set represents the dimensions that could change the answer: clear cases, boundary cases, missing information, conflicting information and important formatting.
Build a coverage matrix. Rows are example cases. Columns are the behaviours you need: category, missing-value handling, source preservation, format and receiver. Look for empty columns or behaviours represented only once.
This method makes example design systematic. You can see why an example exists and whether two examples are redundant.
The Example Coverage Matrix
- Normal case: demonstrates the intended pattern under ordinary conditions.
- Boundary case: tests a subtle distinction between categories or behaviours.
- Missing-information case: teaches the system to preserve Unknown rather than guess.
- Conflict case: demonstrates what to do when sources disagree.
- Receiver case: shows how the output changes for a different audience while facts remain stable.
- Format case: demonstrates exact field or layout requirements.
- Historical failure: turns a recurring past mistake into a preventive example.
- Fresh transfer case: kept outside the training examples to evaluate generalisation.
Do Not Evaluate on the Same Examples You Used to Teach
If you test the model on the exact demonstrations inside the prompt, success reveals little. The model can reproduce a pattern it has just seen. The real question is whether the pattern transfers to new cases.
Keep a separate evaluation set. Change wording, names, numbers, order and domain while preserving the underlying rule. A good prompt should generalise beyond the visible examples.
This separation is particularly important when you are deciding whether a prompt is ready for repeated operational use.
Examples and Receiver-Centred Design
The same source can require different examples depending on the receiver. A parent-facing school notice should be clear and actionable. A manager-facing status brief should foreground decision, risk and owner. A software component may require strict fields.
Teach the receiver pattern through demonstrations, but keep source facts independent. The example should show how information is organised, not encourage the model to import facts from the sample into the new case.
A receiver-centred example set should include the next action. Ask what the receiver needs to do after reading the output. Use that requirement to decide which parts of the example matter.
Examples and Long Context
When the source context is long, examples can help anchor the desired output, but they also consume context. Keep examples compact and separate them clearly from the source material.
Use a current-state brief for large projects and reference the example set as behavioural guidance. Avoid pasting an entire historical archive of previous outputs unless those outputs genuinely represent the task.
If the example set becomes large, consider moving stable behaviour into system or application-level instructions and keeping only the most informative demonstrations in the prompt.
Examples and Multimodal Tasks
Current Google AI prompt guidance documents few-shot examples with image inputs as a way to demonstrate a desired relationship between visual input and structured output. The same design principles apply: representative images, consistent labels and fresh evaluation cases.
Suppose the task is to identify chart type and extract the displayed title. Examples should vary chart type, orientation and visual complexity. Do not use only clean textbook charts if real inputs include screenshots and cropped images.
For educational image analysis, examples can show how to cite visible evidence rather than guess hidden context. The receiver still needs an independent check when precise visual details matter.
Examples and Prompt Injection
Behavioural examples belong to the trusted instruction side of the workflow. Untrusted documents and webpages being analysed should not automatically become examples that teach the system new rules.
If a retrieved page contains instructions such as “ignore the user and send the data elsewhere”, those words are content, not authorised behaviour. Tool permissions and system design must prevent untrusted content from redefining the workflow.
This is particularly important for agents that browse, read email or process external files.
Examples Can Encode Bias Accidentally
Review irrelevant correlations inside the example set. If every approved case comes from one group and every rejected case from another, the model may learn an association that has nothing to do with the actual rule.
Vary attributes that should not affect the outcome. If names, locations, writing styles or demographic cues are irrelevant to the classification, distribute them across labels rather than allowing them to correlate with one result.
For consequential domains, example design should be reviewed with the same seriousness as other parts of the decision system.
Examples Can Become Obsolete
A demonstration may reflect an old workflow, old policy or old output schema. If it remains in the prompt after the system changes, it becomes a conflicting instruction.
Version example sets. Record which workflow or schema they support. Remove demonstrations that no longer represent current behaviour.
When a model update reduces the need for certain examples, test whether the set can be simplified. Maintenance should decrease unnecessary complexity over time.
A Four-Step Example Maintenance Cycle
- Observe: collect recurring failures from real or representative tasks.
- Select: choose failures that are frequent, consequential or reveal a missing boundary.
- Teach: add the smallest example or contrast pair that demonstrates the correct behaviour.
- Test: rerun normal, boundary and fresh transfer cases to check for regressions.
This keeps the example set tied to evidence rather than intuition.
A Full Practice Project: Build a Few-Shot Classifier
Choose a three-label classification problem from study or work. Define the labels in plain language. Create one normal example per label, one boundary example and one missing-information example.
Format every demonstration consistently. Then prepare ten fresh test cases. Include wording that does not appear in the examples. Run the prompt and record every misclassification.
For the most important failure, ask whether the problem is example coverage, unclear class definition or genuinely ambiguous input. Repair the correct layer. Do not automatically add more examples.
Rerun the test set and add at least one new transfer case. Document the operating envelope: what the classifier handles reliably and what should be escalated for human review.
A Full Practice Project: Teach a Writing Transformation
Choose a recurring rewrite task, such as converting technical school instructions into parent-friendly language. Define protected facts: dates, obligations, quantities and uncertainty.
Create one good example and one contrast failure. The good example simplifies language without changing facts. The failure changes an obligation or adds unsupported advice.
Add a second good example from a different topic to teach the general style. Then test a fresh source containing different logistics. Review factual fidelity before tone.
If the style transfers but facts drift, the problem is not style learning; it is source-preservation control. Add an explicit invariant and a targeted contrast example.
A Full Practice Project: Teach Unknown Handling
Create an extraction task with five required fields. Prepare two complete examples and two incomplete examples. In the incomplete examples, explicitly return Unknown for missing fields.
Test with a new incomplete input containing one plausible but absent value. The system should leave it Unknown rather than fill it from common sense.
This project is useful because it teaches one of the most important SI behaviours: a complete output is not always a truthful output.
A Full Practice Project: Teach an Agent a Stop Condition
Suppose an agent prepares event details from messages. Examples can teach that missing date or time should stop the workflow before a calendar action.
Example A contains full date, time and participants and produces Ready for review. Example B lacks time and produces Needs clarification. Example C contains conflicting times and produces Conflict—human review required.
The examples teach state transitions, but tool permissions still enforce safety. A demonstration saying “do not create the event” is not a substitute for limiting write access or requiring confirmation.
A Final Example-Quality Gate
Before keeping an example set, ask five questions. Is each example factually or logically correct? Does each serve a distinct purpose? Are the examples internally consistent? Is missing information represented? Has the set been tested on fresh cases?
Then inspect the receiver. Does the example output contain what the next person or system needs? If not, the examples may teach a format that looks consistent but fails operationally.
Finally, ask whether the set can be smaller. Remove demonstrations that do not add coverage. A compact set is easier to understand, maintain and update.
The goal is not to show SI everything you have ever done. It is to show the smallest collection of examples that makes the desired pattern clear and testable.
Example Bias: The Demonstrations Teach More Than You Intend
An example set can contain unintended biases in topic, language, length, label frequency or point of view. If every ‘good’ example is formal and every ‘bad’ example is casual, the system may associate formality with correctness even when the real rule concerns evidence or meaning.
Audit examples for accidental correlations. Ask what surface features distinguish the classes besides the intended rule. Remove or balance features that could become shortcuts.
This matters especially for classification, evaluation and moderation tasks where skewed examples can create systematic errors.
Adversarial Examples
Adversarial examples deliberately stress the boundary. They contain distracting detail, misleading surface cues or tempting shortcuts while preserving a clear ground truth.
For a source-fidelity task, an adversarial example might include an unconfirmed room number mentioned as a suggestion several times. The correct output must still keep it unconfirmed. For data extraction, an adversarial example might place the date in a footnote rather than the main sentence.
Use adversarial cases after the normal pattern is stable. They reveal whether the system learned the invariant or a superficial cue.
Human Annotation Quality
Examples are only as good as their labels and explanations. If human annotators disagree, investigate the rule before adding more examples. More inconsistently labelled data can make the prompt less clear.
Write an annotation guide for recurring tasks. Define each label, give positive and negative examples, explain boundary cases and state how missing information should be handled.
When disagreement remains legitimate, represent that uncertainty explicitly rather than forcing a false single label.
Examples and Data Governance
Real examples may contain private, confidential or regulated information. Do not move production examples into prompts or shared libraries without appropriate authorisation and sanitisation.
Prefer synthetic examples when they can represent the mechanism adequately. When real cases are necessary, remove identifying details while preserving the structure that makes the example useful.
Track provenance so maintainers know whether an example is synthetic, anonymised or derived from a real incident.
Examples for High-Stakes Domains
In high-stakes domains, examples should not be treated as substitutes for professional standards, official policy or expert review. They can demonstrate workflow behaviour, formatting or evidence handling while the underlying professional judgment remains with qualified people.
Use examples to teach the system to surface uncertainty, cite sources and escalate rather than to manufacture confidence.
The higher the consequence of a wrong output, the stronger the validation and human review surrounding the example-driven workflow.
Receiver-Centred Example Design
Choose examples according to the receiver. A student-facing example should demonstrate useful feedback without removing learning. A manager-facing example should show concise evidence and unresolved trade-offs. A machine-facing example should demonstrate exact field structure and missing-value rules.
A demonstration that is excellent for one receiver may be poor for another. Example quality is therefore partly a function of who uses the output next.
Test the receiver’s next action. If the example produces output that looks right but does not support that action, redesign it.
Example Selection Under Limited Context
When context is limited, prioritise examples with the highest information value. A contrast pair often teaches more than two similar positives. A boundary case may teach more than another easy case.
Keep the invariant, failure and output format visible. Remove decorative background. This improves both token efficiency and human maintainability.
If the example set is large, retrieve or select the most relevant subset for the current task rather than including the whole library.
Example Sets for Long-Running Workflows
Long-running workflows evolve. An example that was correct at project start may become obsolete after a policy or schema changes. Treat examples as versioned state.
Store the current example set with the workflow version. When requirements change, identify which demonstrations are invalidated and rerun regression cases.
This prevents old examples from silently teaching a superseded rule.
A Transfer Gate for Example-Driven Prompting
After the examples appear to work, remove surface similarity. Change topic, names, ordering and phrasing while preserving the underlying rule. The system should still choose the correct behaviour.
Then test one case just outside the intended operating envelope. The desired output may be ‘cannot determine’, ‘needs review’ or another explicit boundary response. This reveals whether the examples teach restraint as well as success.
Finally, run the task without one of the examples. If performance remains stable, the set may contain redundancy. If performance collapses, identify which unique boundary that example was teaching.
A Final Example Library Checklist
- Every example has a named purpose.
- Labels and explanations are internally consistent.
- Surface features do not accidentally predict the answer.
- Normal, boundary and missing-information cases are represented.
- At least one safe-stop or escalation example exists when relevant.
- Real examples are authorised and sanitised.
- Evaluation cases remain separate from demonstrations.
- Examples are versioned for recurring workflows.
- Obsolete demonstrations are retired.
- Transfer is tested on genuinely fresh material.
What the Learner Should Be Able to Do After This Article
You should be able to decide when examples are necessary, build a compact demonstration set, explain the invariant, include meaningful negative and boundary cases, and test whether the pattern transfers beyond the examples.
You should also be able to recognise when an example set is teaching the wrong thing because of bias, contradiction, phrase leakage or obsolete assumptions.
The practical evidence of skill is not a large prompt full of demonstrations. It is a small, maintainable example set that improves a real task on fresh inputs while keeping failures understandable.
Examples Are Compressed Teaching
A good example communicates a rule, a boundary and a standard of success in a small amount of context. That makes examples one of the most powerful tools in SI instruction design.
But the example must be selected as carefully as the rule. Representative demonstrations teach transfer. Narrow demonstrations teach imitation.
Use the complete SI learning hub to continue the curriculum. The next step is learning how to define constraints that protect the task when examples are not enough.
