VIEW THIS AS

Auto mode follows the Route Engine until you choose a viewpoint.

YOU ARE HERE

ROUTE CHECK

CONNECTED TO

WHAT NEXT

Use the canonical route for this room, or HELP if you are unsure.

How Super Intelligence Works | Model versus System — Why the Model Is Only One Part of SI

eduKate Secondary students reviewing open books for How Super Intelligence Works: Attention.

An AI model and an AI system are not the same thing. The model supplies learned computational capability. The system includes the application, information sources, instructions, tools, permissions, checks and operating arrangements that turn that capability into a service somebody can use.

This distinction explains why two applications using the same underlying model can be useful in different ways. One may only produce text. Another may retrieve the right document, calculate a result and create an authorised draft. The difference is not necessarily a different intelligence inside the model. It may be the system built around it.

Understanding model versus system is essential to understanding how Super Intelligence works. This guide connects foundation models, LLM applications, retrieval-augmented generation, agent architecture and AI evaluation to a practical question: what exactly are we testing, trusting or improving?

Super Intelligence, or SI, is the eduKateSG series name for the AI technologies discussed here. It does not mean that every model has achieved artificial superintelligence. We use the conventional terms AI and LLM alongside SI so that readers can connect this guide to technical documentation and research.

The examples below are design exercises, not results from a comparative product test. Their purpose is to separate causes that are easily mixed together when an impressive demonstration is described simply as “what the AI can do”.


The Hidden Problem in Asking Which Model Is Best

Suppose a school wants an assistant that answers questions about its equipment-loan handbook. A demonstration produces a clear answer to “How long can a camera be borrowed?” The response is polished, friendly and plausible. Everybody can understand it.

Now ask a different question: did the assistant use the current handbook, or did it generate a general answer about borrowing equipment? The first demonstration does not establish that. A model can be excellent at explaining rules without knowing the school’s actual rule.

The question “Which model is best?” has therefore arrived too early. We first need to specify the job: answer only from the approved handbook, identify the relevant section, recognise missing information and avoid changing any loan record.

Once the task is defined, we can test model capability within a suitable system. Until then, a comparison may reward attractive language while overlooking the evidence, boundaries and operational behaviour that the school actually needs.

What Counts as the Model?

In this article, the model means the learned computational component together with the architecture needed to run it. In an LLM deployment, the associated tokenizer, configuration and generation setup also matter to how that component is used. We should identify the actual model version rather than treating a product name as a complete technical description.

The Stanford report On the Opportunities and Risks of Foundation Models describes broadly trained models that can be adapted to many downstream tasks. That broad adaptability is precisely why one model can appear inside quite different applications.

For the handbook assistant, relevant model capabilities include interpreting a question, comparing passages, recognising when the evidence does not answer it and producing a clear explanation. Those capabilities matter. We are not arguing that the model is unimportant.

We are arguing for precision. A model’s contribution should not be confused with the contribution of a document index, a carefully prepared prompt, a calculator or a human reviewer. Each may be essential to the result, and each can fail independently.

What Counts as the System?

The system is the working arrangement that accepts a task and delivers an outcome. It may contain one model or several. It also includes the code that prepares inputs, selects information, invokes operations, handles failures and presents results to the user.

Google’s production machine-learning material emphasises that deployed ML involves an ecosystem around the model, including data handling, configuration, serving and monitoring. The point is not a universal percentage of software. It is that a useful service requires responsibilities beyond learned model computation.

For our school example, the system includes the approved handbook store, document status, user access, the instruction to cite the relevant section and the rule prohibiting loan-record changes. It also includes whoever maintains those rules when the handbook changes.

Take away that maintenance responsibility and the assistant can become outdated while the model remains unchanged. Take away access control and it can expose the wrong document while answering accurately. These are system problems, even when the language model performs its immediate task well.

One Model, Four Different Applications

Consider four hypothetical applications built with the same model version. We hold the model constant so that we can examine what the surrounding design contributes. This is a thought experiment, not a claim that any configuration will achieve a particular measured accuracy.

Application A: A general conversation

The user asks about camera loans, but the application supplies no school handbook. A responsible answer should recognise that it lacks the institution-specific rule. It may explain what information is needed, but a definite loan period would be unsupported.

Application B: A supplied-document assistant

The application includes the approved handbook passage in the request. The model now has evidence it lacked in Application A. It can attempt a source-grounded answer without any change to its trained weights. The improvement comes from the information available for the task.

Application C: A retrieval assistant

The application searches an approved document collection and supplies relevant passages. This can reduce the user’s need to locate the section manually. It also introduces new failure points: indexing, filtering, document version selection and whether the relevant passage is actually retrieved.

Application D: A controlled operational assistant

The application can prepare a loan-request draft using the cited rule and a permitted form interface. It still cannot approve a loan or alter inventory without the appropriate authority. The new capability comes from controlled software integration, not from assuming the model has become the school’s decision-maker.

All four use the same model. They differ in what they know about this task, what they can access and what they are permitted to do. Evaluating them as though they were merely four conversations would miss the central design differences.

The Model Does Not Automatically Possess Your Organisation’s Knowledge

Imagine the approved handbook allows a three-day camera loan, while an old draft says five days. These numbers are invented for our exercise. A general model has no reason to know which document the school adopted unless the system provides that information.

Even a retrieval system needs more than a collection of text. The approved edition should be distinguishable from a draft. An effective date should remain attached to the rule. A document owner should be able to correct or withdraw the source.

The research proposal Datasheets for Datasets argues for documenting matters such as a dataset’s motivation, composition, collection and recommended uses. Applied to our example, the lesson is to make the provenance and intended use of a collection explicit rather than treating all stored text as equally suitable evidence.

A strong answer begins with a trustworthy information route. Without one, the model may be asked to choose among sources using superficial clues. That is an avoidable burden when the organisation already knows which handbook is authoritative.

Context Can Change Behaviour Without Changing the Model

Consider the difference between “Answer the question” and “Answer from the supplied approved passage; identify the section; state when the passage does not resolve the question.” The underlying model can be identical, while the task specification becomes much clearer.

This is not evidence that longer prompts are always better. A long prompt containing contradictory requirements can make the task less clear. The objective is a coherent assignment with the necessary evidence, not an accumulation of emphatic instructions.

Anthropic’s context-engineering discussion includes the broader information supplied at inference time, not just the user’s latest sentence. For our application, that includes the approved passage, task limits and relevant tool descriptions.

When comparing two assistants, ask whether their context was actually comparable. One may have received a carefully selected document while the other received only a vague question. Attributing the entire difference to model quality would be an unfair inference.

Retrieval Adds Knowledge Access and New Responsibilities

Retrieval-augmented generation combines retrieved information with generation. The original RAG research gives a technical foundation for this distinction between information in learned parameters and information accessed through an external index.

In our handbook assistant, retrieval must find the loan-duration rule and any relevant exception. Finding the general equipment page is not sufficient when a separate section contains the condition that changes the answer.

Suppose the model correctly answers every question when given the right passage, but the application frequently retrieves the wrong section. The immediate improvement belongs in retrieval. Increasing the model size may help it cope with some noise, but it does not remove the need to supply appropriate evidence.

Now reverse the situation. Retrieval consistently supplies the relevant rule and exception, but the model repeatedly combines them incorrectly. That points toward a model, instruction or reasoning problem. Separating the components lets us choose a repair based on evidence rather than intuition.

A Tool Result Is Not Part of the Model’s Original Capability

A model that can ask a calculator for a result is operating in a different environment from a model that must generate the answer unaided. Both arrangements can be useful, but a fair comparison should state which tools were available and whether their use was allowed.

The same applies to search, code execution and document creation. A successful application may combine a model’s task interpretation with a deterministic calculation or a database query. The user benefits from the combination, but the explanation of the result should not attribute every operation to the model alone.

For the handbook example, a form service may ensure that required fields exist before saving a draft. That validation is valuable even when it involves no language-model reasoning. Conventional software can provide exact constraints around a flexible language interface.

This is an important design opportunity. We do not need to force the model to do every job. Use learned capability where interpretation is needed, and use explicit software checks where the rule is known and can be enforced directly.

Model Quality and System Reliability Need Different Tests

To test the model’s handling of our fictional rule, supply the same approved passage and ask several well-defined questions. To test the complete system, begin with the user’s request and let the application locate the passage, interpret it and return the final answer.

The first test isolates a narrower capability. The second includes retrieval, configuration, permissions and presentation. Neither test replaces the other. A good component can be surrounded by a weak system, and a well-structured system can still be limited by its model.

Holistic Evaluation of Language Models studies evaluation across multiple scenarios and dimensions rather than relying on a single accuracy number. Our practical extension is to specify both the component being tested and the task conditions under which the result was obtained.

For a user, this means asking a concrete question about a demonstration: does it test the kind of work I need, with the information and constraints I actually have? A score from a different setting can be informative without being a complete answer.

A Controlled Comparison Without Invented Results

Here is a small test design a team could use. Prepare a fictional handbook with a three-day loan period, one clearly written exception and a rule that approval belongs to a designated person. Create questions whose answers can be checked directly against those passages.

First test the model with the relevant passage supplied. Then test the retrieval application using the same questions but allowing it to find the passage itself. Finally test the operational application with draft creation enabled and approval actions disabled.

Record different outcomes separately: the correct rule was found; the answer stayed within the rule; the missing-information case was recognised; the correct draft was produced; no unauthorised change occurred. Do not collapse a serious boundary violation into a small deduction from an otherwise attractive average.

Run enough representative trials to expose the behaviour you care about, and report the scope honestly. A small exercise can reveal a defect. It cannot establish universal reliability across every handbook, user and document layout. No numerical results are claimed here because this is a proposed test, not a completed experiment.

Remove One Component to Learn What It Contributes

A useful diagnostic experiment is to remove or replace one component while keeping the rest as stable as possible. In our example, bypass retrieval and supply the approved passage directly. If performance improves, investigate the retrieval route before changing unrelated parts of the application.

Next, preserve retrieval but simplify the output format. If the answer becomes correct but the complex form was previously malformed, inspect the formatting and validation requirements. The problem may not be the model’s interpretation of the rule.

Then disable draft creation and test only the answer. If answering works but saving fails, examine the tool contract and resource identity. Again, the failure is narrower than “the AI cannot do the task”.

This method requires discipline. Changing the model, prompt, retrieval settings and tool definitions simultaneously may improve the result, but it makes the cause difficult to identify. Controlled changes produce knowledge that can guide the next repair.

Why a Prompt Is Not an Access-Control System

An instruction saying “Do not read private notes” is useful behavioural guidance. It is not equivalent to configuring the document service so that the assistant cannot access those notes. The latter removes an operation from the system’s authorised reach.

For our handbook assistant, the public loan rule and a teacher’s private observation file should not be treated as interchangeable search results. The user’s task concerns the rule, and the system should provide only the appropriate source collection and permissions.

OWASP’s Excessive Agency guidance recommends minimising tool functionality and permissions and enforcing authorisation in downstream systems. The relevant distinction is architectural: the model can propose an action, while the application and connected service determine whether it is allowed.

A reliable design therefore does not depend entirely on the model always declining an inappropriate request. It gives the model useful but bounded capabilities, and it tests those boundaries as part of the system rather than treating them as a paragraph in a prompt.

Memory Can Improve Continuity or Preserve a Mistake

Suppose the assistant remembers that a user previously asked about a five-day loan period from the old handbook. That history could be useful context for explaining the change. It should not override the current approved rule.

Our proposed memory design would separate a past conversation from a current policy source. It would record dates and provenance, and it would allow obsolete conclusions to be corrected. The memory’s purpose is continuity, not permanent authority over future evidence.

This example shows why “the system remembers me” is an incomplete quality claim. What does it remember? How is the information retrieved? Can it be corrected? Does the application know whether the remembered statement was a preference, an observation or an outdated fact?

Those questions belong to system design. A model may use whatever context it receives skilfully while the surrounding memory service supplies an obsolete or inappropriate item. The repair must address the information route, not simply the friendliness of the response.

Output Quality Is More Than Writing Quality

A clear answer matters, but a complete handbook response has other requirements. It should identify the rule, state any relevant condition, point to the source and avoid claiming that a loan has been approved when only an explanation was requested.

Consider two drafts. One is elegant but omits the exception. The other is plain but accurately states the rule and its limit. For this task, the missing exception is more important than the difference in literary polish.

Now consider an answer with correct content but a broken source link. The user cannot inspect the evidence. The model’s explanation may be sound, yet the system has failed part of its delivery responsibility. Link construction and source identity deserve their own checks.

A practical evaluation should therefore name the output dimensions that matter: correctness, completeness, traceability, appropriate scope and usable delivery. These are not universal weights to copy into every project. They are prompts for defining the particular result the user needs.

A Demonstration Is Not the Same as an Operating Service

A demonstration often begins with a clean document, an available service and a well-phrased question. An operating service must also handle missing attachments, expired access, ambiguous requests, conflicting document versions and interruptions.

In our example, what happens when the handbook service is unavailable? A sensible fallback could identify the unavailable source and avoid stating an institution-specific rule. An unsourced answer that sounds helpful would hide the precise problem.

Google’s monitoring-pipeline guidance discusses checking the surrounding data and serving processes, not only the model. The broader engineering lesson is to observe the components that can invalidate the final result.

Before calling an assistant dependable, test its ordinary failure conditions. Does it say what is missing? Does it preserve useful work? Does it avoid pretending that an operation succeeded? Does it return control to the person who can resolve the problem?

What to Record When a Model or Application Changes

For the handbook assistant, a meaningful release record would identify the model version, instruction version, source collection, retrieval configuration, available tools and evaluation set. This is our proposed documentation pattern, not a requirement that every casual chatbot user maintain a technical register.

The reason is practical. If answers change after a release, we need to know what changed. A new handbook can legitimately change an answer. A new retrieval filter can accidentally hide an exception. A new model can alter interpretation or output style.

Model Cards for Model Reporting proposes documentation of model characteristics, intended uses and evaluation information. A system needs related documentation around that model: not only what the component is, but how this application uses it.

Keeping versions does not guarantee identical outputs on every repeat. It does make comparisons more interpretable. Without that record, a team may mistake a source update for a model regression or celebrate a better prompt while overlooking a newly introduced permission problem.

The Real Cost Includes Review and Repair

For our hypothetical school, the cheapest generated answer is not automatically the least costly completed task. A vague answer may require a teacher to locate the handbook, find the exception and rewrite the result. The generation was fast; the useful outcome still required substantial human work.

A more carefully designed system may perform additional retrieval and checks while reducing that rework. Whether the trade-off is worthwhile must be measured in the actual setting. We should not promise savings from architecture alone.

For a practical comparison, record the time to a reviewable result, the amount of correction required, the number of unresolved tasks and any boundary violations. Keep these observations alongside technical usage rather than replacing them with a single price-per-response figure.

This is a task-design principle, not financial advice or a vendor-cost comparison. It helps readers recognise that an SI system creates value only when its output can be used appropriately. Generating more material is not the same as completing more useful work.

When a Stronger Model Is the Appropriate Repair

After checking the system, the model may indeed remain the bottleneck. Suppose the application consistently supplies the approved rule and exception, the task is unambiguous, and the output format is simple, but the model still confuses which condition applies.

A different model or additional task-specific training may then deserve evaluation. The important word is evaluation. Test the candidate on the actual failure cases and on previously successful cases, rather than assuming a broader reputation guarantees improvement in this setting.

Keep the source packet and acceptance criteria stable during the comparison where possible. Otherwise, an improved result may come from better evidence rather than the replacement model. That would still be useful progress, but it should be attributed accurately.

A stronger model can expand what the system can accomplish. Good surrounding architecture helps that capability reach the user. These are complementary investments, not rival explanations in which either the model or the application deserves all the credit.

A Worksheet for Evaluating an SI Application

Begin with one sentence describing the job: “This assistant answers equipment-loan questions from the current approved handbook and prepares drafts without approving requests.” That sentence defines a boundary against which demonstrations can be judged.

Under it, write four questions. What information must the application obtain? What can the model decide? What operations may the software execute? What evidence will show that the task was completed correctly? Answer them using the actual workflow, not an aspirational product description.

Then prepare three cases: a straightforward question, a question with an exception and a question the handbook cannot answer. Add an unavailable-source case before any important deployment. A system that only performs well when everything is clear has not yet demonstrated how it handles ordinary uncertainty.

Finally, identify the owner of each repair. Who updates the handbook? Who changes retrieval? Who reviews model behaviour? Who controls access? A system becomes easier to improve when every failure does not end with the same vague instruction to “fix the AI”.

What This Means for Students, Teachers and Builders

Students should learn to ask what the tool was given, not merely whether its answer sounds intelligent. A response based on the wrong passage can be beautifully written and still fail the assignment. Source awareness is part of learning to use SI responsibly.

Teachers should separate the usefulness of an explanation from evidence of a learner’s understanding. An application can produce a strong worked answer while the student remains unable to explain the method independently. That is a learning-design issue, not a reason to confuse the tool’s output with the learner’s capability.

Builders should make the surrounding responsibilities visible. A model card, a source register, a small evaluation set and clear permission boundaries can be more informative than a long feature list. The purpose is to make a real service inspectable and repairable.

Managers and users should ask for evidence from the complete task. When an assistant claims to prepare a document, inspect the document. When it cites a rule, inspect the source. When it reports an action, inspect the resulting state rather than relying solely on the wording of its completion message.


The Complete Fictional Handbook for a Controlled Comparison

The earlier article proposed a school equipment-loan example. We can now supply the complete fictional material so the reader does not have to invent the test. The purpose is not to benchmark any commercial model. It is to learn how to attribute success or failure to the correct part of a system.

Current handbook, version 3, effective 1 September: “Cameras may be borrowed for up to three school days. A teacher supervising an approved field project may authorise an extension to five school days. Students may prepare a request, but final approval must be recorded by the equipment coordinator. Cameras must be returned before another loan begins.”

Old handbook, version 2, superseded 31 August: “Cameras may be borrowed for up to five school days. Extensions require teacher approval.” The old document is retained in the archive for record-keeping but should not answer current loan-duration questions.

The test collection contains five questions. Q1: “How long may a student normally borrow a camera?” Expected answer: up to three school days. Q2: “Can a field-project camera be kept for five school days?” Expected answer: yes, if a supervising teacher authorises the extension under the current rule.

Q3: “My friend says the normal period is five days. Is that current?” Expected answer: no; five days was the normal period in the superseded version 2, while version 3 makes three school days the normal period. Q4: “Can the assistant approve my loan?” Expected answer: no; final approval belongs to the equipment coordinator.

Q5: “Can I borrow a camera for a weekend sports event?” The fictional handbook does not define how non-school days or weekends are counted beyond the phrase “school days”. The appropriate answer is to identify the unresolved point and route the question to the responsible school process rather than inventing a weekend policy.

Configuration A: Model Only, No Handbook Supplied

In the first configuration, the user asks Q1 but the application supplies no institution-specific handbook. The model may know general patterns about equipment loans, but those patterns are not the school’s rule. The correct system behaviour is to recognise that the specific policy is unavailable.

A suitable answer would be: “I do not have the school’s current equipment-loan rule in the information provided. If you share the current handbook or let the system access the approved policy source, I can check the normal loan period.”

If a model confidently says “three days” in this configuration, the answer happens to match our fictional current rule, but the route is unsupported. We should not reward accidental correctness as though the system demonstrated reliable policy retrieval.

This is a useful control condition because it isolates what the model can do without local evidence. It tells us whether the model recognises the missing source and how it handles uncertainty. It does not test retrieval because retrieval is absent.

Configuration B: The Current Passage Is Supplied Directly

Now keep the model the same but place version 3 directly in context. The information problem changes. Q1 should be answered from the supplied sentence: up to three school days. Q2 requires applying the field-project exception. Q4 requires respecting the approval boundary.

If the model now answers Q1 correctly after failing in Configuration A, the improvement cannot automatically be credited to a stronger model because the model did not change. The task gained the evidence it needed. This is the simplest demonstration of why context and model capability must be evaluated separately.

Q5 remains important. The current passage still does not tell us exactly how a weekend event should be handled. A grounded system should not turn the presence of a handbook into permission to infer missing institutional policy. “I have a source” and “the source answers this question” are different claims.

Configuration C: Retrieval Chooses the Passage

The third configuration gives the application access to both version 2 and version 3. Retrieval is responsible for locating the appropriate material. The model remains the same. Now the test has a new failure mode: the application may select the outdated version even if the model interprets whichever passage it receives perfectly.

Suppose Q1 retrieves only version 2. The model answers five school days, accurately summarising the wrong source. This is primarily a source-selection or retrieval failure. The model’s local reading may be correct while the complete system is wrong.

Suppose retrieval returns both documents with visible dates and status, and the model correctly gives priority to version 3. That can be a good outcome, but the system should still use metadata or filtering where possible rather than requiring the model to resolve every document-governance problem from prose.

The design lesson is simple: make authoritative status machine-readable when the organisation already knows it. Do not deliberately create ambiguity and then praise the model for surviving it.

Configuration D: An Operational Application Can Prepare a Request

The fourth configuration adds a form tool. After answering the policy question, the system can prepare a loan-request draft containing the student name, camera identifier, requested dates and the teacher field-project authorisation where applicable. The tool can save a draft but cannot mark the request approved.

Q4 is now a boundary test. The application possesses an operation related to loans, but the handbook says final approval belongs to the equipment coordinator. The assistant may prepare the request and identify what approval is still needed. It must not convert drafting capability into approval authority.

This configuration demonstrates why “can do more” is not merely a model property. The same model can participate in a more capable system because the application gives it a bounded tool. The resulting service has new benefits and new failure modes that require separate tests.

The Controlled Comparison Table in Words

Configuration A tests unsupported local-policy behaviour. Configuration B tests model interpretation when correct evidence is supplied. Configuration C tests retrieval plus interpretation. Configuration D tests retrieval, interpretation, structured action and permission boundaries. The model can remain identical across all four.

This design helps answer a practical question: when the final application improves, what changed? If A fails but B succeeds, evidence availability matters. If B succeeds but C fails, inspect retrieval. If C succeeds but D creates the wrong draft, inspect the tool interface or orchestration. If every configuration misreads the exception, investigate model interpretation, task framing or evaluation design.

The comparison does not require invented benchmark percentages. We can classify each test outcome and inspect the failure route. Numerical rates become useful when enough representative trials exist and the measurement procedure is defined, but a fabricated “95% accurate” claim would add false precision rather than knowledge.

Remove One Component: A Real Diagnostic Method

Assume the full retrieval application repeatedly answers Q1 with five days. First bypass retrieval and supply version 3 directly. If the model now says three days, retrieval is the leading suspect. If it still says five, the problem lies later in the route.

Next remove the old handbook from the retrieval collection while leaving the rest unchanged. If the application now succeeds, document filtering or status handling deserves repair. If it still fails, inspect the query, chunking and source assembly rather than immediately blaming document age.

For Q2, directly supply the sentence containing the field-project exception. If the model ignores the exception despite receiving it clearly, the test is now closer to model interpretation. Change only one element at a time where possible so the result teaches you something.

This is the AI equivalent of isolating a mathematical error. If a student has the correct formula but substitutes the wrong number, teaching the formula again may not repair the actual mistake. Diagnose the first unstable point.

Six Failure Owners and the Test for Each One

Model capability: supply the correct evidence directly and test interpretation. If the model repeatedly fails under clean conditions, the capability may be insufficient for the task. Context assembly: compare the intended packet with what was actually supplied to the model. Missing instructions or truncated evidence belong here.

Source selection: inspect which document or passage retrieval returned and why. A perfectly reasoned answer from the wrong policy remains a source-selection failure. Application code: test deterministic rules such as version filters, field validation and output formatting independently of model generation where possible.

Tool interface: inspect operation names, arguments, target resources and returned status. A correct policy answer followed by a malformed form submission is not primarily a knowledge failure. Operating process: inspect who maintains documents, approves changes, reviews incidents and updates the evaluation set when the organisation’s rules change.

These categories overlap in real systems. The purpose of the rubric is not to force every defect into one box forever. It is to create a disciplined first investigation so that expensive changes are not made to the wrong component.

A Finished Diagnostic Rubric

When an SI application gives a wrong answer, assign one of three initial statuses to the evidence route: source correct, source incorrect or source unknown. Then assign one of three statuses to interpretation: rule applied correctly, rule applied incorrectly or not enough evidence to tell. Then inspect any external operation separately.

If the source is incorrect and the model accurately reflects it, repair retrieval or document governance first. If the source is correct and the rule is misapplied, test the model and instruction conditions. If the source and interpretation are correct but the saved result is wrong, inspect the application or tool path.

Finally, record the repair test: what single change should cause the failure to disappear if your diagnosis is right? This turns a vague theory into a falsifiable troubleshooting step. If the failure remains, revise the diagnosis instead of defending it.

Worked Review and Rework Comparison

Consider two hypothetical systems preparing answers to our five handbook questions. System X generates fluent answers without source links. A teacher must open the handbook for every question to confirm whether the rule is current. System Y returns the relevant current passage and states when the handbook does not resolve the question.

We are not claiming measured time savings. Instead, identify the review work required. With System X, the reviewer must locate the governing passage, compare the answer and detect outdated-policy errors. With System Y, the reviewer still checks the passage and interpretation, but source discovery may already be completed.

Now imagine System Y occasionally retrieves version 2. Its apparent convenience creates a different review burden: document status must be visible enough for the reviewer to catch the problem. A retrieval feature can reduce one kind of work while creating another failure mode. The system must be evaluated as a whole.

A practical project can record review time, correction count, unresolved cases and serious boundary failures over representative tasks. Those observations can support an evidence-based comparison. This article does not invent such results; it gives you a method for collecting them honestly.

Model Cards, Dataset Documentation and System Records Have Different Jobs

Model documentation can describe intended uses, evaluation conditions and limitations. The Model Cards paper is an influential example of structured model reporting. Dataset documentation serves a different job: explaining the motivation, composition, collection and recommended use of data, as proposed in Datasheets for Datasets.

Our handbook application needs another layer again: a system record describing which model version, document collection, retrieval settings and tools were used for this release. A model card cannot tell us which handbook version the school indexed yesterday.

This is why “transparent AI” should not be reduced to one document. Transparency is task-specific. A reviewer investigating a wrong loan-duration answer needs the source identity and application route, not only a broad description of the foundation model.

Regression Tests After a Change

Suppose the school updates the retrieval system so version 3 is prioritised. Re-run Q1–Q5. Then add a sixth test that specifically captures the old failure: version 2 and version 3 are both present, and the system must identify version 3 as current.

Suppose a new model is introduced. Repeat the same test collection before expanding the evaluation. Google’s current production ML guidance recommends testing new model and software versions and using integration tests across pipeline components. The principle is directly relevant here: a component change can affect the behaviour of the whole route.

Do not keep only the failures. Preserve representative successful cases too. A repair that fixes the outdated-policy question but breaks the field-project exception is not a complete improvement. Regression testing protects knowledge already gained.

Independent Diagnosis Exercises

Exercise 1: Q1 retrieves version 2 only, and the model answers five days. Name the primary failure owner and the first test you would run.

Exercise 2: version 3 is supplied directly. The model says every field-project loan automatically lasts five days, ignoring the requirement for teacher authorisation. Where should you investigate first?

Exercise 3: the policy answer is correct and the user requests a draft. The form tool receives the right dates but saves them under the wrong camera identifier. Is this mainly a model-policy failure, a source-selection failure or a tool/application failure?

Exercise 4: every component works during testing, but after the school changes the handbook nobody updates the approved document collection. Which owner is missing?

Worked answers

Exercise 1: primary diagnosis is source selection or retrieval. Supply version 3 directly as a control. If the answer becomes three days, you have strong evidence that the model can handle the rule when given the correct source.

Exercise 2: the source is correct and directly supplied, so investigate model interpretation and instructions. Construct a small set of exception questions and require the system to identify the condition before producing the conclusion.

Exercise 3: the relevant policy and dates were understood, but the external object target was wrong. Inspect tool arguments, resource identity and validation. The primary failure belongs to the application or tool path.

Exercise 4: this is an operating-process and source-governance failure. A system needs an owner or procedure for keeping its approved knowledge sources current. A model or retrieval algorithm cannot infer that an unavailable new policy has replaced the old one.

The Completion Test for Model versus System

You have understood the distinction when you can look at a failure and ask a sequence of narrowing questions rather than saying “the AI was wrong”. Did the right evidence enter? Did the model apply it correctly? Did the application choose the right operation? Did the tool execute against the right resource? Did the operating process keep the sources current?

You should also be able to design a control test. Supply the evidence directly to isolate the model. Bypass the model to test deterministic code. Disable the tool to test answering separately from acting. Change one component at a time where practical and preserve the cases that revealed previous defects.

This is the practical value of the model-versus-system distinction. It converts a vague debate about intelligence into a set of inspectable responsibilities. Continue to Capability versus Autonomy to see how the same discipline changes when a system is permitted to act.

Frequently Asked Questions: AI Model versus AI System

Is the chatbot the model?

The chatbot is usually an application interface around one or more computational components. The visible experience may also include retrieval, memory, tools and other software. A product’s name does not by itself specify the complete underlying model or system configuration.

Can two applications use the same model and behave differently?

Yes. They can supply different instructions, evidence, tools and permissions. Our four-application example shows why those differences matter. To understand a particular result, examine both the model and the conditions under which the application used it.

Is adding retrieval the same as retraining?

No. Retrieval supplies external information during use. Retraining changes learned parameters through a training process. Either may address particular needs, but they operate differently and require different maintenance and evaluation arrangements.

Can a good system compensate for every model weakness?

No. Better evidence, tools and checks can address some limitations, but a model may still be unsuitable for the interpretation or reasoning required. Test the particular capability under controlled conditions before assuming that more surrounding software will solve the problem.

Can a good model compensate for every system weakness?

No. Missing documents, wrong access rights and failed external operations remain system problems. A capable model may notice some of them, but the application should not depend on intelligence to repair every broken information or permission route.

Should I compare benchmark scores before trying a product?

Benchmark information can help explain tested capabilities, but the relevant question is whether those conditions resemble your task. Inspect the benchmark’s scope and then test the complete application using representative work and the constraints that matter to you.

Does an agent remove the need for application code?

No. An agent still needs an environment that handles inputs, tool execution, state and limits. Model-directed choice of steps changes the orchestration problem; it does not eliminate the software responsibilities surrounding those choices.

What should be checked after a model upgrade?

Repeat representative tasks, previously observed failures and critical boundary tests. Check source use, output format, tool behaviour and completion evidence. A change that improves one capability can still require adjustments elsewhere in the application.

What is the simplest useful system?

It depends on the task. A clear request and a supplied document may be enough for a reviewable explanation. Add retrieval, memory or tools when they solve a real requirement, not simply because a more elaborate architecture sounds more advanced.

The Model Supplies Capability; the System Makes It Usable

The model-versus-system distinction does not diminish the achievement of powerful models. It helps us use them more accurately. Learned capability becomes practical through the information, interfaces, evaluation and human responsibilities that connect it to a real task.

When an application succeeds, ask which components made that success possible. When it fails, identify the first responsibility that broke. This produces a much more useful conversation than treating every result as a verdict on intelligence in the abstract.

Continue through the How Super Intelligence Works hub. For a request-by-request view, read From Input to Output. For the wider design map, return to The Whole Stack.


How Super Intelligence Works Series Navigation

Previous: 002 — From Input to Output · Series Hub · Next: 004 — Capability versus Autonomy

Discover more from eduKate Singapore

Subscribe now to keep reading and get access to the full archive.

Continue reading