VIEW THIS AS

Auto mode follows the Route Engine until you choose a viewpoint.

YOU ARE HERE

ROUTE CHECK

CONNECTED TO

WHAT NEXT

Use the canonical route for this room, or HELP if you are unsure.

How Superintelligence Works | SI versus Traditional Software — Learned Models, Explicit Rules and the Power of Hybrid Systems

eduKate Secondary students reviewing open books for How Super Intelligence Works: Attention.

Super Intelligence and traditional software solve problems in different ways. Traditional software is usually built from explicit instructions written by programmers: if this condition is true, perform this operation; store this value in this field; reject an invalid input; return a defined result. Machine-learning systems learn patterns from data and can make predictions or generate outputs that were not individually hand-coded.

The most useful real systems are often hybrids. A language model can interpret flexible human input, while traditional software validates fields, calculates exact values, checks permissions, saves records and enforces hard constraints. Understanding SI versus traditional software is therefore not about choosing one side. It is about deciding which part of a task should be learned and which part should remain explicit.

This article develops that distinction through worked examples, testing methods, failure modes, cost, security, maintainability and system design. We will repeatedly ask one Clementi-style diagnostic question: where is the first unstable point, and which kind of computation should own the repair?

Google’s current Production ML Systems module emphasises that real-world machine-learning systems are large ecosystems in which the model is only one component. That is the central engineering frame for this article.

Previous: Why Language Models Can Do More Than Language. Here we compare learned models with explicit software rules and show how to combine them without making either component responsible for work it does poorly.


The Hidden Transition: From Written Rules to Learned Behaviour

Imagine a library loan system. Traditional software can calculate a due date from a checkout date and a rule such as “14 days”. It can store the borrower ID, mark the item returned and prevent two simultaneous active loans for the same physical copy. These rules are explicit and testable.

Now imagine a user writes: “I brought the blue book back yesterday but the system still says I owe it.” Traditional software does not automatically know which book the user means or whether “yesterday” refers to a return record. A language model can interpret the message, extract likely intent and ask for missing information.

The model is better at flexible language; the database and transaction logic are better at authoritative state. A robust system lets each component do the job that matches its strengths.

What Traditional Software Means in This Article

Traditional software here means programs whose main behaviour is defined by explicit code, algorithms, rules, data structures and interfaces rather than learned statistical parameters. It includes databases, web servers, compilers, calculators, business rules, validation code and deterministic workflows.

Traditional software can still contain complex algorithms. Search algorithms, encryption, route planning and database query execution are not “simple” merely because they are hand-specified. The distinction is about how behaviour is defined, not about whether the system is sophisticated.

A program can also use randomness intentionally. Traditional software is not automatically deterministic in every run. The useful contrast is explicit programmed logic versus learned statistical mapping.

What SI Adds

SI adds models whose behaviour is learned from data. Instead of specifying every mapping from input to output, training adjusts parameters so the model performs well on examples and generalises to new cases.

This is valuable when the input space is too varied for practical hand-written rules: natural language, images, audio, complex pattern recognition and open-ended generation. A programmer can write rules for valid JSON more easily than rules for every possible way a person might ask for help.

The trade-off is that learned behaviour is approximate. The model can produce unexpected outputs, respond differently to small wording changes and fail on cases outside its evaluation distribution.

Worked Example: A Library Loan Assistant

We will use one fictional system throughout the article. The authoritative library database stores item_id, borrower_id, checkout_date, due_date, return_date and status. The public loan rule is 14 calendar days. An item marked returned cannot remain in active-loan status.

The SI layer receives messages such as “I returned the blue book yesterday” or “Can I keep this until next Friday?” It can extract likely intent, identify referenced dates, ask for the item ID when ambiguous and generate a clear explanation.

The traditional software layer performs the actual query and update. It checks the borrower’s authenticated identity, verifies that the item belongs to the active loan, records the return transaction and calculates the due date according to the explicit rule.

The assistant should never invent database state. If no returned transaction exists, it can say that the current record still shows the item active and explain the next step. It cannot make the book returned merely because the user says it was returned, unless the workflow explicitly authorises a correction process.

Rule-Based Calculation: Where Traditional Software Wins

Suppose a book is checked out on 1 October and the rule is exactly 14 calendar days. The due-date calculation can be implemented with a date library and tested against known cases. There is no benefit in asking a language model to “reason” about every date from scratch.

The explicit function can also handle time zones, leap years and date arithmetic consistently. If the policy changes to 21 days, the rule can be updated in a controlled configuration or code path.

A language model can explain the due date to the user, but the date calculation should come from the authoritative rule engine or database when exactness matters.

Natural-Language Interpretation: Where SI Helps

Now the user says, “I need the book for the project presentation next Fri; can I extend?” The model can recognise that the request is about an extension, identify the relative date phrase and gather the information needed for the policy check.

Traditional software can then apply the actual extension policy. Perhaps extensions are allowed only when no reservation exists. The model does not need to encode that policy in its prose memory; the application can query the live reservation status and explicit rule.

This is a powerful hybrid: learned interpretation at the edge, explicit state and rules at the core.

Deterministic Validation Around Probabilistic Generation

Suppose the model extracts a structured request: item_id = B431, requested_extension_days = 7. The application can validate that the item ID exists, the borrower owns the active loan and the requested extension is within allowed bounds.

If the model outputs an impossible date or malformed ID, deterministic validation rejects the operation. The system can ask the user to clarify rather than passing bad data into the database.

This is one of the strongest design patterns in modern SI: let the model interpret flexible input, then enforce hard constraints with ordinary software.

Traditional Software Is Better at Invariants

An invariant is a condition that must remain true. For a loan system, one physical item cannot be simultaneously marked as available and actively loaned to two borrowers. A database transaction can enforce that invariant.

A prompt saying “never create duplicate loans” is weaker than a database constraint that makes duplicate active loans impossible. Behavioural instructions help the model; structural constraints protect the state.

Where a rule can be expressed exactly and enforced cheaply, traditional software is often the stronger owner.

SI Is Better at Fuzzy Boundaries

Many human requests are not exact. “Find the part where the teacher explains why the experiment failed” is not a rigid database query. “Rewrite this so a parent understands it” has no single correct output string.

Learned models handle these fuzzy mappings because they can generalise across varied phrasing and examples. The result still needs evaluation, but explicit rules for every linguistic variant would be expensive and brittle.

The Difference Between Code and Weights

Traditional software behaviour is largely encoded in source code and configuration. A developer can inspect an if-statement, formula or SQL query directly. Machine-learning behaviour is also influenced by learned weights—millions or billions of numerical parameters adjusted during training.

This changes debugging. If a due-date function adds 13 days instead of 14, the relevant line of code may be obvious. If a model misclassifies one ambiguous message, there may be no single parameter that corresponds to that mistake.

ML debugging therefore relies more heavily on datasets, evaluation cases, prompt/context inspection and statistical behaviour across examples.

Testing Traditional Software

Traditional unit tests specify an input and expected output for a function. Example: checkout_date = 1 October, loan_period = 14 days, expected due date = 15 October under the defined counting convention. The test passes or fails.

Integration tests check components together: does the return endpoint update the database and make the item available? End-to-end tests check the user workflow.

These tests are powerful because many software functions are intended to behave identically for the same input and state.

Testing SI

SI requires a broader evaluation set because inputs vary and outputs can have multiple acceptable forms. For message classification, build representative examples across clear, ambiguous, noisy and adversarial cases.

For generation, define acceptance criteria: factual support, completeness, format, tone, prohibited claims and tool behaviour. Exact string matching may be too strict for prose, while vague “looks good” review is too weak.

A regression set preserves failures and representative successes so model, prompt or pipeline changes can be compared over time.

The Same Task Can Have Two Different Kinds of Test

Take “explain the current loan status”. The traditional software test checks that the database query returns the right row and status. The SI test checks that the explanation accurately reflects that row without inventing a reason.

If the row says active and due 15 October, the generated explanation may vary in wording. The factual fields should not vary. Separate exact state validation from flexible language evaluation.

Failure Pattern 1: Using SI for an Exact Rule

Suppose the application asks the model to calculate late fees entirely from prose. The rule is $1 per day after the due date, capped at $10. The model sometimes miscounts days or forgets the cap.

Repair: implement the fee calculation in deterministic software. Let the model explain the result using the returned fields. The model can still parse the user’s question, but the rule engine owns the exact arithmetic.

Failure Pattern 2: Using Traditional Rules for Open Language

Suppose programmers attempt to detect extension requests with keyword rules: if message contains “extend”, “longer”, “extra days” or “keep”. Users write “Can I hang onto it until Friday?” and the rule misses the intent.

Repair: use a classifier or language model to infer intent, then validate the extracted request. Learned models are useful where language variation makes hand-written rules explode.

Failure Pattern 3: Letting the Model Own Authoritative State

A user asks whether a book is returned. The model replies from conversation memory instead of checking the database. The answer may be stale.

Repair: query authoritative state at runtime. The model can explain the state, but the database remains the source of truth.

Failure Pattern 4: Letting Software Treat Model Output as Trusted Input

The model generates a borrower ID that looks valid. The application updates that borrower’s record without checking authentication or resource ownership.

Repair: validate identity and permissions in conventional software. Generated output is untrusted input to downstream systems until checked.

A Hybrid Architecture in Full

User message → language model interprets intent → structured request → validator checks fields → database reads current state → rule engine calculates allowed action → model explains result → user approves any consequential change → transaction executes → database returns new state → model reports verified outcome.

Every arrow is a handoff. Each handoff can fail. The architecture is strong because responsibilities are explicit. The model is not asked to be a database, and the database is not asked to understand unrestricted human language.

State: Traditional Software Usually Owns the Record

State is what the system currently knows about loans, users, files, balances or workflow status. Databases and transactional systems are designed to preserve state with explicit consistency rules.

Language models can carry temporary context and applications can provide persistent memory, but that is different from authoritative transactional state. A conversation saying “I returned it” should not silently overwrite the library database.

A system can store user preferences in memory while keeping loan status in the library database. Different kinds of state deserve different owners.

Versioning: Code Releases and Model Releases

Traditional software versions can change behaviour through code and configuration updates. Model versions can change learned behaviour. SI applications often change both at once, plus prompts, retrieval settings and tools.

A release record should therefore identify the model version, relevant application code, instructions, data or source collection and evaluation set. Otherwise, a behavioural change can be hard to attribute.

Google’s production ML guidance emphasises testing new model versions and monitoring deployed systems because model and software changes interact.

Maintenance: Rules Drift and Data Drift Are Different

Traditional software can become wrong when business rules change but code does not. Machine-learning systems can become wrong when input distributions shift or labels change.

Our library example might change the loan period from 14 to 21 days. A deterministic rule must be updated. The language model may continue explaining old examples if stale documents remain in retrieval. Both code and information governance need maintenance.

A model classifier may also drift if users adopt new phrasing. The repair could involve new examples, retraining or prompt updates. Different failure mechanisms require different maintenance plans.

Security: Explicit Boundaries Still Matter

Traditional access control should remain outside model discretion. Authentication, authorisation, data permissions and write restrictions are software and infrastructure responsibilities.

The model can help interpret what the user wants, but it should not decide that a user “sounds like an administrator”. Identity comes from the authenticated system, not generated inference.

This is especially important for tool-using agents. The more flexible the model, the more important it becomes to bound what downstream software permits.

Cost and Latency

Traditional code can be extremely fast and cheap for fixed operations. A database lookup or arithmetic function should not be replaced with a large model call simply because models are available.

SI calls can add computational cost and latency, especially for long context or repeated agent loops. The trade-off is worthwhile when flexible interpretation or generation reduces substantial human work.

A good architecture routes simple deterministic work to simple deterministic components and reserves model computation for tasks that benefit from learned flexibility.

Explainability Is Different for Rules and Models

An explicit rule can often be explained by showing the code or decision table. A model output may require evidence from examples, source context, evaluation and post-hoc analysis rather than one readable rule.

This does not make all traditional software automatically transparent. Large codebases can be difficult to understand. It means the source of behaviour differs: explicit logic versus learned parameters.

Worked Example: Rule Engine Plus Language Model

Policy: a book may be renewed once for 7 days if no reservation exists and the item is not already overdue. User: “Nobody else seems to want this book. Can I keep it another week?”

Model output: intent = renewal request; desired extension = 7 days; item reference = current active book. Software checks current item ID, overdue status, previous renewal count and reservation table.

If no reservation exists, the item is not overdue and renewal count is zero, the rule engine returns allowed = true, new_due_date = old_due_date + 7 days. The model explains: “Your renewal is allowed for seven more days because the book is not overdue and no reservation is recorded.”

If a reservation exists, the model does not override it because the user wrote “nobody else seems to want it”. The database state controls the rule.

Decision Tables: A Bridge Between Human Policy and Code

A decision table can make explicit software logic readable. Conditions: overdue? reservation exists? already renewed? Outcomes: allow renewal or deny.

For example, overdue = no, reservation = no, already renewed = no → allow. Any condition yes → deny under our fictional rule. A programmer can implement and test this table directly.

The model can map natural language into the condition values, but the application should verify those values from the authoritative database before applying the decision.

Why Not Replace the Rule Engine With More Prompting?

Prompts are useful for model behaviour, but exact policy enforcement benefits from explicit constraints. A prompt can be forgotten, misinterpreted or overridden by conflicting context. A database constraint or rule function can make prohibited states impossible.

The goal is not “never put rules in prompts”. Many conversational behaviours belong there. The goal is to match enforcement strength to consequence.

Why Not Replace the Model With More Rules?

Hand-written rules become difficult when inputs are open-ended. Every new phrase, spelling mistake, language variation and indirect request creates another branch.

A trained model compresses many linguistic patterns into learned parameters and can generalise to unseen phrasings. That is exactly where learned systems earn their complexity.

The Boundary Between Learned and Programmed Can Move

A team may begin with rule-based intent classification because there are only five phrases. As usage grows, a model may become worthwhile. Another team may begin with model-generated calculations, then move arithmetic into deterministic code after observing mistakes.

Architecture is not ideology. Responsibilities can migrate as requirements and evidence change.

A Practical Selection Framework

Use traditional software when: the rule is explicit; exactness is required; the state is authoritative; the operation is safety- or permission-critical; the task is easily specified; or the behaviour must be highly repeatable.

Use SI when: the input is variable language or media; the task involves fuzzy classification; useful outputs have many valid forms; synthesis across unstructured material is needed; or hand-written rules would be impractically large.

Use a hybrid when: a flexible human request must trigger exact operations against real systems. This is the default pattern for many useful SI applications.

Independent Exercise 1: Which Component Should Own It?

Task: calculate a tuition invoice total from fixed rates and recorded lesson counts. Should a language model own the arithmetic?

Answer

No. Traditional software or a calculator should compute the total. A model can explain the invoice or interpret a natural-language question about it.

Independent Exercise 2: Which Component Should Own It?

Task: classify a parent message as scheduling, payment, curriculum or feedback despite varied phrasing. Should this be only hand-written keywords?

Answer

A model or trained classifier can be useful because language varies. The output should still be constrained to allowed categories and evaluated on representative messages.

Independent Exercise 3: Hybrid Design

Task: “Move my lesson from Tuesday to Thursday if there is an available slot.” Divide the work.

Answer

The model interprets the request and relevant preferences. Scheduling software queries actual availability, validates the student identity and applies booking rules. If a change is permitted and authorised, the scheduling system performs it. The model reports the verified result.

Independent Exercise 4: Debugging

A system books the wrong Thursday slot even though the user’s intent was correctly extracted. Where should you investigate first?

Answer

Inspect scheduling logic, slot identity, timezone and tool arguments before changing the language model. The interpretation may be correct while the transactional step is wrong.

A Complete Hybrid Workflow: From Parent Message to Verified Schedule Change

A hybrid design becomes easiest to understand when we follow one request all the way through. Consider a fictional tuition schedule. A parent writes: “Can we move Thursday’s 5 PM lesson to sometime after 6 on Friday? Same tutor if possible.” The task contains fuzzy language, real availability, identity, booking rules and an external schedule change.

The language model first interprets the message. Candidate structured intent: current booking = Thursday 5 PM; preferred day = Friday; earliest acceptable time = after 6 PM; tutor preference = same tutor if possible. This is a learned language task because the request could be phrased in hundreds of ways.

Traditional scheduling software then checks the authenticated student account and identifies the actual Thursday booking. It queries Friday slots after 6 PM for the same tutor. Suppose 6:30 PM is free and 7:30 PM is occupied. The software returns a factual availability result.

The model can generate a response: “Friday at 6:30 PM with the same tutor is available. Shall I move the lesson?” If the workflow requires explicit approval before changing bookings, the system stops here. The model’s ability to see a suitable slot does not grant authority to move it.

After the parent approves, ordinary software performs the transaction. It verifies that the original booking is still present, that 6:30 PM is still available, and that the user is authorised to change the booking. A transaction updates the schedule. The system reads back the new booking and the model reports the confirmed result.

This one example contains nearly the entire article: fuzzy intent belongs to SI, authoritative availability belongs to the scheduling database, exact booking rules belong to software, permission belongs to the authenticated workflow, and natural-language explanation belongs to the model.

What Happens if We Give Every Responsibility to the Model?

Imagine replacing the schedule query with a prompt that includes last week’s timetable. The model may infer that Friday 6:30 PM is free, but the timetable can already be stale. It may calculate the time correctly yet book a slot that was taken ten minutes ago.

Next imagine letting the model “remember” the student ID from an earlier conversation without verifying the authenticated account. It could apply the correct scheduling logic to the wrong student. The failure is not about language fluency; it is about authoritative identity.

Finally imagine asking the model to update the schedule by generating SQL directly against a live database. A malformed or over-broad query could affect multiple rows. Even if the model usually behaves correctly, the database should expose narrow, validated operations instead of unrestricted write access.

The lesson is not that models are dangerous by nature. It is that flexible probabilistic components should not be made responsible for exact state transitions when ordinary software can enforce safer boundaries.

What Happens if We Give Every Responsibility to Traditional Rules?

Now remove the language model and build a form. The form asks: current lesson, preferred day, preferred time, tutor preference. This can work perfectly when users are willing to fill the fields.

But messages arrive by chat: “Friday evening is easier”, “same teacher if possible”, “not too late because of school next morning”. A rigid keyword parser may fail to translate these preferences into structured constraints.

Programmers can keep adding rules, but natural language is open-ended. The number of variants grows faster than the usefulness of a keyword list. Learned language models earn their place by compressing many surface variations into a reusable interpretation capability.

The best design may still include a form after interpretation. The model converts the message into candidate fields; the user sees and confirms them; traditional scheduling code executes the exact operation.

A Responsibility Matrix for Hybrid SI Systems

Human intent: expressed in language and clarified through conversation. Model responsibility: interpret intent, extract candidate parameters, explain choices, generate drafts. Software responsibility: authenticate, validate fields, query authoritative state, enforce rules, execute transactions and return verifiable results.

Database responsibility: preserve current records and constraints. Tool interface responsibility: expose narrow operations with typed arguments. Evaluation responsibility: test both model behaviour and deterministic software paths. Human responsibility: define policy, authorise consequential changes and review exceptional cases.

This matrix does not require every organisation to use the same architecture. Its value is diagnostic. When a failure occurs, the team can ask which responsibility was violated instead of calling the entire application “the AI”.

Unit Tests, Integration Tests and Evals Are Different

A unit test for the scheduling system might check that a function rejects a requested start time outside opening hours. The input and expected output are exact. A unit test for fee calculation can verify arithmetic to the cent.

An integration test might create a temporary booking, attempt a reschedule, then confirm that the old slot becomes free and the new slot becomes occupied. This tests multiple traditional components together.

An SI evaluation asks different questions: does the model correctly interpret “sometime after 6 on Friday”? Does it preserve the preference for the same tutor? Does it avoid inventing a booking when no slot exists? Multiple phrasings may be acceptable, so the evaluation criteria cannot always be exact string equality.

An end-to-end evaluation combines both. Send the natural-language request, let the system interpret it, query the schedule, ask for approval, execute the transaction and verify the resulting state. The end-to-end test proves the workflow under that scenario, not every possible input.

A Concrete Test Suite for the Hybrid Scheduling Example

Test 1: “Move Thursday 5 PM to Friday 6:30 PM.” Expected intent fields are exact. Test 2: “Friday evening, same tutor if possible.” Expected constraints include Friday, evening and same-tutor preference while leaving the exact time open.

Test 3: “Anytime after six except 7:30.” Expected exclusion = 7:30. Test 4: “Move it to tomorrow.” The system must resolve tomorrow from the current date and timezone, not guess.

Test 5: no available slot. The system should report unavailability and perhaps offer alternatives. It must not fabricate an opening. Test 6: slot becomes occupied between proposal and approval. The transaction should fail cleanly and the assistant should refresh availability.

Test 7: user is not authorised to modify the referenced booking. The model may understand the request perfectly; the backend should refuse the transaction. Test 8: tool response times out. The system must inspect booking state before retrying a state-changing operation.

These tests protect different layers. A language-model benchmark alone would not cover transaction safety. Backend unit tests alone would not cover ambiguous natural-language interpretation.

Traditional Software Gives Strong Guarantees When the Rules Are Known

A type checker can reject a string where an integer is required. A database can enforce a unique key. A payment system can require a valid currency. A scheduling service can forbid overlapping bookings. These guarantees are valuable precisely because they are explicit.

Language models are useful before and after these constraints: they help humans express intent and understand outcomes. The most dependable architecture often places hard guarantees around the flexible model rather than asking the model to simulate those guarantees through prose.

SI Gives Graceful Handling When the Rules Are Incomplete

Traditional software usually needs a defined branch for each recognised situation. A language model can still produce a useful response when the request is incomplete, such as “Friday evening works better.” It can identify missing information and ask the next question.

This makes SI valuable at interfaces. Humans rarely express requests as perfect API arguments. They explain goals, constraints, preferences and exceptions in uneven language. A model can translate that messiness into a structured conversation with the exact backend.

The Anti-Pattern of “AI All the Way Down”

One architectural mistake is routing every small operation through a large model: calculate a date, validate an email address, compare two IDs, parse a fixed code, check whether a number is within a range.

This adds latency, cost and probabilistic failure without gaining useful flexibility. If a rule can be expressed in a few lines of deterministic code and must be exact, use the code.

The model should be reserved for the parts where generalisation is valuable: language interpretation, synthesis, fuzzy classification, content generation or reasoning across unstructured evidence.

The Anti-Pattern of “Rules for Everything”

The opposite mistake is trying to solve human language with thousands of fragile regular expressions and keyword lists. The system becomes difficult to maintain and still misses ordinary paraphrases.

A model can replace a huge surface-level rule set with learned representations. But after it extracts the meaning, the application can return to exact software for the operation itself.

Schema Contracts Make the Boundary Explicit

Suppose the language model must output a reschedule request with fields booking_id, preferred_day, earliest_time, excluded_times and tutor_preference. A schema defines their types and allowed values.

The application rejects unknown fields or malformed values. If earliest_time is “after dinner”, the model may need to ask the user for a specific time or map the phrase under a documented convention. The schema makes ambiguity visible.

This is much stronger than passing free-form model text directly to the scheduling system.

Database Transactions Protect State

A scheduling change usually involves more than changing one field. The system must ensure the target slot is still free, release the old slot and reserve the new one without leaving an inconsistent intermediate state.

Databases provide transaction mechanisms for this kind of exact state change. The model can request the transaction, but the transaction system should own consistency.

This pattern generalises to orders, accounts, inventory, enrolments and many other domains. SI interprets and coordinates; transactional software commits authoritative changes.

Idempotency and Retry Safety Belong to the Software Contract

If a reschedule API times out, the assistant may not know whether the booking changed. Blindly repeating the request can create duplicate operations or confusing state.

The backend can support idempotency keys or status checks so repeated attempts do not create unintended duplicates. These mechanisms are properties of the software interface, not of the model’s intelligence.

A good SI agent knows how to use them because the tool description and workflow expose the contract. The safety comes from cooperation between model behaviour and traditional API design.

Logging: What Should Be Observable?

For deterministic software, logs often record requests, errors, database operations and latency. SI systems need additional observability: model version, prompt or instruction version, retrieved sources, tool calls, validation results and evaluation outcomes.

The goal is not to record private chain-of-thought. It is to preserve operational evidence that explains what the system received and did. This makes debugging possible without turning the model’s generated self-description into the source of truth.

A Failure-Triage Table

Wrong intent extracted: inspect model, prompt, examples and input representation. Correct intent, wrong database row: inspect identity mapping and query code. Correct row, wrong calculation: inspect deterministic function or tool arguments. Correct calculation, unauthorised write: inspect permissions and transaction boundary.

Correct write, wrong explanation: inspect model context and output verification. Correct explanation, stale source: inspect retrieval or authoritative data. Intermittent duplicate actions: inspect idempotency and retry handling. This table turns “AI bug” into a useful starting diagnosis.

Changing Rules: Who Gets Updated?

Suppose the library loan period changes from 14 to 21 days. The authoritative rule engine or configuration must change. The documentation source must change. Evaluation cases must change. Cached or retrieved old policy should be retired or clearly marked archived.

The foundation model may not need retraining at all if the current rule is supplied at runtime. This is a major advantage of separating business policy from learned weights.

If the model has memorised the old rule and sometimes ignores the new source, then model behaviour becomes part of the repair. But updating the policy should not depend on waiting for a global model retraining cycle.

Changing Language: Who Gets Updated?

Now the rule stays the same, but users adopt new slang or multilingual phrasing. The backend does not need a new due-date formula. The language layer may need better examples, a stronger model or updated routing.

This shows why maintenance ownership should be split. Rule owners maintain policy. ML owners maintain interpretation. Application owners maintain interfaces and transactions.

Performance Engineering: Route Cheap Work Away From the Model

A model call may be orders of magnitude more expensive than checking whether two IDs are equal or adding two numbers. A high-volume application benefits from routing exact low-level work to ordinary code.

The model can also reduce its own workload by requesting only relevant database fields rather than receiving an entire record dump. Traditional query systems can filter first; the model interprets the smaller result.

This improves cost, latency and privacy simultaneously because less irrelevant information enters the model context.

Traditional Software Can Also Be Wrong

It is important not to romanticise explicit code. Programmers make mistakes. Requirements are misunderstood. Edge cases are omitted. Databases contain bad data. A deterministic bug can fail the same way for every user.

The difference is not “software correct, AI unreliable”. The difference is failure shape. Traditional bugs often reproduce under defined conditions; model failures can be more statistical and input-sensitive. Both need testing and monitoring.

SI Can Help Maintain Traditional Software

The relationship is not one-way. Language models can read logs, explain code, generate tests, propose refactors and help developers navigate large codebases. Traditional software remains the deployed mechanism while SI accelerates human maintenance.

This creates a recursive toolchain: models help write code that constrains models. The quality of the overall system still depends on review and testing.

A Migration Strategy for Existing Software

Do not begin by replacing the whole application. Identify one interface where language variation causes human work. Add an SI layer that produces structured candidate inputs for the existing backend.

Keep the backend rules and transactions unchanged initially. Compare extracted fields with human-entered fields. Measure correction rate and failure types. Only automate the handoff after the model behaviour is understood.

Then add additional capabilities one at a time: retrieval, drafting, tool execution, memory. Each expansion gets its own acceptance tests and permission boundary.

A Migration Strategy for Prototype AI Apps

The opposite situation is a prototype where the model currently does everything in prompts. Start extracting hard rules into software. Move arithmetic to tools. Move authoritative state to a database. Add schema validation and narrow action APIs.

This reduces prompt complexity and makes the remaining model task clearer. A smaller, more focused context can improve both reliability and maintainability.

Worked Exercise: Redesign an “AI-Only” Expense Assistant

Prototype behaviour: user writes an expense description, model guesses category, calculates tax, decides whether policy permits reimbursement and writes “approved” into a spreadsheet.

Redesign: model extracts merchant, amount, date and candidate category. Validation checks types. Tax calculation uses explicit rules. Policy engine checks documented limits. Approval remains with the authorised person. Spreadsheet update uses a narrow transaction tool after approval.

The model still adds substantial value by reading messy receipts and descriptions. The exact financial and permission rules move into components that can be tested directly.

Independent Exercise 5: Find the Wrong Owner

A model correctly extracts a date as 12 October, but the scheduling API interprets it in the wrong timezone and books the previous day. Which component owns the first repair?

Answer

The first repair belongs in the scheduling or date-handling software because the extracted date was correct and the timezone conversion changed the meaning afterwards.

Independent Exercise 6: Find the Wrong Owner

A database returns the correct policy row, but the model explains the opposite rule. Which component should you isolate?

Answer

Isolate the model interpretation and context. Supply the row directly in a controlled evaluation and test whether the model can state the rule accurately.

Independent Exercise 7: Find the Wrong Owner

A model generates an invalid room code “LAB-9” even though only LAB-1 to LAB-4 exist. What should prevent the write?

Answer

Traditional validation should restrict allowed room codes. The model can be asked to correct the value, but the backend should reject invalid codes regardless.

Independent Exercise 8: Decide Whether SI Is Needed

Task: compute a 7% tax on a known subtotal and round to two decimals. Do you need an LLM?

Answer

No. A deterministic calculation is simpler, cheaper and more exact. An LLM may explain the calculation if the user asks, but it is not necessary for the computation itself.


Deep Worked Build: One Parent Request Through a Hybrid SI System

To reach the Clementi floor, the distinction between SI and traditional software has to survive a complete task, not only a list of component names. Consider this fictional request: “My daughter is Secondary 3, I paid yesterday, and I would like to register her for Saturday’s Mathematics workshop if a seat is still available. Please confirm the booking.”

The application receives one natural-language request containing identity, level, a claimed payment state, an event, a conditional request and permission to complete a booking. A language model is useful because the message is flexible. But no part of the sentence should be treated as authoritative proof that payment cleared, that a seat remains, or that the student record really belongs to Secondary 3.

Step 1: SI interprets intent and candidate fields

The model can extract candidate intent = register for workshop; stated level = Secondary 3; claimed payment = paid yesterday; requested action = confirm booking if available. This is a linguistic interpretation. The model should preserve uncertainty around fields that need a system-of-record check.

A structured candidate might be represented as student_reference = supplied name or account identity; workshop = Saturday Mathematics; intent = register; stated_level = Secondary 3; claimed_payment = yes; conditional = seat available. Conventional validation can then check whether required fields are present before the system proceeds.

Step 2: Traditional software checks authoritative state

The student database returns official level = Secondary 3. The payment system returns status = settled. The workshop database returns capacity = 20 and confirmed bookings = 19. The current timestamp is before the registration deadline. These are authoritative state checks, not language-model guesses.

Every check can be tested independently. The eligibility rule accepts Secondary 3. The payment rule accepts settled. The capacity rule sees one remaining seat. The deadline rule remains open. None of these decisions needs a generative model once the correct structured fields and current records are available.

Step 3: A transaction protects the final seat

The dangerous moment is not the explanation. It is the state change. Between reading “19 bookings” and writing the new booking, another request could take the last seat. A dependable booking system therefore needs an atomic transaction, conditional write or equivalent concurrency control rather than trusting a previously read number.

The booking operation succeeds and returns booking ID MATH-SAT-020. The database now records 20 confirmed bookings. That returned state is the evidence that the booking exists. If the transaction fails because another request took the seat first, the assistant must report the failure or alternative process rather than repeat the earlier availability claim.

Step 4: SI turns verified state into a human answer

Only after the transaction succeeds should the model produce language such as: “Your daughter’s registration is confirmed for the Saturday Mathematics workshop. Payment is recorded as settled and booking MATH-SAT-020 has been created.” The model makes the result readable; the database and transaction establish that the result happened.

This four-step flow—interpret, validate, transact, explain—is a reusable hybrid pattern. Google’s current Production ML Systems material likewise treats the model as one component inside a larger production system whose pipelines, serving and monitoring matter to the outcome.

Run the Same Request Through Four Broken Architectures

Broken architecture A: model-only. The model reads the parent’s claim “I paid yesterday” and assumes payment is complete. It also assumes a seat remains because the message asks “if a seat is available”. The system can produce a beautiful confirmation with no authoritative state. The repair is not better prose; it is connection to the appropriate records and transaction.

Broken architecture B: rules-only language parser. A keyword rule sees “paid” and “register” but cannot distinguish “I paid yesterday” from “I tried to pay yesterday but it failed”. The system routes both to confirmation. The repair can be a richer parser or learned language model before deterministic validation.

Broken architecture C: model plus database but no transaction. Two users simultaneously see the final seat and both are told it is available. Separate writes then create an over-capacity state. Interpretation and database lookup worked; concurrency control failed.

Broken architecture D: correct transaction, wrong completion message. The booking fails because the final seat is taken, but the language layer reuses a success template and says “confirmed”. The state is correct while reporting is wrong. This is why completion language should be generated from observed operation results rather than from the model’s expectation.

A Concrete Responsibility Matrix

Natural-language understanding belongs primarily to the model layer. Required-field and type validation belongs to conventional code. Eligibility and deadline rules belong to explicit policy enforcement. Payment and booking state belong to authoritative databases. Concurrency belongs to transaction logic. User-facing explanation can return to the model.

Security spans the entire system. The model should only see the minimum data necessary for the task. The application should authenticate the user. The database should enforce access controls. The booking service should reject unauthorised changes even if the model asks for them.

Evaluation also spans the system. Model tests cover intent and field extraction. Unit tests cover policy functions. Integration tests cover database and transaction behaviour. End-to-end tests start from a parent message and inspect the final booking state and response.

When to Prefer Traditional Software First

Prefer explicit software when the rule is stable, enumerable and consequential: permission checks, deadlines, numeric bounds, uniqueness constraints, state transitions, money arithmetic, cryptographic operations and database integrity. These tasks benefit from exact semantics and deterministic tests.

Even here, SI can help developers write tests, explain rules or diagnose failures. Assistance with the software is different from replacing the software’s runtime responsibility with an LLM call.

When to Prefer an SI Layer First

Prefer a learned model when the input variation is difficult to enumerate: free-form language, document interpretation, fuzzy classification, summarisation, extraction from inconsistent wording, draft generation and semantic matching. These tasks benefit from learned representations and broad training.

Then surround the model with explicit validation and state systems. A flexible interpretation can produce a candidate structure; conventional software decides whether the candidate is syntactically valid, permitted and consistent with authoritative records.

The Strong Hybrid Principle

The strong hybrid principle is: let SI handle ambiguity; let software enforce invariants; let authoritative stores own state; let tools perform bounded operations; verify the resulting state before reporting completion. This is not a universal architecture diagram, but it is a useful default for deciding where responsibilities belong.

Once this principle becomes intuitive, “AI versus software” stops being a useful argument. The real engineering question becomes: which mechanism gives this part of the task the clearest specification, strongest evidence and most repairable failure mode?

Frequently Asked Questions About SI versus Traditional Software

Is traditional software deterministic?

Often, but not always. Traditional programs can use randomness, concurrency and external state. The key distinction is that their behaviour is primarily defined through explicit code and rules rather than learned model parameters.

Is SI better than normal software?

No universal ranking is meaningful. Each is better suited to different responsibilities. Flexible interpretation favours learned models; exact invariants and transactions favour explicit software.

Will SI replace programming?

SI can help write, explain and maintain code, but deployed systems still require data structures, interfaces, security boundaries, tests and operations. Programming changes rather than simply disappearing.

Why do AI systems still need databases?

Because databases preserve authoritative current state and support exact queries, transactions and constraints. Model parameters are not a replacement for transactional records.

Why do agents need ordinary code?

Agents need software to expose tools, validate calls, maintain state, enforce permissions, schedule work and observe results. Model-directed planning is one layer inside a larger application.

Can rules and models disagree?

Yes. The model may infer one intent while a rule engine determines that the requested action is not allowed. The system should preserve the distinction and explain the binding rule.

What should be tested after a model upgrade?

Repeat model-evaluation cases and end-to-end workflows. A new model may change extracted fields, tool selection or output formatting even when traditional backend code is unchanged.

What should be tested after a software change?

Repeat deterministic tests and SI integration cases. A new API schema or validation rule can break a previously working model-tool handoff.

The Strongest SI Systems Use the Right Kind of Computation for Each Job

Traditional software gives us explicit rules, authoritative state, transactions and strong invariants. SI gives us learned flexibility across language, images, code and unstructured information. Neither eliminates the other.

The architectural skill is deciding where to draw the boundary. Let models interpret what is fuzzy. Let software enforce what is exact. Connect them with validated interfaces, explicit permissions and observable outcomes.

Continue through the How Super Intelligence Works hub. Previous: 006 — Why Language Models Can Do More Than Language. Next: 008 — SI versus Search Engines.

Discover more from eduKate Singapore

Subscribe now to keep reading and get access to the full archive.

Continue reading