VIEW THIS AS

Auto mode follows the Route Engine until you choose a viewpoint.

YOU ARE HERE

ROUTE CHECK

CONNECTED TO

WHAT NEXT

Use the canonical route for this room, or HELP if you are unsure.

Super Intelligence | Productivity: Why Better Technology Does Not Instantly Transform an Economy

eduKate Secondary students reviewing open books for How Super Intelligence Works: the SI Failure Map.

A reader’s guide to AI, productivity and the journey from faster tasks to a stronger economy.

Follow the output. Count the resources. Test the bridge from a demonstration to the whole economy.

Three students studying together with open books at a classroom table
Learning to connect technology claims with evidence and economic reasoning.

A technology can become much better at a task before an economy becomes much better at producing useful goods and services. Between those two events sit working practices, complementary equipment, skills, demand, investment, prices and measurement. Understanding that distance helps us interpret impressive demonstrations without dismissing real progress or promising a transformation that has not yet occurred.

This guide follows the distance step by step. First we distinguish the quantities being measured. Then a completely fictional print workshop shows why a large improvement in document preparation becomes a smaller improvement in completed output. A second fictional packet builds an entire three-sector economy, with enough information to reproduce every calculation. The later exercises change the inputs so that you can test whether you understand the mechanism rather than remember a conclusion.

In this series, “Super Intelligence” is an editorial umbrella for studying AI and its possible implications. Artificial superintelligence, or ASI, means a hypothetical broadly superhuman capability. The studies discussed here concern particular AI systems and observed work, not an economy operating with demonstrated ASI. A separate scenario chapter asks what stronger capabilities could change. It supplies assumptions rather than an arrival date or a forecast disguised as arithmetic.

1. Choose the output before calculating productivity

Productivity is a relationship between an output and the inputs used to produce it. The question “How much faster is the technology?” is incomplete until we know what it is doing and what resources the comparison includes. A model may generate twice as many draft paragraphs while a team completes the same number of acceptable reports. The paragraphs and the reports are different outputs. Measuring the first cannot settle what happened to the second.

For a narrow, uniform activity, we can count accepted units per labour hour. If a workshop finishes 100 comparable packs in 100 hours, its measured rate is one pack per hour. If it finishes 120 equally acceptable packs in the same hours, the rate rises by 20%. This is useful operational evidence. It does not yet tell us whether more machinery, purchased services or energy were needed, or whether the packs are economically valuable to customers.

Official labour-productivity measures use real output relative to labour input. Total factor productivity, also called multifactor productivity in some contexts, compares output with a combination of measured inputs. The precise output and input coverage varies by statistical series. These distinctions are explained in the U.S. Bureau of Labor Statistics overview of productivity. A labour-productivity increase may reflect more capital available to workers; it should not automatically be described as an equal improvement in the efficiency of every resource.

“Real” matters because higher prices can raise money receipts without increasing the volume or quality of production. If the workshop sells the same packs at a higher price, revenue per hour rises. A calculation that holds prices constant would not treat that price increase as more physical output. For diverse goods and services, statistical agencies use methods for constructing output-volume measures rather than simply adding unlike items. Ten consultations plus ten bicycles are not twenty interchangeable units of national production.

We also need an appropriate boundary. A company’s revenue includes what it paid other companies to supply. Summing all company revenues can count intermediate goods repeatedly. Value added removes intermediate inputs from gross output, as explained by the U.S. Bureau of Economic Analysis. Our economy packet will supply real value added directly, so that readers do not have to invent price indices or accidentally count the same component twice.

A good opening sentence therefore names the measure: accepted packs per actual staff hour, real value added per hour, or output relative to a specified combined-input index. The discipline is simple but powerful. It prevents a discussion from beginning with one numerator and ending with another while calling everything productivity.

Back to contents · Next: 2. There are several bridges between a benchmark and an economy

2. There are several bridges between a benchmark and an economy

Imagine an assistant that prepares a correct page layout in five minutes rather than twenty. That is a result about one activity under the tested conditions. To reach an economic conclusion, we need to ask whether ordinary employees can reproduce the result, how much of their work resembles the test, and what happens before and after page preparation. A benchmark can be excellent evidence about the benchmark while leaving those questions unanswered.

The first bridge is from capability to use. Suitable tools must be available, affordable and usable with the inputs that actually arrive. A system that handles clean documents may be less useful when files are inconsistent or incomplete. The second bridge is from use to an improved workflow. Employees may still need to collect information, resolve contradictions, inspect outputs and coordinate the next stage. These activities can absorb part of the apparent saving.

The third bridge is from a better workflow to a better firm-level outcome. A printing machine, delivery schedule or shortage of demand may limit completed orders. Released time can be used well, left idle or consumed by a new reporting burden. The fourth bridge is from a successful firm to its sector. A firm winning customers from rivals does not by itself prove that the entire market produces more. The fifth bridge is from a sector to the economy, where other sectors and the resources needed to supply the technology must be included.

These bridges are not five compulsory waiting periods. Sometimes several change together. A new service might arrive with suitable software, existing distribution and strong demand, producing gains quickly. In another setting, a small technical change may require extensive institutional adjustment. The correct inference is conditional: a demonstrated gain reaches a wider boundary when the relevant connections also work. There is no universal delay built into the word technology.

Nor are the bridges an argument that capability is unimportant. Better capability can reduce review, broaden the set of suitable tasks and make previously impossible products feasible. It can change the constraints themselves. The point is to locate that contribution precisely. If better reasoning removes a costly checking stage, show the new checking evidence. If easier integration allows smaller firms to adopt, measure their adoption and results rather than continuing to use a demonstration from a specialist team.

The same logic applies in reverse. Weak national productivity over a short period does not prove that no individual AI application is useful. A small successful application can be economically real yet too limited, too recent or too offset by other developments to dominate an aggregate statistic. Local success and modest aggregate change can coexist without either observation being fraudulent.

Back to contents · Next: 3. Complementary investment turns an available tool into a useful method

3. Complementary investment turns an available tool into a useful method

A complement is something that makes another input more useful. A bicycle works better with a safe route, repair skills and a place to park it. In an organisation, an AI system may need coherent records, a suitable interface, trained staff, reliable equipment and a redesigned sequence of work. Buying access to the model purchases one component of this arrangement. It does not automatically purchase the arrangement itself.

Consider a workshop with three versions of each customer’s specification. Faster document preparation could amplify confusion unless someone establishes which version controls. That improvement in record organisation may require no new model capability, yet it changes what the existing tool can accomplish. Similarly, a staff member who understands the print process can recognise an impossible layout requirement. Their domain knowledge complements the assistant’s ability to prepare alternatives quickly.

Complementarity also explains why copying a successful firm’s software purchase may not reproduce its success. The original firm might have spent years standardising records, coordinating suppliers and training employees. A case study that reports only the final subscription price leaves those resources out of view. The apparent miracle can partly be the return on investments made earlier, including investments that the firm itself now takes for granted.

The 2015 OECD research on frontier firms, technology diffusion and public policy documents uneven diffusion and the importance of adaptation between frontier technologies and firms using them. This is evidence about differences and diffusion in the economies studied, not a measured adoption schedule for hypothetical ASI. Its useful lesson here is to examine the route by which a method becomes usable elsewhere.

Not every complement is expensive or slow. A shared template, a clarified field definition or the removal of an unnecessary handoff can be cheap. Others require major investment, such as a power connection or production equipment. The right response is to identify the actual constraint, estimate what would relax it and test the result. Calling every difficulty “culture” or every delay “regulation” conceals mechanisms that might have very different remedies.

There is also a sequencing problem. Training staff before the workflow is settled can teach a process that soon changes. Buying equipment before demand is understood can leave spare capacity unused. Waiting for every uncertainty to disappear can forgo useful learning. A bounded trial can reveal which complements matter, but the trial must record their costs and outcomes. Otherwise the firm may repeatedly rediscover the same hidden work while describing each attempt as effortless adoption.

Back to contents · Next: 4. Read current AI evidence at its actual scale

4. Read current AI evidence at its actual scale

A useful empirical study identifies a technology, a population, a task, an outcome and a comparison. It may support a strong causal claim within that setting while supporting a much weaker claim about another setting. The distinction is especially important when a study’s striking percentage travels through headlines without its denominator or date.

In the research version of Generative AI at Work, Brynjolfsson, Li and Raymond examine the staggered introduction of an assistant using data from 5,172 customer-support agents. They report a 15% average increase in issues resolved per hour, with substantial variation across workers. This is an important finding about the studied support setting. It is not a 15% increase in national GDP or an experiment on superintelligence. The version matters: earlier summaries can report different numbers from earlier versions of the research.

Evidence can also reveal costs that demonstrations miss. METR’s early-2025 study of experienced open-source developers randomised AI access for 246 tasks undertaken by 16 developers working in familiar projects. It found that AI-allowed tasks took 19% longer in that setting. This does not establish that all developers, all tasks or all later tools become slower. It establishes why real completion time deserves measurement alongside users’ impressions.

The authors subsequently reported that their later experiment had serious selection and measurement problems. Their February 2026 update says the newer data provide an unreliable signal of the current productivity effect and describes why developers and tasks selecting out of the experiment could bias results. That update matters when interpreting the earlier study. The responsible summary preserves the dated finding and the later uncertainty rather than freezing the technology at its first measured result.

These studies should not be averaged into a universal AI multiplier. Their workers, tools, tasks, designs and output measures differ. A positive effect in a support centre and a negative effect in familiar open-source repositories can both be informative. Together they motivate a question: which features of the work explain where assistance helps, and which measurements would reveal those features in another setting?

For a new claim, read beyond the average. Does the result include checking and correction? Were unsuccessful attempts counted? Did quality change? Are experienced and novice workers affected differently? Was the tool optional, randomly assigned or adopted by volunteers? Did the study compare equivalent tasks? The answers determine what travels to a new application. They also prevent a research result from being used as permission to assume away the costs that the next organisation will actually face.

This guide uses empirical findings to motivate careful measurement. It does not use them to set the fictional parameters below. Those parameters are supplied teaching assumptions, deliberately separate from the studies, so that a reader can reproduce the examples without believing that any real company achieved the same outcome.

Back to contents · Next: 5. A productivity level is different from a growth rate

5. A productivity level is different from a growth rate

Suppose productivity is indexed to 100 before a change and rises to 110 after it. That is a 10% increase in the level. If it then stays at 110, productivity does not grow by another 10% in the next period. The improvement remains valuable, but it is not a permanently higher growth rate. Confusing a one-time level gain with a recurring annual gain can make a modest scenario appear explosive.

A second economy might start at 100 and grow by 2% each year. After one year it reaches 102, after two years 104.04, and after five years approximately 110.408. Compounding means the same percentage is applied to a changing base. The total five-year increase is about 10.408%, not exactly 10%. These arithmetic distinctions become important when someone compares a task study conducted over weeks with a forecast stated as additional annual national growth.

Time saved and output per hour also use different denominators. Reducing a fixed task from 100 minutes to 80 minutes saves 20% of the original time. At unchanged quality, output per minute rises from one-hundredth to one-eightieth of a task, an increase of 25%. Neither figure is wrong. Calling both a 20% productivity improvement would obscure which ratio was computed.

Likewise, saying an activity is four times as fast means its time per unit falls to one quarter. It does not mean the whole organisation becomes four times as productive if the activity occupies only part of its work. In the workshop example, a fourfold improvement in preparation affects a category accounting for 30% of baseline staff hours. We will then add checking and maintenance rather than treating them as invisible.

Finally, distinguish the transition from the eventual operating state. Installing a new process can use resources now while creating benefits later. A monthly comparison may record the installation burden; an operating trial may omit it. Both can be useful if clearly labelled. A fair assessment asks whether the relevant benefits over time justify the relevant resources over time, while admitting uncertainty about how long benefits last. An appealing eventual ratio is not evidence that the transition has already paid for itself.

Back to contents · Next: 6. The supplied workshop packet: Alder Print

6. The supplied workshop packet: Alder Print

Alder Print is entirely fictional. It makes standardised booklet packs for local organisations. The example does not describe eduKate, a real supplier or a recommended business investment. Every number below is a teaching input. A pack counts as accepted only when its pages match the approved order, text is legible, binding is secure, quantities are correct and it reaches the dispatch stage without an unresolved defect. All central-scenario packs are assumed comparable in size and difficulty.

In the baseline month, Alder completes 1,000 accepted packs using 1,000 actual staff hours. Preparation uses 300 hours, ordinary checking uses 100, printing and binding use 500, and dispatch uses 100. These categories are mutually exclusive and sum to the total. Each scales proportionately with accepted packs within the range modelled. Baseline staff time is therefore 0.300, 0.100, 0.500 and 0.100 hours per pack respectively, or one hour altogether.

An AI-assisted preparation process reduces preparation time by 75%, from 0.300 to 0.075 hours per pack. Checking rises from 0.100 to 0.160 hours because staff inspect generated layouts and correct routine mistakes. This new checking figure includes the existing checking work; do not add the old 0.100 again. Printing and binding remain at 0.500, and dispatch remains at 0.100. Operating the new system also requires 30 fixed staff hours per month for template maintenance and evaluation.

All packs use the new process in the central scenario. The system has adequate digital capacity. Quality is unchanged unless a later exercise explicitly changes it. Staff can perform the stated tasks in the required proportions, so the first calculation deliberately ignores a qualification bottleneck. The workshop has 1,000 actual staff hours available for these operations and associated improvement work. Its physical printing-and-binding equipment can handle at most 1,080 comparable accepted packs per month under the stated operating conditions.

Demand is 1,000 packs in the initial comparison and 1,060 in the mature-month scenario. Demand above capacity is not delivered in advance or borrowed from a later month. There is no inventory change, outsourcing or double counting of packs. Fractional capacity calculations describe thresholds; actual finished packs must be whole numbers, so a capacity threshold is rounded down when necessary. Unless otherwise stated, money, wages, employment changes and environmental effects are outside this packet.

The information is intentionally sufficient to solve the example without asking an assistant to invent missing operating facts. Copy the four baseline categories, four assisted categories, fixed hours, staff availability, physical capacity and demand into your own notes. Before reading on, predict which constraint will bind. The prediction is useful because the later arithmetic can correct a specific intuition rather than simply introduce an impressive number.

Back to contents · Next: 7. Solve the whole workflow, not just the accelerated task

7. Solve the whole workflow, not just the accelerated task

The assisted variable time per accepted pack is 0.075 + 0.160 + 0.500 + 0.100 = 0.835 staff hours. Including fixed operating work, required hours for Q accepted packs are H = 0.835Q + 30. This equation applies while the process is operating and while the proportional assumptions hold. If Alder never starts the system, the baseline remains H = Q and the additional 30 hours do not appear.

At 1,000 packs, the assisted process requires 0.835×1,000 + 30 = 865 hours. The independent ledger route gives the same answer: preparation falls from 300 to 75, checking rises from 100 to 160, printing and binding stay at 500, dispatch stays at 100, and fixed work adds 30. Adding 75 + 160 + 500 + 100 + 30 again gives 865. Two equivalent calculations are a useful safeguard against overlooking a category.

Net required hours fall by 135, or 13.5% of baseline. The preparation category alone saves 225 hours, but 60 are used by extra checking and 30 by fixed work. Reporting the full 225 hours as a net saving would overstate the effect. Reporting a 75% reduction in the entire workflow would be a much larger error: only preparation received that reduction.

If the denominator is the 865 hours required by this operating process, productivity is 1,000/865, approximately 1.15607 accepted packs per hour. Relative to the baseline rate of one, that is a 15.607% increase. The 13.5% time saving and 15.607% productivity increase are consistent because their denominators differ. Keep more digits during calculation and round only the final presentation.

But required operating hours are not automatically observed total firm hours. Suppose Alder still records 1,000 actual staff hours and delivers the same 1,000 packs, using released time on training and improving its records. Its pack-output-per-total-hour ratio remains one. The operating method has created capacity, while this month’s narrow output ratio has not increased. Training may be valuable, but we must measure or describe that value separately rather than pretending additional packs were already delivered.

This distinction prevents a common argument from going in circles. The operating trial can truthfully show reduced labour requirements, and the firm-level monthly ratio can truthfully show no change. They measure different denominators and activities. To resolve the apparent disagreement, reconcile the time ledger and say what happened to released hours. Do not choose whichever denominator produces the most flattering percentage.

Back to contents · Next: 8. Demand and physical capacity decide how much potential becomes output

8. Demand and physical capacity decide how much potential becomes output

Under an idealised labour-only calculation, 1,000 staff hours can support Q = (1,000 − 30)/0.835, approximately 1,161.677 packs. Alder could complete at most 1,161 whole packs within that labour budget if no other constraint mattered. This is a capacity result, not a demand forecast. It says nothing about whether customers want those packs or whether the equipment can produce them.

The physical limit is 1,080, and mature-month demand is 1,060. Actual completed output is limited by the smallest of labour capacity, physical capacity and demand: 1,060 packs. The assisted operating process requires 0.835×1,060 + 30 = 915.1 hours. Alder still records 1,000 actual staff hours, with the remaining 84.9 allocated to the improvement work specified in the packet. Accepted packs per total staff hour rise from one to 1.06, a 6% increase.

This result is smaller than the 15.607% operating-productivity increase at fixed output, and much smaller than the 75% preparation-time reduction. Nothing has gone wrong with the arithmetic. Each number describes a different point in the chain. The tool accelerates one task, the workflow needs fewer hours per pack, and the firm delivers more packs only to the extent that customers and production resources support them.

Now increase demand to 1,160 while holding everything else constant. The equipment becomes binding, so output stops at 1,080. Required operating hours are 0.835×1,080 + 30 = 931.8. The pack-per-total-hour ratio becomes 1.08. Asking the preparation assistant to work even faster would not raise completed output under these constraints. The next relevant question is whether increasing physical capacity would be worthwhile and feasible, not whether the digital task can win another speed benchmark.

Bottlenecks can move after an improvement. If equipment capacity later rises, demand or specialised labour might become the limit. There is no reason to expect one permanent bottleneck. Nor should every unused minute be treated as waste: slack can support maintenance, learning and resilience. The analytical requirement is to name its purpose and avoid counting the same time as both released capacity and extra completed output.

A queue is another possible outcome. If the front end accepts more work than the physical process can finish, unfinished orders accumulate. A dashboard counting newly prepared layouts may celebrate growth while customers wait longer. In this packet, unfinished orders are explicitly excluded from accepted output. That definition keeps the result attached to what customers receive rather than the speed of an intermediate stage.

Back to contents · Next: 9. Quality and outsourcing can change the apparent gain

9. Quality and outsourcing can change the apparent gain

A faster process is comparable only if the relevant output remains comparable. Suppose a different assisted trial records 1,000 attempted packs and 865 staff hours, but 50 packs fail the acceptance standard and remain unresolved at month end. There are 950 accepted packs, not 1,000. Accepted output per hour is 950/865, approximately 1.09827, which is a 9.827% increase over baseline. Attempted output would exaggerate what was delivered.

Now suppose those 50 packs are repaired during the same month using 70 additional staff hours, with no other changes. All 1,000 are eventually accepted, but total hours become 935. The ratio is 1,000/935, approximately 1.06952, or a 6.952% improvement. The failed intermediate units must not be counted once when first attempted and again when repaired. Nor should the extra repair work vanish because a different team performed it.

These are alternative quality scenarios, not additions to the central packet. The central assisted checking allowance already contains routine correction under its stated unchanged-quality assumption. The extra 70 hours are introduced only in the changed scenario. Keeping scenario boundaries explicit prevents double counting both failures and improvements. It also makes the calculation reproducible for a reader who has not watched the original trial.

Outsourcing creates a related boundary problem. Suppose Alder transfers dispatch to a supplier. Its own staff-hour denominator falls, which may raise its measured labour-productivity ratio. Yet the dispatch work has not necessarily disappeared from the economy. The supplier uses labour and other resources. To assess the whole production chain, include the supplier’s contribution and use an appropriate value-added boundary, rather than declaring that fewer hours on Alder’s payroll prove an equal economy-wide efficiency gain.

Outsourcing can genuinely improve efficiency if the supplier performs the work more effectively. The lesson is not that outsourcing gains are fictitious. The lesson is that a change in organisational boundary and a change in underlying resource use are separate possibilities. Evidence about the supplier determines which explanation applies. A firm-level ratio alone cannot distinguish them.

Quality can also improve in ways that pack counts miss. A clearer layout may help readers, or better checking may prevent an expensive mistake. Those benefits deserve evidence and an appropriate outcome measure. They should not be forced into an arbitrary “quality multiplier” chosen after seeing the results. Define the dimension, establish how it is observed and show the comparison. Honest measurement can recognise value without manufacturing precision.

Back to contents · Next: 10. Investment can make the first months look worse

10. Investment can make the first months look worse

Alder’s mature operating equation leaves out one-time implementation. Add a transition month with 160 staff hours for training and 120 for record cleanup and template conversion. These 280 hours are genuinely used during the month. Keep total available staff hours at 1,000, include the recurring 30 hours, and retain the assisted variable requirement of 0.835 hours per accepted pack. Demand and equipment are ample for this reduced production level.

The available capacity for variable production is 1,000 − 280 − 30 = 690 hours. Dividing by 0.835 gives approximately 826.347 packs, so at most 826 whole packs are completed. Their variable work uses 689.71 hours; adding implementation and recurring work gives 999.71, leaving 0.29 hours. The narrow pack-output-per-total-hour ratio is approximately 0.826 if all 1,000 hours are recorded. It is below the baseline even though the assisted operating method is faster.

This transition is not automatically a failure. It may be a deliberate investment in a better later process. It is also not automatically a success waiting to happen. The training could be ineffective, the records could become obsolete, or the tool could fail to deliver the expected benefits. A credible investment story specifies what reusable capability is being created and what later evidence would show that it is useful.

The research on the productivity J-curve by Brynjolfsson, Rock and Syverson adds a particular measurement mechanism. When complementary intangible investment is not fully measured as output, conventional productivity can understate growth during its creation. Later, services from omitted intangible capital can make measured productivity growth overstate growth relative to a more comprehensive accounting treatment. This is more specific than the slogan that every technology must dip before succeeding.

Alder’s pack-count example illustrates a real allocation of staff time during a transition. It is not a calculation of the paper’s TFP measurement gap. We have not priced an intangible asset, estimated depreciation or measured its capital services. Calling all 280 hours an asset worth a chosen amount would not establish its economic value. Some implementation work may build durable capability; some may simply correct avoidable problems.

The practical conclusion is to keep two views. One records current production and resources actually used. The other records the proposed investment, its expected useful life and evidence that it improves later work. Neither view should erase the other. A long-term story should not conceal current costs, and a single transition month should not be mistaken for a complete evaluation of a process intended to operate for years.

Back to contents · Next: 11. The supplied economy packet: Mere Island

11. The supplied economy packet: Mere Island

Mere Island is a fictional closed teaching economy with three sectors. The name is not a substitute for Singapore or any other country. The units are deliberately small enough to calculate by hand. Output is already supplied as real value added in millions of constant-price currency units, abbreviated CU. Labour is supplied as millions of actual hours. We assume comparable quality, unchanged base prices and additive sector outputs for this exercise. Real statistical systems can use more complicated volume measures.

The digital-services sector produces CU40 million of real value added using 0.4 million hours. Goods production contributes CU35 million using 0.7 million hours. Personal services contribute CU25 million using 0.5 million hours. Baseline totals are CU100 million and 1.6 million hours. Digital services therefore produce CU100 per hour, while each other sector produces CU50 per hour. These are accounting inputs, not judgments about the social importance or effort of the people doing different work.

Within digital services, adopting firms initially account for exactly half the sector’s real value added and half its hours. They are identical to non-adopters in baseline output per hour. Adopters achieve a 20% increase in real value added while keeping their hours unchanged. Non-adopters remain unchanged. Goods production and personal services also remain unchanged. There are initially no spillovers, supplier changes, entry, exit, labour movements or implementation hours outside the supplied totals.

That last sentence is intentionally strong. It gives us an exact small accounting exercise. It is not a description of how a real economy would remain frozen while a technology changes. Once we understand the simple case, we will relax particular assumptions one at a time. A scenario is easier to interpret when a changed assumption has a visible place in the ledger.

Keep adoption shares attached to their denominator. Half of firms would not necessarily mean half of sector output. A single large adopter could account for most of production; many small adopters could account for little. Here the packet explicitly supplies equal output and hour shares, so multiplication is legitimate. If that information were missing, the correct response would be to request or estimate it transparently, not silently assume every firm has the same weight.

The packet also supplies no forecast probability. The 20% gain, 50% adoption share and sector composition are chosen inputs. They do not come from the 15% support-agent study, and they do not represent an official prediction. Their purpose is to expose how the same local percentage changes when it is carried into a larger denominator.

Back to contents · Next: 12. Aggregate the economy without multiplying a headline by everyone

12. Aggregate the economy without multiplying a headline by everyone

Baseline economy-wide labour productivity is CU100 million divided by 1.6 million hours, or CU62.50 per hour. An unweighted average of the three sector rates would be (100 + 50 + 50)/3, approximately CU66.67, which is wrong for the economy. The sectors use different numbers of hours. The correct rate can be obtained by dividing totals or by weighting sector rates by their hour shares.

The adopting half of digital services initially contributes CU20 million. Its 20% gain adds CU4 million, so it now contributes CU24 million. The non-adopting half remains at CU20 million. Digital services therefore reach CU44 million, a 10% sector-level increase. The economy reaches CU104 million with the same 1.6 million hours. Labour productivity becomes CU65 per hour, a 4% increase over CU62.50.

The bridge can also be checked by multiplication: the affected sector’s baseline output share is 40%; adoption covers 50% of that output; adopters improve by 20%. Under the packet’s assumptions, 0.40×0.50×0.20 = 0.04, or 4%. This is an exact result for this deliberately fixed accounting example. It is not a general formula for every change in a real economy with shifting prices, inputs and networks.

Now add 50,000 actual implementation hours to the economy during the period, without changing the supplied CU104 million output. Total hours become 1.65 million. Productivity is 104/1.65, approximately CU63.0303 per hour, or 0.8485% above baseline. The impressive local gain still exists, but transition resources change the period’s aggregate ratio. Do not subtract those same hours again from output; they already enter the denominator.

A different extension lets goods production rise by 10% at unchanged hours, perhaps because a specified downstream improvement is assumed. Goods output becomes CU38.5 million. Without the extra implementation hours, total output is CU107.5 million and productivity is CU67.1875 per hour, a 7.5% increase. The additional CU3.5 million is a new scenario input. We cannot infer it merely from digital firms’ success.

This exercise teaches both restraint and openness. Partial adoption and limited economic weight can make a large local gain small in aggregate. Complementary improvements and spillovers can enlarge the effect. The right response is to measure or model those connections, not to assume they are either zero forever or automatically enormous. Every additional effect needs a source, a mechanism or an explicitly labelled scenario assumption.

Back to contents · Next: 13. An average can change because activity moves

13. An average can change because activity moves

Return Mere Island to its baseline before AI adoption. Now move 100,000 labour hours from goods production to digital services. Assume, only for this scenario, that both sectors preserve their original output per hour and have all the complementary resources required. Digital hours rise from 0.4 to 0.5 million; goods hours fall from 0.7 to 0.6 million. Personal-service hours stay at 0.5 million.

Digital output becomes CU50 million, goods output becomes CU30 million and personal-service output stays at CU25 million. Total output is CU105 million using the original 1.6 million hours. Aggregate productivity rises to CU65.625 per hour, a 5% gain. Yet no sector improved its own productivity rate. The aggregate increased because a larger share of activity occurred in the higher-output-per-hour sector.

This is a composition effect, not evidence that the measurement is broken. The economy really produces more under the scenario’s assumptions. But it is a different mechanism from an AI tool making the same worker and equipment more efficient inside a firm. A national statistic combining both mechanisms cannot by itself tell us which one dominated.

The scenario’s hard assumption is the ease of movement. The receiving sector needs suitable skills, equipment, customers and organisational capacity. Workers may face retraining costs, location constraints and different working conditions. Its original output per hour might not hold when it expands. Treating the movement as frictionless is useful for identifying an accounting possibility, but poor evidence that the gain is immediately available in practice.

Composition can also raise an average during an undesirable contraction. If low-output-per-hour activity disappears and its workers do not find new work, total output can fall while average output per remaining hour rises. That observation would not show that society became better off. We must examine output, hours, employment and distribution together. The productivity ratio describes one relationship; it is not a complete score for the economy.

When reading a claim that AI caused an aggregate acceleration, ask whether the analysis separates improvements within firms from movements between firms and sectors. Also ask whether changes in hours and output reflect a business-cycle recovery or contraction. The same headline ratio can emerge from different stories. Policy and personal decisions depend on the story, so identifying the decomposition is substantive work rather than a statistical technicality.

Back to contents · Next: 14. Labour productivity can rise because more capital is used

14. Labour productivity can rise because more capital is used

An organisation can produce more per staff hour by giving staff more equipment or purchased services. That can be an excellent improvement, but the additional resources matter. The purpose of a combined-input measure is to ask how output changes after accounting for more than labour. The BLS explanation of labour productivity and total factor productivity provides a useful introduction to why both measures are needed.

Use a separate fictional index exercise. Baseline output, capital services and labour are each normalised to one. In the next period, output rises to 1.10, capital services rise to 1.20 and labour remains at one. For teaching purposes only, define a production relation Y = A×square-root(K×L). Capital and labour receive equal exponents. This is a chosen model, not an estimate of Mere Island’s production function.

Labour productivity rises by 10%, because output rises by 10% while labour is unchanged. The combined-input index rises to square-root(1.20×1), approximately 1.095445. The implied A ratio is 1.10/1.095445, approximately 1.004158. Under this model, the residual efficiency term rises by about 0.416%, much less than the 10% labour-productivity increase. Most of the output increase is accounted for by the expanded capital-services input.

Do not read that residual as a pure AI score. An empirical residual can reflect omitted inputs, utilisation, measurement error and other factors as well as technology or organisation. Nor should the toy exponents be applied to a real country without evidence. The exercise demonstrates why one metric cannot answer every question; it does not supply a ready-made national growth-accounting model.

Capital services also differ from a purchase invoice. A machine bought this year can supply production services across several years, while its usefulness depends on utilisation, quality and depreciation. Treating the full purchase price as this month’s labour hours would mix units. Similarly, a software bill is a cost, not automatically a quantity of capital services. Appropriate measures require consistent definitions and prices.

For AI, this distinction keeps the infrastructure in view. More output per human hour may involve additional compute, energy, equipment and supplier labour. Those inputs do not invalidate the gain. They help explain its resource requirements and whether the improvement remains attractive at scale. A cheap demonstration with subsidised computing and a production service carrying its full costs answer different economic questions.

Back to contents · Next: 15. Prices, income and wellbeing can move differently

15. Prices, income and wellbeing can move differently

Return to the central Mere Island adoption scenario: real output is CU104 million and hours are 1.6 million. Now suppose the digital sector’s value-added price deflator falls from one to 0.9, while other sector deflators remain at one. Digital real value added of CU44 million corresponds to CU39.6 million at current prices. Current-price totals become 39.6 + 35 + 25 = CU99.6 million, while constant-price output remains CU104 million.

Current-price value added has fallen by 0.4% from the original CU100 million, even though real output has risen by 4%. There is no contradiction. The fictional economy supplies more output at lower prices in the affected sector. Using current receipts as the only productivity measure would miss the distinction. This example uses simple additive constant-price values; it should not be substituted mechanically for official chain-volume aggregation methods.

Lower prices can make a useful service accessible to more customers, but the response depends on demand and budgets. A lower price does not guarantee an unlimited increase in quantity purchased. If families need one document rather than ten, producing drafts more cheaply may reduce spending or free time rather than multiply demand tenfold. New uses can emerge, but they must be identified rather than assumed from supply capability alone.

Income distribution is another question. Some benefits may reach customers through prices, workers through compensation or improved conditions, and owners through returns. The productivity ratio contains no rule assigning those benefits. Two economies with the same measured efficiency improvement could distribute its gains differently because their ownership, competition and institutions differ. A fuller analysis of wealth and inequality would address that distribution question; this guide establishes the production arithmetic first.

Wellbeing is broader still. Reduced waiting, improved accessibility, greater leisure and less frustrating work can matter even when they are imperfectly captured by a production measure. Conversely, more measured output can coexist with environmental damage, intrusive monitoring or an unequal transition. It is better to maintain explicit additional indicators than to insist that GDP either captures all value or captures nothing useful.

For a student or parent, this is a useful reading habit: translate “the economy benefited” into several questions. Was there more real output? Were fewer resources needed? Did prices or access improve? Who received the gains? What costs were displaced elsewhere? Clear answers can support a positive conclusion without demanding that one number carry the entire ethical and economic argument.

Back to contents · Next: 16. A before-and-after improvement does not identify its cause

16. A before-and-after improvement does not identify its cause

Suppose Alder’s output rises after it introduces AI. That timing alone does not establish the contribution of AI. The firm may also have received easier orders, replaced a printing machine, hired an experienced supervisor or recovered from a temporary disruption. A causal question asks what would have happened under an appropriate alternative without the intervention. That counterfactual is not directly observable for the same firm at the same moment.

Here is a fully supplied fictional comparison. Two branches, North and South, record accepted output per hour over four periods. North records 8, 9 and 10 in three pre-intervention periods, then 12 afterward. South records 7, 8 and 9 before, then 10 afterward. Only North receives the new process. Quality standards, case mix and measurement definitions are assumed unchanged, and there are no spillovers between branches for this exercise.

North’s before-after increase is two units per hour. South’s increase is one. A difference-in-differences calculation in levels is (12 − 10) − (10 − 9) = one additional unit per hour. If North would have followed South’s one-unit increase without treatment, its counterfactual is 11. Its actual 12 is approximately 9.091% above that counterfactual. The one-unit effect is also 10% of North’s pre-period level of 10. These percentages use different bases.

The crucial assumption is parallel untreated trends in levels, not merely the existence of a control branch. The three pre-period observations are consistent with parallel one-unit increases, but they cannot prove what would have happened next. A new customer contract at North, different seasonal exposure or a change in measurement could invalidate the comparison. We have supplied no sampling distribution, so the example yields no confidence interval or statistical significance claim.

Random assignment can strengthen a local comparison when feasible, but design still matters. Workers may share methods between groups, choose which tasks to attempt or change behaviour because they know they are observed. Longer-term effects can differ from short-term trials. An evaluation should therefore specify the intervention, outcome, comparison, period and likely threats, rather than treating the word experiment as a guarantee that every wider conclusion follows.

The lesson is not to wait for perfect evidence before learning. It is to match confidence to the design. A well-recorded pilot can identify a promising mechanism. A controlled comparison can make a causal claim more credible. Broader data can show whether the mechanism spreads. The chain is stronger when each piece is labelled honestly than when a single before-after chart is asked to prove an economy-wide transformation.

Back to contents · Next: 17. What would a superintelligence scenario change?

17. What would a superintelligence scenario change?

A hypothetical system with much broader capabilities could change several assumptions at once. It might improve preparation and review, develop better equipment designs, accelerate research or help organisations adapt. It might also create services that the original output categories do not describe. These possibilities make it inappropriate to assume that a present tool’s modest effect is a permanent ceiling on every future system.

They do not make economic constraints disappear by definition. A superior design must still be translated into usable production. Physical construction, material availability, energy, customer needs and legitimate decisions can remain relevant. Some of those constraints might themselves become easier to address with stronger intelligence. The question is how, on what timescale and using which resources, rather than whether the word superintelligence permits us to omit them.

To make the distinction concrete, alter Alder’s packet again. Assume a hypothetical future system reduces all variable staff time to 0.2 hours per accepted pack, while recurring system work rises to 50 hours. Assume quality remains acceptable and staff availability stays at 1,000 hours. Labour-only capacity becomes (1,000 − 50)/0.2 = 4,750 packs. This is a fictional conditional result, not a demonstrated capability or an ASI forecast.

If demand is still 1,060 and equipment capacity still 1,080, output remains 1,060. If new customers want 2,000 packs and physical capacity is expanded to 2,000, output can reach 2,000, but the expansion is an additional assumption requiring resources. The scenario makes a useful point: radically better cognition could create large possibilities while the realised outcome depends on the rest of the production system. A different scenario could change those constraints too, provided it states the changes.

Economic forecasts disagree partly because they choose different affected-task shares, cost reductions, adoption paths and innovation mechanisms. Acemoglu’s The Simple Macroeconomics of AI illustrates a task-based route from cost savings to aggregate effects. Its 2024 analysis reports a 0.66% ten-year TFP gain in one baseline calculation using its stated evidence and assumptions. That is a dated conditional estimate, not a universal ceiling on future ASI or an observed outcome.

When comparing forecasts, align the quantity, baseline and horizon before judging disagreement. A permanent annual growth-rate increase is not a one-time level effect; GDP is not TFP; a broad research-acceleration scenario is not a calculation limited to existing tasks. Ask which assumption would have to change to reconcile the results. That question is more productive than choosing the largest or smallest headline as if confidence were determined by optimism or pessimism.

Back to contents · Next: 18. Diagnose the missing step rather than demanding more speed

18. Diagnose the missing step rather than demanding more speed

The examples suggest different repairs for different failures. If the task becomes faster but required workflow hours do not fall, inspect the new work created around it. Alder’s extra checking is the relevant category; another system might create reconciliation, correction or support. Reducing the burden requires understanding its cause. Deleting necessary checking from the measurement would only make the reported result better.

If workflow hours fall but completed output stays flat, examine demand, downstream capacity and the use of released time. A firm may have a valuable capacity improvement without immediate output growth. It can test whether a waiting list is genuine, whether another stage is constrained, or whether the released hours can improve a different useful outcome. The repair depends on which of these explanations is supported.

If successful pilots do not spread, inspect whether later users have the same complements as the pilot team. A prototype built by specialists with clean data can underestimate the cost faced by ordinary branches. Documentation, simpler integration or targeted training may help, but the intervention should address the observed difference. Requiring everyone to copy the pilot’s enthusiasm is not a technical explanation of why it worked.

If company gains do not appear in national statistics, check economic weight, adoption coverage, transition resources and offsetting developments. Also check whether the company gained market share or moved work outside its boundary. These are testable possibilities. “The statistics must be wrong” and “the technology must be useless” are both premature when the accounting bridge has not been examined.

The literature on the modern productivity paradox distinguishes explanations including disappointing capability, mismeasurement, redistribution and implementation lags. These are competing or overlapping possibilities, not a licence to assume that every missing gain is merely delayed. A useful investigation asks which observation would discriminate between explanations. Repeated failure to improve accepted outcomes should weaken the investment story; evidence of reusable complements and improving later results can strengthen it.

For independent use, write one claim in a sentence, identify the ratio it requires, and draw the shortest ledger connecting the current evidence to that ratio. Circle the first unsupported step. That circle is your next research question. This approach produces a smaller, answerable investigation than asking whether AI will transform everything or nothing.

Back to contents · Next: 19. Independent transfer packet

19. Independent transfer packet

Try these problems before reading the next chapter. All organisations and economies are fictional, all data are supplied, and every unchanged assumption is stated. Use a calculator if helpful. Your explanation of the denominator matters as much as the final percentage. Do not introduce a current AI benchmark as a missing parameter.

Problem A: A map-printing team completes 200 accepted sets in 400 staff hours. Drafting uses 160 hours, validation uses 80 and production uses 160. The new process reduces drafting to 40 hours, raises validation to 100 and leaves production at 160, all at the original 200-set volume. It adds 20 fixed hours. Variable categories scale proportionately. Quality is unchanged. Find required hours at 200 sets, the time-saving percentage and the operating-productivity increase. Then use demand of 220 sets, physical capacity of 210 and 400 available hours to find completed output. Assume total actual recorded hours remain 400, including improvement work, and calculate the observed set-per-total-hour change.

Problem B: Use Mere Island’s baseline of CU40 million and 0.4 million hours in digital services, CU35 million and 0.7 million hours in goods, and CU25 million and 0.5 million hours in personal services. Only 25% of digital output and hours adopt, but those adopters increase output by 40% at unchanged hours. Everything else stays constant. Find aggregate output and productivity. Then add 100,000 implementation hours without adding output. Explain whether rising output necessarily produces a rising output-per-hour ratio.

Problem C: A baseline team completes 500 accepted cases in 300 hours. An assisted trial attempts 500 cases in 250 hours, but 50 remain defective. All 50 can be repaired in 30 additional hours during the same period; no other output or input changes. Calculate accepted output per hour before repair and after repair, each relative to baseline. Explain why counting the repaired cases twice would be wrong and why the final ratio can be lower than the pre-repair accepted-output ratio.

Problem D: Start with Mere Island’s baseline. Remove 100,000 hours from personal services, whose output remains CU50 per hour for the hours retained. Those removed hours are not used elsewhere. The other sectors are unchanged. Calculate total output, total hours and aggregate productivity. Does a higher average prove that the economy expanded or that the people who lost work benefited?

Problem E: A treated branch moves from 12 to 15 accepted units per hour. A comparison branch moves from 10 to 12. Assume parallel untreated trends in levels for the calculation, comparable quality and no spillovers. Find the difference-in-differences estimate, the treated branch’s counterfactual and its percentage gain relative to that counterfactual. Name one reason the two-period data alone cannot establish the causal assumption.

Problem F: A hypothetical future process uses 0.2 variable staff hours per accepted pack and 50 fixed hours, with 1,000 available hours. Demand is 3,000 packs and physical capacity is 2,000. Quality is assumed unchanged. Calculate labour-only capacity, completed packs and operating hours for those packs. If total actual staff hours remain 1,000, compare output per total hour with a baseline of 1,000 packs in 1,000 hours. Explain why the result is not evidence that a real superintelligent system exists.

Back to contents · Next: 20. Answers, independent checks and interpretation

20. Answers, independent checks and interpretation

For A, required hours are 40 + 100 + 160 + 20 = 320. Time saved is 80/400 = 20%. Operating productivity rises from 200/400 = 0.5 to 200/320 = 0.625 sets per hour, a 25% gain. Variable time is 300/200 = 1.5 hours per set, so H = 1.5Q + 20. Labour-only capacity is 380/1.5, approximately 253.333 sets. Physical capacity of 210 is below demand of 220 and below labour capacity, so 210 sets are completed. They require 335 operating hours. With 400 total actual hours, output per total hour is 0.525, a 5% increase over 0.5. The additional 65 hours have the supplied improvement-work destination; they are not silently removed from the denominator.

For B, adopters initially account for CU10 million of digital output. Their 40% gain adds CU4 million. Aggregate output is again CU104 million, and with 1.6 million hours productivity is CU65 per hour, 4% above baseline. A smaller adoption share and larger adopter gain happen to produce the same aggregate result as the central packet. Adding 0.1 million implementation hours changes productivity to 104/1.7, approximately CU61.1765 per hour. Relative to CU62.50, that is a decline of about 2.118%. Output has risen by 4%, but hours have risen by 6.25%. The ratio can fall when its denominator grows faster than its numerator.

For C, baseline productivity is 500/300, approximately 1.66667 accepted cases per hour. Before repair, the assisted trial has 450 accepted cases in 250 hours, or 1.8 per hour. That is 8% above baseline. After repair, there are 500 accepted cases in 280 hours, or approximately 1.78571, which is 7.143% above baseline. Repair adds 50 accepted cases in 30 hours, a marginal rate of approximately 1.66667, below the trial’s pre-repair average of 1.8. It therefore lowers that average while increasing accepted output. The original defective attempts are not accepted cases, so adding 500 attempted plus 50 repaired would fabricate 50 extra outputs.

For D, personal-service output falls by CU5 million because 0.1 million hours at CU50 per hour are removed. Total output is CU95 million and total hours are 1.5 million. Productivity is approximately CU63.3333 per hour, about 1.333% above CU62.50. Yet total output fell by 5%, and hours fell by 6.25%. The higher ratio arises from composition. It does not prove an expansion, a technological improvement within any sector or a good outcome for affected workers. You would need evidence about employment, income, alternatives and wellbeing to assess their experience.

For E, the estimate is (15 − 12) − (12 − 10) = one unit per hour. Under the supplied parallel-trend assumption, the treated branch would have risen by two without treatment, reaching 14. Its observed 15 is 1/14, approximately 7.143%, above that counterfactual. Two periods provide no pre-intervention trend history. Different demand changes, a simultaneous management change or a measurement change could also undermine the comparison. Supplying an assumption lets us perform the exercise; it does not turn the assumption into observed evidence.

For F, labour-only capacity is (1,000 − 50)/0.2 = 4,750 packs. Physical capacity of 2,000 binds before demand of 3,000 or labour capacity. Completed output is 2,000, requiring 0.2×2,000 + 50 = 450 operating hours. With 1,000 total actual hours, the observed ratio is two packs per hour, double the baseline and therefore 100% higher. The other 550 hours need a stated destination in any real assessment; the exercise simply holds the total fixed. All capability and quality assumptions were supplied fictionally, so the calculation demonstrates their implications, not their truth.

A strong independent response shows the inputs, formula, denominator and interpretation. Check your work by reconstructing at least one answer from both a category ledger and a per-unit equation. Then alter one input and predict the direction of the result before recalculating. If you can explain why an output increase, an efficiency increase and a welfare improvement can differ, you have learned the transferable structure rather than memorised a percentage.

Back to contents · Next: 21. Questions readers often ask

21. Questions readers often ask

If AI is genuinely useful, why might national productivity barely move?

Its uses may cover a small part of the economy, adoption may be incomplete, and gains may be offset by transition work or unrelated changes elsewhere. Some benefits may also be difficult to measure. None of these explanations should be assumed without evidence. The Mere Island packet shows how a 20% adopter gain becomes a 4% aggregate gain under explicit weights, and how extra hours reduce that period’s measured improvement further.

Does a slow transition prove that the technology is overhyped?

No, but a transition story is not an unlimited excuse. Ask whether the proposed complements are actually being built, whether accepted outcomes improve in later periods, and whether remaining costs are falling or simply being renamed. A useful technology can have a difficult introduction. A poor application can also consume resources indefinitely. Distinguishing them requires milestones tied to outcomes rather than faith in an eventual curve.

Could superintelligence make adoption almost immediate?

It could conceivably reduce some design, learning and coordination delays if those capabilities were achieved. That is a scenario to analyse. It would still be necessary to specify physical resources, deployment scope and the outcomes being produced. Assuming away every constraint may define an interesting theoretical case, but it does not establish a calendar date or show that observed current-AI results already describe that case.

Should students conclude that learning ordinary economics is becoming unnecessary?

No. The examples depend on understanding ratios, denominators, opportunity cost, comparison groups and assumptions. An assistant can calculate quickly while answering the wrong question. Being able to explain what a number means remains useful when the tools improve. Practise checking a complete ledger and distinguishing evidence from a scenario before asking the tool for a conclusion.

Where should a Singapore reader look for actual productivity data?

The Singapore Department of Statistics productivity dashboard provides productivity trends by industry using value added per actual hour worked and per worker. Read the series definition, reference period and revision information before comparing figures. This article does not attribute a particular Singapore productivity change to AI; such a claim would require additional evidence about adoption, outcomes and alternative explanations.

What is the shortest defensible conclusion?

Better AI can expand productive possibilities, sometimes substantially. Realised productivity depends on which tasks improve, the resources required, the share of activity affected and what the surrounding economy does in response. A task demonstration is a starting piece of evidence. A convincing economic claim shows the bridges, supplies the denominator and remains clear about which future capabilities are hypothetical.

Back to contents · Next: 22. Continue with the right question

22. Continue with the right question

For the broader foundations, read How Productivity Works and How Economics Works. Those guides introduce the wider subject; this article has focused on the specific bridge from AI capability to measured economic productivity.

For employment effects, continue with Will Superintelligence Replace Jobs? Tasks, Occupations and the Missing Assumptions. For the conceptual distinctions behind hypothetical ASI, return to What Is Super Intelligence?.

Back to contents · Next: Sources and scope

Sources and scope

The workshop, economy, branch comparisons, quality cases and transfer exercises are original fictional teaching examples. Their numerical parameters are supplied assumptions, not estimates borrowed from the research below. Calculations can be reproduced using the equations and inputs printed in the article. No model test, company deployment, national causal result or ASI capability is claimed from those examples.

Back to contents · Return to the Super Intelligence guide

Discover more from eduKate Singapore

Subscribe now to keep reading and get access to the full archive.

Continue reading