What Is Super Intelligence? · Energy, evidence and deployment

Choose your reading route
Understand what AI can improve in energy, reproduce the numbers and recognise what remains unproved. All worked packets are fictional educational examples.
Understand the claim
1. What AI can change, and what an energy system still needs · 2. Read energy evidence by its stage, date and boundary · 3. Power, energy, capacity and timing · 4. Better designs need a path from candidate to asset
Assess the system
5. Forecasting is a support function before it is a control claim · 6. Reliability requires more than a good average · 7. Interconnection and construction do not disappear when studies become faster
Check the quantities
8. Computation creates demand even when individual tasks become more efficient · 9. Energy savings, peak reduction and load shifting are different outcomes · 10. A storage packet: matching energy, power and the service promised
Make and test a decision
11. Measure outcomes against a fair counterfactual · 12. A complete fictional review: the Riverbend proposal · 13. Build an evaluation brief that someone else can check · 14. Independent practice, with complete answers · 15. Questions readers should be able to answer
1. What AI can change, and what an energy system still needs
AI can help people discover promising materials, compare engineering designs, forecast demand, detect equipment problems and prepare decisions about an electricity system. None of those achievements, by itself, establishes that a new asset is ready to supply dependable power. The useful question is not simply whether a model is intelligent. It is which part of the energy problem the model improves, what evidence supports that improvement, and what must happen before people receive a better service.
This guide uses Super Intelligence as eduKate’s practical editorial umbrella for AI and possible future capabilities. Technical artificial superintelligence, or ASI, means a hypothetical system with broadly superhuman intellectual capability. An accurate electricity forecast, an impressive battery candidate or an effective planning assistant does not demonstrate that ASI exists. The examples below concern claims we can inspect today; discussion of future capability remains conditional.
Energy is a particularly useful subject for learning this distinction because the consequences are physical. A drawing must become equipment. Equipment must be manufactured, delivered, installed, connected, tested and maintained. A forecast must arrive early enough for an authorised person or validated control system to use it. An efficiency improvement must survive changes in workload, weather and operating conditions. A model can improve one link while the final result remains constrained elsewhere.
Consider a fictional team that announces an AI-designed storage device. Its simulation predicts excellent performance. The team has learned something worth investigating, but it has not yet established manufacturing yield, lifetime, safety, repairability or installed cost. Another team improves tomorrow’s solar forecast. That may be operationally useful without inventing any new hardware. A third team makes an AI service use less electricity per accepted task. It may nevertheless consume more electricity overall if many more tasks are performed. These are three different claims and should receive three different evaluations.
A reader can keep them separate by following a chain: proposed capability, checked evidence, permitted use, physical deployment and measured outcome. Each transition has a question. Does the capability work on genuinely new cases? Does the evaluation resemble its intended use? Who is allowed to act on the result? Are the necessary assets available? Did the service improve after costs and failures were counted? Missing answers should remain visible rather than being filled by enthusiasm.
The purpose here is energy literacy and evidence evaluation. The worked packets are entirely fictional, with deliberately simple numbers that can be reproduced on paper. They are not equipment specifications, grid operating instructions, investment advice or permission to control infrastructure. Real switching, dispatch, protection settings, emergency response and electrical work belong to qualified, authorised professionals operating under the applicable procedures.
Back to contents · Continue: 2. Read energy evidence by its stage, date and boundary
2. Read energy evidence by its stage, date and boundary
A useful evidence claim names its object. A model may predict a material property; a laboratory may measure that property; a manufacturer may demonstrate repeatable production; an operator may demonstrate performance in service. These findings can strengthen one another, but they are not interchangeable. A report of discovery should not quietly become a report of deployment. A report of a pilot should not become a promise about every electricity system.
The United States Department of Energy’s 29 April 2024 AI for Energy overview identified potential applications in grid planning, permitting, operations and reliability, and resilience. Examples included accelerated models for capacity and transmission studies, assistance with review documents and renewable-generation forecasts. That is an authoritative account of opportunity areas and their associated risks, not evidence that every proposed application has delivered benefits at commercial scale. DOE: AI for Energy, 29 April 2024
The International Energy Agency’s report published on 16 April 2026 estimated global data-centre electricity consumption at 485 TWh in 2025 and projected about 950 TWh in 2030. The first is an estimate for a past year; the second is a projection. Both concern data centres as a category, not exclusively AI. The same summary describes rising use alongside improvements in task efficiency. These quantities should retain their original scope when repeated. IEA: Key Questions on Energy and AI, executive summary
Dates matter because a forecast is a statement made with information available at a particular time. A later revision can reflect new evidence rather than dishonesty in the earlier analysis. When comparing reports, preserve the publication date, reference year, geography, unit and scenario. A 2030 projection should not be presented as today’s measured consumption. Nor should an estimate for all data centres be relabelled as the energy cost of chatbots.
A strong reading method separates four categories in ordinary language. “Measured” means an observation made under stated conditions. “Estimated” means a quantity inferred from incomplete observations or a model. “Projected” means a conditional view of the future. “Proposed” means an intended action or capability. A single announcement can contain all four. Marking them separately often reveals that the most confident headline rests on the least mature part of the evidence.
Also ask who produced the evidence and what they can know directly. An equipment supplier can provide detailed test results but may have an interest in favourable presentation. A system operator can describe a collaboration without yet knowing its final effect. An independent laboratory can test a device without establishing whether its supply chain will scale. Source quality is partly about matching the source’s access to the claim being made.
Finally, do not treat uncertainty as a blank permission to choose a preferred story. Name the missing information and the decision it could change. If actual operating demand is uncertain, that affects infrastructure sizing. If the failure rate is uncertain, that affects the permitted role of an assistant. If a prototype’s lifetime is uncertain, that affects whether its apparent efficiency survives replacement costs. Uncertainty becomes useful when it changes the next test.
Back to contents · Continue: 3. Power, energy, capacity and timing
3. Power, energy, capacity and timing
Power is a rate of energy transfer. Energy is the quantity transferred over time. A megawatt, MW, measures power; a megawatt-hour, MWh, measures energy. A steady 4 MW load operating for three hours consumes 12 MWh. A 4 MW peak lasting only a minute does not imply the same total. The US Energy Information Administration explains the corresponding kilowatt and kilowatt-hour distinction, with one kilowatt-hour representing one kilowatt used for one hour. EIA: kilowatt-hour definition
For a variable load, calculate energy for each interval and add the results. In a fictional four-hour period, a site draws 2 MW for the first hour, 3 MW for the second, 5 MW for the third and 2 MW for the fourth. Its energy consumption is 2 + 3 + 5 + 2 = 12 MWh. Its average power is 12 MWh divided by four hours, or 3 MW. Its peak is 5 MW. The average and peak answer different planning questions.
Suppose someone proposes a connection rated for 3 MW because the average is 3 MW. The arithmetic is correct but the conclusion is unsupported: the third hour exceeds that rating in the supplied profile. An assessment would need an authorised, technically valid way to change the profile or provide an appropriate connection. Annual energy totals cannot establish that a particular network can carry a local peak. Neither an AI forecast nor an annual purchase contract changes this arithmetic.
Generating capacity is also not the same as constant production. A fictional 20 MW generator operating at a 30% average capacity factor over a 24-hour teaching day would produce 20 × 0.30 × 24 = 144 MWh. Its average output would be 6 MW. That daily average says nothing about whether output arrives in the hours when a particular customer needs it. Capacity factor is a ratio over a period, not a promise that each hour has the same output.
Storage adds two distinct quantities. A battery might have a 5 MW discharge limit and 20 MWh of usable output energy under stated conditions. In an idealised exercise, it can deliver 5 MW for four hours, or 2 MW for ten hours. It cannot deliver 10 MW merely because it holds enough energy for two hours at that rate. The power limit forbids that inference. Real specifications also require conditions for losses, operating reserves, degradation and temperature.
Always write units through the calculation. MW multiplied by hours gives MWh. MWh divided by hours gives MW. Dividing MWh by MW gives hours. If an expression leaves MW where the answer is supposed to be annual energy, something is missing. This simple habit catches errors that polished prose can hide. Large prefixes do not change the principle: one GWh is one thousand MWh, and one TWh is one million MWh.
Timing has another dimension: the difference between a forecast interval and an operating decision. A daily total may help with accounting while being too coarse to answer a question about an evening peak. A monthly average may conceal a difficult hot afternoon. Good evaluation begins by choosing the time resolution that matches the decision, then checking whether the measurements and predictions actually use that resolution.
Back to contents · Continue: 4. Better designs need a path from candidate to asset
4. Better designs need a path from candidate to asset
AI-assisted design is attractive because many engineering problems contain more possible combinations than a team can test exhaustively. A model can rank candidates, propose promising structures or approximate an expensive calculation. The benefit is often better allocation of experimental effort. It does not require believing that the model has discovered a final answer without further evidence. A good candidate is valuable because it tells the team what is worth checking next.
A 2024 research paper on AI and cloud high-performance computing described screening more than 32 million candidate materials, identifying promising solid-state electrolyte compositions and proceeding to experimental validation. The scope is a discovery pipeline. It is not a demonstration that every screened candidate works, that the resulting chemistry is a commercially qualified battery, or that electricity storage has ceased to face manufacturing constraints. Research paper: computational materials discovery and experimental validation
To evaluate a design claim, identify the property being optimised and the properties being constrained. A battery material selected for one conductivity measure might have other unresolved requirements. A turbine component may appear efficient while being difficult to manufacture or inspect. A building design may reduce one form of energy use while shifting costs to another season. An optimisation result is meaningful only with its objective and limits attached.
Here is a fictional evaluation packet for a thermal-storage component. The current design is called Reference. Two AI-selected candidates are Amber and Birch. All three are tested by the same fictional team with the same measurement procedure. On three laboratory runs, Reference delivers 96, 98 and 97 units of the chosen output measure. Amber delivers 108, 112 and 110. Birch delivers 103, 104 and 102. The units are deliberately generic: this packet teaches evidence interpretation rather than a construction method.
The averages are 97 for Reference, 110 for Amber and 103 for Birch. Amber’s average improvement over Reference is 13 divided by 97, approximately 13.4%. Birch’s is 6 divided by 97, approximately 6.2%. Based on this one measure, Amber appears stronger. But the packet also reports that eight of ten Amber pilot units meet manufacturing inspection criteria, while nine of ten Birch units and ten of ten Reference units meet them. There is no long-duration service evidence for either candidate.
The correct result is not “Amber is the best deployable technology.” It is “Amber leads this small laboratory comparison on the specified output measure, while its pilot inspection results and missing service evidence require further work.” Birch may deserve investigation because its apparent manufacturing consistency is better in this tiny packet, but ten units cannot establish a dependable population failure rate. The information supports a next-stage comparison, not a procurement claim.
The team’s next evaluation should be specified before the next attractive result arrives. It would require an agreed test population, defined operating conditions, independent measurement checks, a meaningful comparison and treatment of failed units. It would also need a plan for tracking changes between the laboratory sample and production design. Those are questions for qualified researchers and engineers; they are not steps for readers to build or operate equipment.
A common failure is survivor-only reporting. If the report measures only units that pass an initial screen, the headline can hide manufacturing difficulty. Another failure is changing several variables at once and crediting all improvement to the AI-selected feature. A third is treating an accelerated test as a complete substitute for service experience without explaining the relationship. Repair each failure by restoring the excluded cases, preserving the comparison and stating which inference remains untested.
The reader’s practical skill is to label the stage accurately. Candidate generation, simulation, laboratory measurement, prototype integration, qualified production and field performance are distinct milestones. Progress through them may be faster with AI, but each milestone supplies evidence the previous one did not. The contribution of intelligence is strongest when it makes those transitions more testable rather than less visible.
Back to contents · Continue: 5. Forecasting is a support function before it is a control claim
5. Forecasting is a support function before it is a control claim
Electricity demand and renewable output vary. A forecast gives a view of what may happen, with a particular horizon and information set. Its value depends on whether it improves a real decision relative to a credible alternative. A more complicated model is not automatically a better forecast, and a better forecast is not automatically a safer operating system. The entire chain from input to use needs examination.
Begin with a baseline that people could actually use. For some exercises, that might be a persistence forecast using a recent observation. For others, it might be an established weather-informed method. Do not compare a new system only with an intentionally weak guess. A fair test also protects the time boundary: information arriving after the forecast was due cannot be included in the model’s historical inputs.
Consider a fictional four-interval solar forecasting packet. Each interval lasts one hour. Actual output is 18, 26, 20 and 8 MW. The existing method predicts 20, 24, 24 and 12 MW. The candidate predicts 19, 25, 22 and 9 MW. Both sets of forecasts were recorded before their respective intervals began. The packet contains no grid constraints, dispatch prices or authority to operate equipment; it is a forecasting exercise only.
For the existing method, absolute errors are 2, 2, 4 and 4 MW. Their sum is 12 MW, and the mean absolute error is 12 divided by four, or 3 MW. For the candidate, absolute errors are 1, 1, 2 and 1 MW. Their sum is 5 MW, and the mean absolute error is 1.25 MW. The reduction in mean absolute error is 1.75 MW, or about 58.3% relative to the baseline’s 3 MW.
That improvement is a score on four observations. It is not a demonstrated 58.3% reduction in operating cost, reserve requirements, emissions or outages. Those outcomes have additional causes. Even the forecast claim must be limited: four intervals cannot establish performance across seasons, unusual weather, missing sensors or different plants. The result justifies further evaluation under the packet’s conditions, not a universal statement of superiority.
The direction of error also matters. Define signed error here as forecast minus actual. The baseline’s signed errors are +2, −2, +4 and +4 MW, averaging +2 MW. The candidate’s are +1, −1, +2 and +1 MW, averaging +0.75 MW. Both overpredict on average in this tiny sample. A lower absolute error is helpful, but a systematic bias in a consequential period can still matter. The sign convention must be written down so another reader can reproduce it.
Because each interval is one hour, the baseline’s absolute error magnitudes correspond to 12 MWh across the four intervals, while the candidate’s correspond to 5 MWh. These are sums of absolute forecast discrepancies, not measured electricity savings. The actual solar generation is 72 MWh under either forecast. Better information does not retroactively create seven MWh of generation. This distinction is one of the simplest checks against inflated AI-benefit claims.
A stronger evaluation adds many untouched periods, reports performance by conditions and shows whether uncertainty intervals are useful. It also examines missing-data behaviour, delayed forecasts, version changes and human workload. A model that improves the average but fails to produce a result on the hardest days may be less useful than its average score suggests. Count absent forecasts rather than deleting them from the denominator.
Finally, specify the permitted role. A forecast can be displayed to an analyst, incorporated into a reviewed planning study or consumed by a validated operational system. Those uses have different assurance requirements. “The AI predicts demand” describes a capability. “The AI may directly change equipment behaviour” describes authority and system design. Evidence for the first does not silently grant the second.
Back to contents · Continue: 6. Reliability requires more than a good average
6. Reliability requires more than a good average
An electricity service must work when conditions become difficult, not only when the average day resembles training data. Reliability evaluation therefore asks about failures, unusual combinations of events, recovery and the consequences of being wrong. A tool that is useful for retrospective analysis may need much stronger evidence before it influences a time-sensitive decision. The difference is not whether the tool uses AI; it is what depends on its result.
For public understanding, distinguish resource adequacy from the many details of operating reliability. Resource adequacy asks whether a system has sufficient resources to meet demand under an assessed range of conditions. An adequate annual energy total does not answer every question about short-term delivery, equipment outages or local network limitations. Conversely, an isolated equipment problem does not automatically demonstrate that an entire region lacks annual energy resources. The appropriate assessment depends on the failure being investigated. DOE: resource adequacy
A fictional maintenance-alert packet illustrates why an accuracy headline is insufficient. During an evaluation, 1,000 equipment observations are reviewed. Ten are confirmed to require the defined follow-up inspection. An AI system raises 30 alerts. Eight correspond to the ten confirmed cases; 22 do not. It misses two confirmed cases. The packet assumes the follow-up labels are complete and correct, an assumption that would itself require attention in a real study.
The system finds eight of ten confirmed cases, giving 80% recall. Only eight of its 30 alerts correspond to confirmed cases, giving about 26.7% precision. It correctly leaves 968 of the 990 non-cases unflagged. Overall accuracy is therefore (8 + 968) divided by 1,000, or 97.6%. A system that never raised an alert would achieve 99% accuracy in this imbalanced packet while finding none of the cases. Higher raw accuracy would conceal worse performance on the purpose of the task.
The maintenance team should not choose between those systems using accuracy alone. It needs to understand the consequence of a missed case, the burden of unnecessary inspections, the timing of detection and the quality of the labels. If inspection capacity is limited, too many low-value alerts can distract attention. If the missed cases are particularly consequential, an apparently efficient alert system may still be inappropriate. There is no universal score that decides those trade-offs without an intended use.
The safe role in this fictional packet is a reviewed prioritisation aid. The model’s alert does not certify a component as defective, and its silence does not certify equipment as safe. Qualified personnel must interpret the evidence through the established maintenance process. Readers should not convert the example into an instruction to inspect live electrical assets or alter maintenance schedules. It shows how to question a claim, not how to perform the work.
Reliability testing also needs unfamiliar conditions. What happens when a sensor stops updating? When equipment has been replaced? When weather exceeds the range represented in training? When two systems share the same erroneous data feed? Testing each model separately may miss a common-cause failure in the wider system. A fallback that depends on the same unavailable information is not truly independent of that failure.
A practical evaluation record should preserve the troublesome cases. Record why an alert was dismissed, whether an output arrived too late, what information was missing and which people carried the recovery burden. The relevant outcome may be less unplanned downtime, earlier justified inspections or better preparation, rather than simply more alerts. The evaluation should also count new work created by the tool. Otherwise the model can appear helpful by transferring effort out of the metric.
Back to contents · Continue: 7. Interconnection and construction do not disappear when studies become faster
7. Interconnection and construction do not disappear when studies become faster
Electricity infrastructure is a chain of related projects. A new generator may need a connection study and network upgrades. A new customer may need a suitable service connection. Equipment procurement, land agreements, permitting, financing, construction and commissioning can progress at different speeds. AI may help with some analysis or document work, but a faster analysis is not the same as an energised asset.
Berkeley Lab’s 1 July 2026 summary reported 2,061 GW of generation and storage capacity actively seeking US transmission interconnection at the end of 2025. Its historical analysis found that 13% of capacity in requests submitted during 2000–2020 had come online by the end of 2025, while 75% had withdrawn. These are capacity shares in the stated historical cohort, not probabilities that any particular new project will succeed. Berkeley Lab: 2026 queue findings
That evidence teaches an important reading rule: a queue is a pipeline of proposals, not an inventory of available power. Some proposals overlap, change or disappear. Storage power ratings and generator power ratings can be added for a carefully labelled capacity total, but the sum does not describe the energy that will be produced or guaranteed at a particular time. It is especially misleading to subtract a data-centre load from a queue total and call the remainder spare supply.
There are real efforts to improve the analytical part of this process. On 10 April 2025, PJM announced a multiyear collaboration with Google and Tapestry to develop AI-enhanced tools for generation-interconnection planning, with the aim of reducing processing times. The announcement establishes the collaboration and its purpose. It does not, on its own, establish a measured reduction in construction time, completed connections or customer bills. PJM: AI-enhanced planning collaboration
A fictional project schedule shows why this boundary matters. A project begins three parallel workstreams at month zero. The study takes six months, the permit process takes eight months and equipment procurement takes ten months. Construction takes four months and cannot begin until all three are complete. Commissioning then takes two months. Under these simplified dependencies, the earliest completion is month 16: ten months for the longest prerequisite, four for construction and two for commissioning.
Suppose an AI-assisted study process reduces the study from six months to three, with no change in quality. The earliest completion remains month 16 because equipment procurement still takes ten months. The analytical improvement is real, but the project completion gain is zero under this schedule. This does not mean the study improvement is worthless. It could reduce staff effort, uncertainty or the risk of a later delay. Those benefits should be named separately rather than represented as three months of earlier electricity.
Now change the packet. Suppose the original study takes 12 months while the other workstreams remain at eight and ten. Completion is month 18. Reducing the study to seven months moves the longest prerequisite to equipment at month ten, so completion becomes month 16. The study has become five months faster, but the complete project only two months faster. The slowest remaining dependency determines how much local improvement reaches the final outcome.
Real projects have more complicated dependencies and uncertain durations. The example is not a prediction of any permitting or construction process. Its lesson is to draw the sequence, identify which activities can overlap and ask which one currently controls completion. A claim that software will accelerate infrastructure should name the affected stage and the constraints that remain outside it.
This also changes procurement questions. A plan should distinguish an indicative equipment delivery date from a confirmed order, a conceptual connection from approved conditions and a construction target from completed commissioning. AI can organise these facts, flag inconsistencies and compare scenarios. It cannot truthfully convert missing commitments into completed milestones. The most useful report may be the one that makes an unresolved dependency impossible to overlook.
Back to contents · Continue: 8. Computation creates demand even when individual tasks become more efficient
8. Computation creates demand even when individual tasks become more efficient
The energy relationship runs in both directions. AI can support energy research and operations, while the computing used to develop and run AI requires electricity. These sides should be measured separately before anyone claims a net benefit. A useful application can create additional demand. An efficiency improvement can lower energy per task while aggregate consumption rises. Neither outcome is logically contradictory.
Use a fictional service with a complete, clearly stated measurement boundary. During a baseline day, it performs 100,000 accepted tasks and consumes 500 kWh for the included computing and facility overhead. Its energy intensity is 500 divided by 100,000, or 0.005 kWh per accepted task. That is 5 Wh per accepted task. The packet assumes rejected attempts and retries are included in the 500 kWh, so the metric does not hide unsuccessful work.
After an improvement, the service uses 3 Wh per accepted task under comparable conditions. If volume remains at 100,000 accepted tasks, consumption falls to 300 kWh, saving 200 kWh or 40%. If volume rises to 200,000 tasks, consumption becomes 600 kWh. Energy per accepted task has improved by 40%, but total consumption is 20% above the original 500 kWh. Both statements are true and belong in the same report.
The break-even volume is 500,000 Wh divided by 3 Wh per task, or approximately 166,667 accepted tasks. Below that volume, the improved service uses less than the original daily total under the supplied assumptions. Above it, the total is higher. This is a sensitivity calculation, not a forecast of demand. It shows precisely what must be known before efficiency can be translated into an aggregate energy claim.
Task comparability is equally important. A brief answer, a long reasoning workflow and a generated video are not interchangeable units of useful work. Even two text tasks can differ in input size, output length, retries and quality requirements. If the service starts performing harder tasks, the denominator changes. The evaluator should report task mix rather than suggesting that every request has a universal energy cost.
The accepted-task definition prevents another distortion. Imagine two systems that process the same 1,000 requests. System One uses 10 kWh and produces 900 acceptable results. System Two uses 8 kWh and produces 600 acceptable results. Their intensities are approximately 11.1 Wh and 13.3 Wh per acceptable result respectively. The second system uses less total energy in this packet, but more energy per acceptable result. Choosing between them also requires knowing whether the missing results can be tolerated and how they are repaired.
Facility-level metrics answer a related but different question. Power usage effectiveness, PUE, compares total facility energy with IT equipment energy over a matched period and boundary. DOE FEMP: PUE definition A fictional facility with 1,200 MWh total consumption and 1,000 MWh IT consumption has PUE 1.20. Reducing total consumption to 1,150 MWh with the same IT energy gives PUE 1.15. That is a 50 MWh facility saving, not a demonstration that the models produced more useful answers.
If IT consumption then rises to 1,300 MWh at PUE 1.15, facility consumption becomes 1,495 MWh. The PUE has improved while total consumption exceeds the original 1,200 MWh. A lower ratio does not guarantee a lower total. It also does not describe emissions, water use or the quality of the application. Readers should ask for the metric that answers their question rather than allowing one convenient number to stand for every outcome.
For a school project or an organisational pilot, the right response is modest measurement. Specify the workload, count retries, identify which equipment and overhead are included, record the period and retain quality criteria. Do not infer a precise energy bill from a model’s name alone. Where direct measurements are unavailable, label the result as an estimate and explain the assumptions that dominate it.
Back to contents · Continue: 9. Energy savings, peak reduction and load shifting are different outcomes
9. Energy savings, peak reduction and load shifting are different outcomes
An energy-saving intervention reduces the quantity consumed within a defined comparison. Peak reduction lowers the highest demand during a stated period. Load shifting moves consumption from one time to another. These outcomes can overlap, but they do not have to. An AI scheduling tool may be valuable because it moves flexible work away from a constrained period while leaving total energy unchanged. Calling that energy saving would misdescribe its benefit.
Consider a fictional data-processing site over six one-hour intervals. Its fixed demand is 4 MW in every interval. It also has a flexible job requiring 6 MWh. In the baseline, that job runs at 3 MW in hours three and four. Total site demand is therefore 4, 4, 7, 7, 4 and 4 MW. The fixed component uses 24 MWh; the flexible job uses 6 MWh; total consumption is 30 MWh. The peak is 7 MW.
An authorised scheduler instead spreads the flexible job at 1 MW across all six hours, assuming the task permits that timing and the same result is delivered. Site demand becomes 5 MW in each hour. Consumption remains 30 MWh, but the peak falls to 5 MW. The peak reduction is 2 MW, approximately 28.6% of the original 7 MW. The energy saving is zero. A correct report states both outcomes and the flexibility assumption.
The exercise omits real network and operational details deliberately. It does not establish that any actual customer may change its load in this way. A task may have deadlines, dependencies, data-location requirements or equipment constraints. A service may also have an agreed demand-response arrangement that defines what can be changed and by whom. Permission, technical feasibility and user service quality must be checked separately from a mathematically attractive schedule.
Suppose the site’s fictional electricity price is 200 units of currency per MWh in hours three and four and 100 in the other four hours. The baseline flexible job costs 6 × 200 = 1,200 currency units. Under the spread schedule, two MWh occur in the expensive hours and four in the cheaper hours, costing 400 + 400 = 800. The flexible-job energy cost falls by 400 currency units, even though the job still consumes 6 MWh.
That calculation excludes network charges, demand charges, taxes, contractual restrictions and other costs. It is a teaching example, not a tariff quotation. It also leaves the fixed load’s bill unchanged between schedules. A report claiming a one-third reduction in the entire site bill would be wrong: the one-third reduction applies only to the flexible job’s simplified energy charge. The boundary of a percentage matters as much as the arithmetic.
Emissions require another boundary. A lower-price hour is not automatically a lower-emissions hour. Annual average emissions factors and the effect of changing demand at a particular time answer different questions. A scheduler seeking emissions reductions needs appropriate data and a defensible method for the intended claim. It should not simply relabel financial savings as carbon savings. Nor should it count purchased certificates as physical proof that a site was supplied by a particular generator in every hour.
Now suppose the new schedule introduces extra processing overhead of 0.3 MWh while still delivering the same accepted result. Total site energy becomes 30.3 MWh. The peak may still be lower, depending on when that overhead occurs, and the simplified bill may still fall. An honest evaluation retains all three dimensions: energy, peak and cost. It does not remove the additional consumption because it complicates a favourable story.
The broader lesson is that “optimisation” needs an objective and a constraint set. Lower cost, lower peak, lower emissions, faster completion and greater reliability can pull in different directions. A tool can help compare feasible choices, but it cannot decide that one value may be sacrificed without authority. Good reporting makes the trade-off legible enough for the responsible people to choose.
Back to contents · Continue: 10. A storage packet: matching energy, power and the service promised
10. A storage packet: matching energy, power and the service promised
Storage is often invoked as if it were a single answer to variability. A proper claim specifies the service: shifting a known quantity between hours, supporting a defined load for a duration, absorbing a particular surplus or providing another technically qualified function. Power, usable energy, starting state, losses and recharge conditions all affect the answer. A battery’s presence does not establish that the promised service can be delivered. EIA: energy storage for electricity generation
In a fictional planning packet, a site asks whether a storage system can support a 3 MW load for five hours. The battery’s maximum output is 4 MW, and it starts with 12 MWh of usable energy measured at the output boundary. The packet deliberately defines usable energy after the relevant discharge losses and operating reserve, so those are not subtracted again. The required energy is 3 × 5 = 15 MWh. The power limit is sufficient, but the energy is not.
With 12 MWh available, the idealised duration at 3 MW is four hours. The shortfall against the five-hour request is 3 MWh. A claim that “4 MW is greater than 3 MW, so the battery is sufficient” considers only the rate. A claim that “12 MWh is a large battery, so it should last” avoids the calculation. The correct conclusion is that the supplied packet does not meet the requested duration.
Consider a different request: 5 MW for two hours. It needs 10 MWh, which is less than the stated usable energy. But it exceeds the 4 MW output limit. This second request fails for a different reason. The two examples show why a single capacity number is ambiguous. A complete storage description should keep power and energy visibly separate, together with the conditions under which both ratings apply.
A third request is 2 MW for four hours, requiring 8 MWh. That fits both simplified limits and leaves 4 MWh of usable energy in the packet. The conclusion remains conditional on the starting state and the definition of usable energy. If the battery starts with only 6 MWh available, it cannot fulfil the same request. A model that assumes “fully charged” without evidence can make an otherwise correct calculation misleading.
Recharging also consumes energy. Suppose a separate fictional cycle has an 80% round-trip efficiency, measured consistently from charging input to discharged output. To deliver 8 MWh over that cycle, it requires 8 divided by 0.80 = 10 MWh of input. The difference is 2 MWh. Do not multiply 8 by 0.80 and call that the required input; doing so reverses the relationship. Also do not apply this round-trip factor again to the earlier packet’s already-defined usable output energy.
Storage can therefore enable useful timing changes while increasing total electricity required for the shifted service because of losses. Whether the complete arrangement improves cost, emissions or reliability depends on when and where it charges, what it displaces and what other constraints apply. This is not an argument against storage. It is an argument for measuring the particular service rather than awarding credit to a technology label.
AI may help estimate degradation, compare schedules or detect patterns that deserve expert review. Those functions must respect physical limits and independent safeguards. An evaluator should check that a model does not invent a higher power rating, spend the same stored energy twice or assume recharge during a period when the source is unavailable. These are consistency checks on a proposal, not instructions for controlling a live battery.
For readers, the best final sentence is precise: “Under the stated fictional assumptions, this packet meets the two-megawatt, four-hour request and does not meet the other two requests.” That statement is stronger than a broad claim that the battery is either good or bad. It names exactly what the evidence establishes and makes the calculation reproducible.
Back to contents · Continue: 11. Measure outcomes against a fair counterfactual
11. Measure outcomes against a fair counterfactual
An energy intervention should be compared with what would reasonably have happened without it. That comparison is often called a counterfactual. A before-and-after number alone can be misleading when weather, production, occupancy or service demand changed at the same time. AI is not exempt from this problem. A model can produce a sophisticated report that attributes a difference to the wrong cause.
Consider a fictional building analysis. During a baseline month, the building uses 100 MWh. During a later month with an AI-assisted advisory service, it uses 90 MWh. A headline claims a 10% saving. The supplied packet then adds that the later month had fewer operating days and milder weather. A pre-agreed adjustment method estimates that the building would have used 94 MWh in that later month without the intervention. The estimated intervention-related saving is therefore 4 MWh, about 4.3% of the adjusted 94 MWh comparison.
Even that revised figure is an estimate. It depends on the adjustment method and measurement quality. Suppose the evaluation reports a plausible adjusted-baseline range of 91–97 MWh. Against measured consumption of 90 MWh, the corresponding estimated saving ranges from 1 to 7 MWh. That interval communicates more than a falsely precise claim that the system saved exactly four. It also tells the team how valuable better baseline evidence might be.
The evaluator should not choose the adjustment model after seeing which one produces the largest saving. Define the comparison and exclusions before reviewing the results, or clearly label later analysis as exploratory. Preserve data gaps, unsuccessful recommendations and periods when the system was unavailable. If a recommendation was never implemented, it cannot be credited as a realised saving simply because the model estimated its potential.
Energy savings are not the only outcome worth measuring. The service might reduce time spent preparing reports, improve the consistency of documentation or help staff identify questions earlier. These are legitimate benefits if demonstrated. They should be recorded in their own units rather than converted into energy savings without a causal argument. Likewise, staff effort used to check the AI and repair its errors belongs in the workload account.
A useful report separates gross and net effects. Gross avoided consumption may be reduced by extra computing, additional sensors or other energy used by the intervention. The time boundary should be stated: one day of operating electricity is not a complete life-cycle comparison of hardware manufacture and replacement. A limited operational study can still be worthwhile as long as it does not claim to answer the broader question.
Distribution matters as well. A lower cost for one facility may depend on expenditure by another party. An infrastructure upgrade may benefit several users or be largely dedicated to one. A schedule that improves one customer’s bill may not improve the wider system under the same metric. The evaluator should identify who receives the benefit, who incurs costs and which effects have not been measured. A single net number can obscure important consequences.
This is where energy claims become a lesson in careful reasoning. First establish that a change occurred. Then assess how much of it can reasonably be attributed to the intervention. Then examine whether it produced the intended service without unacceptable trade-offs. Only after those steps should the report discuss transfer to other settings. The strength of the claim should increase with the strength of the evidence, not with the elegance of the presentation.
Back to contents · Continue: 12. A complete fictional review: the Riverbend proposal
12. A complete fictional review: the Riverbend proposal
Riverbend is an invented campus considering three AI-related proposals. Proposal A is a solar forecasting assistant. Proposal B is a flexible-computing scheduler. Proposal C is an AI-designed storage component offered for a future installation. The campus committee wants lower costs and reliable service. It has not authorised any model to control electrical equipment. Its decision is which proposals deserve a bounded next evaluation, not which supplier should receive an immediate contract.
The packet contains the four-hour solar results from Chapter 5, the six-hour computing profile from Chapter 9 and the laboratory comparison from Chapter 4. It also says that the campus’s planned connection upgrade depends on a ten-month equipment delivery, with other prerequisites finishing earlier. No proposal includes independently verified long-term operating results. All numbers are fictional. The committee has enough information to diagnose the evidence, but not enough to approve operational deployment.
For Proposal A, the analyst reproduces the mean absolute errors of 3 MW and 1.25 MW. The candidate is better on the supplied sample. The analyst then rejects the vendor-style sentence “AI saves seven MWh every four hours.” Seven MWh is the difference in summed absolute forecast discrepancies over those one-hour intervals. Actual generation remains 72 MWh. The next useful step is a larger time-separated forecast evaluation, including missing-data periods and difficult conditions.
For Proposal B, the analyst reproduces the reduction from a 7 MW peak to 5 MW with total energy unchanged at 30 MWh. The simplified flexible-job energy charge falls from 1,200 to 800 fictional currency units. The analyst asks whether all six hours are available, whether job quality and deadlines are preserved, and whether the measured overhead changes the result. Because the committee has not authorised control, the next evaluation should compare proposed schedules in a non-operational setting.
For Proposal C, the analyst confirms that Amber has the best average output measure in the small laboratory sample. Its manufacturing-inspection record is eight passes in ten units, and long-duration service evidence is absent. The analyst therefore refuses to combine the laboratory improvement with the campus storage requirement as if commercial performance were established. A research collaboration might be worth considering, but the evidence does not justify treating the component as a dependable installed asset.
The committee then examines the shared connection claim. A supplier says all three proposals will allow the campus to obtain power three months earlier. The schedule gives no basis for that claim. Equipment delivery remains the longest prerequisite. Forecasting quality, shifted computing demand and an experimental component address different parts of the system; none automatically changes that delivery date. The committee requests a dependency-specific explanation instead of accepting a combined promise.
The completed review contains three distinct recommendations. Continue testing the forecast because its small-sample result is promising. Evaluate scheduling feasibility and service quality because the peak and simplified cost calculations are reproducible. Keep the component at a research-evidence stage until the missing manufacturing and service questions are addressed. These recommendations are deliberately different because the proposals have different evidence and consequences.
The review also records what would change its conclusions. Proposal A could weaken if unseen-period errors rise or forecasts arrive too late. Proposal B could weaken if deadlines rule out the assumed flexibility or overhead erodes the benefit. Proposal C could strengthen with independently checked production and service evidence. The connection schedule could change if a verified delivery commitment changes. An evidence-based decision is revisable for named reasons rather than permanently optimistic or pessimistic.
Finally, Riverbend establishes a reporting boundary. Public statements may describe the completed fictional evaluations and their limits. They may not say the campus has reduced outages, completed an upgrade or adopted an autonomous grid system, because none of those things occurred in the packet. This last step is important: accurate calculations can still be used to tell an inaccurate story. Verification must include the sentence that reaches the reader.
Back to contents · Continue: 13. Build an evaluation brief that someone else can check
13. Build an evaluation brief that someone else can check
A useful brief starts with one decision, not an open-ended request to “use AI for energy.” For example: should a forecasting assistant proceed to a larger advisory trial? The brief names the intended users, the setting and the output being assessed. It also states which decisions remain outside the trial. This keeps a modest analytical evaluation from quietly turning into operational authority.
Next, specify the baseline and the unit of success. A forecast might be compared with an existing method on the same timestamps. A design tool might be assessed by the quality of candidates that survive an independent test. A scheduling tool might be evaluated on feasible peak reduction while preserving task deadlines. “More intelligent” is too vague to serve as the result. A reader should be able to identify what would count as an improvement before seeing the model’s output.
The evidence packet should preserve original inputs, units, time zones, missing values and the date the information became available. This matters when daylight-saving changes, different reporting intervals or revised historical data could distort comparisons. Record the model and process version as well. Otherwise a later reviewer may be asked to reproduce a result with a different system and mistake the difference for a calculation error.
Authority belongs in the brief as an explicit field. Is the system producing a draft, a recommendation, an alert or a direct action? Who reviews it, and what independent controls constrain the permitted use? The brief should not include sensitive infrastructure details in a public AI service merely because the system asks for more context. Use approved data-handling arrangements and involve the responsible security team when operational information is involved.
A failure plan should describe how the evaluation recognises an unusable result. Examples include absent data, inconsistent units, a forecast outside the allowed horizon or a recommendation that violates a stated constraint. At the educational level, the repair is to mark the output unusable and return it for qualified review. The brief should not encourage a model to invent missing measurements or bypass an established safeguard to finish the task.
Measurement also needs a denominator. Count all relevant cases, including rejected outputs, unavailable service and periods of human intervention. If a report says “nine successful cases,” ask whether there were ten attempts or one thousand. If it says “the model reduced review time,” ask whether preparation and correction were included. A benefit that disappears when the whole workflow is counted is not the same as a durable improvement.
Before expansion, identify the evidence needed for the next role. A tool that performs well on archived records may progress to a shadow evaluation where outputs are compared with actual practice without controlling it. A later operational role would require the appropriate engineering, safety, security and regulatory processes. These are not boxes that a general-purpose chatbot can certify. The responsible organisation must establish the assurance appropriate to its system.
The finished brief should end with a decision and its limits. Continue, revise, pause or stop, with the reason and the next evidence needed. Avoid an indefinite pilot that keeps collecting impressive examples without resolving a question. The purpose of evaluation is to reduce uncertainty enough for a justified decision. Sometimes the most valuable result is discovering early that a proposed role is not supported.
Back to contents · Continue: 14. Independent practice, with complete answers
14. Independent practice, with complete answers
Exercise One: a fictional site consumes 6 MW for two hours and 3 MW for four hours. Calculate total energy, average power and peak power over the six-hour period. Explain why a connection selected only from the average could be inadequate for the supplied profile. Keep the units visible and do not assume any ability to alter the site’s demand.
Answer One: the first interval contributes 6 × 2 = 12 MWh and the second contributes 3 × 4 = 12 MWh, giving 24 MWh. Average power is 24 divided by six, or 4 MW. Peak power is 6 MW. The profile includes two hours above the average, so a four-megawatt limit would not accommodate it as stated. This is an arithmetic diagnosis, not a design specification for a real connection.
Exercise Two: a service initially performs 50,000 accepted tasks at 4 Wh per task. A revised service uses 2.5 Wh per accepted task and performs 90,000 tasks. Calculate energy before and after, the per-task improvement and the change in total consumption. State the extra information needed before claiming an environmental improvement.
Answer Two: baseline consumption is 200,000 Wh, or 200 kWh. Revised consumption is 225,000 Wh, or 225 kWh. Energy per task falls by 1.5 Wh, a 37.5% improvement relative to four Wh. Total consumption rises by 25 kWh, or 12.5%. An environmental claim also needs appropriate electricity and emissions boundaries, comparable task quality and consideration of other relevant resource effects; the supplied arithmetic alone does not establish it.
Exercise Three: a battery has a 2 MW discharge limit and 9 MWh of usable output energy. Can it supply 1.5 MW for six hours? Can it supply 3 MW for two hours? Assume the battery starts with the stated energy and ignore all limits other than those explicitly supplied.
Answer Three: the first request requires 1.5 × 6 = 9 MWh and stays within the 2 MW limit, so it exactly fits the idealised packet. The second needs only 6 MWh but exceeds the power limit, so it does not fit. The first result has no remaining energy margin in the simplified packet; it should not be promoted into a real reliability guarantee without further conditions and professional assessment.
Exercise Four: an AI-assisted study falls from nine months to four. A permit takes seven months and equipment takes 11 months; all begin together. Construction takes three months after all prerequisites, followed by one month of commissioning. What happens to the earliest completion date? Then repeat with an original study duration of 14 months.
Answer Four: with the original nine-month study, the equipment path already controls the prerequisites. Completion is month 15 before and after the study improvement: 11 + 3 + 1. With a 14-month original study, completion was month 18 and becomes month 15 after the improvement. The study reduction is ten months in the second scenario, but the complete-project reduction is three months. State both the local change and the final outcome.
Exercise Five: a maintenance model reviews 500 observations. Twenty are confirmed cases. It flags 40 observations, of which 15 are confirmed cases. Calculate recall, precision and missed cases. Explain why the 40 alerts are not the same as 40 prevented failures.
Answer Five: recall is 15 divided by 20, or 75%. Precision is 15 divided by 40, or 37.5%. Five confirmed cases are missed and 25 alerts do not correspond to a confirmed case. An alert is an information output. Preventing a failure would require a justified intervention and evidence about what would otherwise have happened. The packet establishes detection results, not prevented failures.
Exercise Six: a report says, “Our model reduced forecast error by 20%, therefore the grid needs 20% less capacity.” Identify the missing inference without calculating a new capacity number. Write a replacement sentence that preserves the legitimate finding.
Answer Six: the report has not shown how forecast error maps to required capacity under the relevant reliability assessment, demand patterns, network limits and resource characteristics. A defensible replacement is: “The model reduced the specified forecast-error measure by 20% on the evaluated dataset; any effect on capacity requirements needs a separate system assessment.” The replacement is useful because it preserves progress while refusing to invent a consequence.
Back to contents · Continue: 15. Questions readers should be able to answer
15. Questions readers should be able to answer
Would a genuinely superintelligent system remove energy bottlenecks?
That remains hypothetical. Greater intellectual capability might improve designs, experiments, coordination or analysis. It would not by definition eliminate the need for evidence, materials, manufacturing, construction or legitimate authority. A claim about a future system should identify which bottleneck it is expected to change and what would demonstrate that change. Intelligence is relevant to physical problems because it can improve decisions and methods, not because the label makes physical dependencies disappear.
Can AI make an existing grid more useful before new lines are built?
Some applications may improve information, planning or the use of existing assets. The outcome depends on the specific technology, system conditions and approved implementation. Readers should distinguish improved visibility of capacity from newly constructed capacity, and a proposed optimisation from validated operating practice. The appropriate question is what usable service has been demonstrated within established limits. An educational article cannot determine those limits for a real network.
Does more accurate renewable forecasting make variable generation firm?
A forecast describes expected production; it does not guarantee production. Better forecasting can support preparation and comparison of options, but resource availability remains a separate issue. A useful report names forecast accuracy, uncertainty and any demonstrated downstream benefits separately. It should not present a statistical improvement as a physical conversion of one resource into another. Dependable service may require a portfolio of resources and arrangements whose adequacy must be assessed in context.
Is a large energy total enough to supply a data centre?
No single annual total answers the complete question. The facility needs an appropriate rate of delivery at its location and at the times it operates, together with the required reliability and supporting infrastructure. An annual energy purchase can be meaningful without proving an hourly match. Ask for the connection, load profile, supply arrangement and status of the relevant assets. Their absence is a reason for further investigation, not permission to invent them.
Is lower electricity use always the best outcome?
The objective depends on the service. An intervention might use slightly more electricity while enabling a valuable result or reducing a consequential risk. Another might reduce consumption by degrading the service people need. Evaluation should show the resource change together with quality, reliability and other relevant effects. This prevents both uncritical celebration of consumption and the assumption that every reduction is automatically beneficial. The decision belongs to the people accountable for those trade-offs.
How should a student assess an impressive AI-energy headline?
Translate it into a sentence with a subject, measure, date and boundary. Then label the evidence as measured, estimated, projected or proposed. Reproduce any arithmetic with units and check whether the conclusion changes category between sentences. A laboratory result becoming a commercial promise, a forecast score becoming a savings claim, or a capacity total becoming guaranteed energy is a signal to slow down. The aim is accurate interpretation rather than reflexive agreement or dismissal.
What would count as real progress?
Real progress is a supported improvement in a specified task or service: better candidates reaching independent tests, forecasts that remain useful on new periods, justified maintenance decisions, feasible schedules that preserve service quality, or completed assets performing as promised. Each result should include its limitations and the work still needed. Energy systems benefit from intelligence when ideas survive contact with measurement, physical constraints, accountable decisions and the everyday needs of the people they serve.
Continue the Super Intelligence series
For the wider definitions and evidence framework, return to What Is Super Intelligence?. For the supply-side constraints on computing, read Why Power Becomes the Limiting Resource. For efficiency measurement, continue with The Economics of Intelligence per Watt. For the investment question, read Can AI Demand Finance the Future of Energy?.