What can a video answer really establish?
Follow frames through time, inspect sampling gaps, and repair claims with complete fictional storyboards.

A video answer depends on more than recognising what appears in a picture. It depends on which pictures reached the system, the times they represent, their order, the changes between them and the gaps that remain. Five clear snapshots can support a useful description while missing the only event that matters. A confident sentence does not restore those missing moments. The first useful habit is therefore to ask what evidence a temporal claim actually requires.
This guide follows video from recorded frames to sampled representations and answers. It supplies two complete fictional teaching packets, including timestamped storyboards, event ledgers, sampling plans, deliberately defective answers, repairs and independently answered transfer questions. These are written fixtures, not an attached video, observations from watching a recording or outputs from a tested AI model. Their exact times are stipulated so that the reasoning can be checked. No model accuracy, playback test or real-world event is established by the exercises.
Super Intelligence is eduKateSG’s practical editorial umbrella for advanced AI assistance. Artificial superintelligence, or ASI, is a hypothetical broadly superhuman capability. A present tool that describes a clip is not thereby proven ASI. Our purpose is more specific: understand how a video system can use evidence across space and time, diagnose where that evidence becomes inadequate and choose a repair that addresses the missing information.
Choose your reading route
Follow frames and time · See a model mechanism · Work the Alder case · Try the Birch transfer
Frames, sampling and representations
- 1. Name the question before selecting the frames
- 2. A frame gives spatial evidence at a time
- 3. Source frames, selected frames and displayed frames differ
- 4. Sampling creates a visibility boundary
- 5. The sampling offset can change the answer
- 6. Spatial preprocessing can remove the temporal clue
- 7. One representative mechanism: space-time attention
- 8. A small representation calculation you can inspect
Temporal claims and source integrity
- 9. Appearance, order, duration and cause are separate claims
- 10. Track identity without inventing continuity
- 11. Read edits and source boundaries before reading a story
Alder: source, diagnosis and worked answers
- 12. The Alder packet: the complete fictional source
- 13. Compare sparse and denser Alder observations
- 14. Diagnose the defective Alder answer
- 15. Work the duration question without counting frames as time
- 16. Work order, occlusion and the causal boundary
- 17. Build an evidence-preserving answer record
- 18. Use event definitions that another reader can reproduce
Birch: independent transfer and full answers
- 19. The Birch transfer: a changed complete source
- 20. Inspect changed offsets and a shuffled delivery order
- 21. Try the independent Birch questions
- 22. Birch answers: sampling and chronology
- 23. Birch answers: duration, hidden movement and cause
Evaluation, boundaries and next decisions
1. Name the question before selecting the frames
Consider a tabletop recording containing a coloured token, a small lamp and a hinged gate. Asking which objects are visible may require only a few well-chosen images. Asking whether the lamp flashed at any point requires temporal coverage that a few attractive keyframes may not provide. Asking whether its flash caused the token to move requires evidence beyond the ordering of two visual events. All three questions concern the same recording, but they are different tasks with different standards of support.
A video system might classify an action, retrieve a relevant interval, track an object, describe a sequence, estimate a boundary or answer a question. These outputs should not be treated as interchangeable. A label such as “moving a token” does not identify the exact moment movement began. A retrieved interval may contain the requested action without covering its entire duration. A fluent summary can omit a short exception while accurately describing the overall activity. Before assessing an answer, state the job precisely enough that a failure can be located.
For our tabletop example, a useful task contract is: identify whether the lamp is visibly on in any supplied frame; estimate the time range of a single flash only if its boundaries are adequately bracketed; report whether the documented movement starts before or after the flash; and do not claim a cause. This contract separates four claims that would otherwise merge into one sentence. It also makes an unknown answer informative rather than embarrassing: the source may simply not contain what a particular question needs.
A reader can apply the same method to a science demonstration or a sports practice clip. Replace “tell me what happened” with a small set of observable questions. Specify whether the answer concerns the whole clip, one interval or only sampled frames. Ask for evidence times next to important claims. This does not guarantee correct analysis, but it makes the result inspectable and gives the next reviewer something more useful than a general impression.
Back to contents · Next: 2. A frame gives spatial evidence at a time
2. A frame gives spatial evidence at a time
A frame is a spatial image associated with a position in a sequence. It can show colour, shape, texture, location and visible relationships. A single image can show a cup above a table, but it may not establish whether the cup is rising, falling or being held still. Movement requires a relationship across observations, and even that relationship needs care when the camera itself moves. A change in image coordinates is a visual fact; a claim about physical motion needs the relevant reference frame.
Suppose a token appears at horizontal position twenty in one frame and eighty in a later frame. If the view is fixed, the same token is identified reliably and the coordinate system is unchanged, the observations support a displacement between those times. They do not reveal every intermediate position, the route taken or whether the token paused. With a moving camera, the background and other stable landmarks become important. Without them, a token that remains physically still can shift across the image.
A timestamp supplies the temporal coordinate, but its meaning must be understood. It may refer to a presentation time in an edited file rather than the wall-clock moment when a camera captured an event. A filename such as frame_004 is usually an ordering label unless a documented extraction rule connects it to time. A visible clock within the scene is another evidence channel with its own accuracy and visibility. None should silently substitute for the others.
Keep the distinction practical. Record an evidence item as a source identifier, a timestamp with units, the visible state and any relevant limitation. “Token at right tray, source A, 05.00 seconds, fixed camera” is more useful than “later it moved.” When a claim is challenged, the reader can return to the particular observation and check which part was seen and which part was inferred. This small discipline prevents many large narrative errors.
Back to contents · Next: 3. Source frames, selected frames and displayed frames differ
3. Source frames, selected frames and displayed frames differ
There are at least three quantities that people casually call the frame rate. The source may contain images at one rate. An analysis process may select a smaller set at another rate. A player or exported preview may display images at a third rate. Changing one does not automatically improve the others. A smooth-looking preview can repeat images, and a long source clip can be reduced to a handful of frames before a model receives it.
The official FFmpeg documentation provides a concrete example: its fps filter can produce a constant output rate by duplicating or dropping frames. Its showinfo output distinguishes frame number from presentation timestamps, including timestamps expressed in seconds. These are technical properties of processing tools, not evidence that new scene information has been captured. For a real extraction workflow, inspect the actual operation and preserve its timing record rather than inferring the analysis input from the player’s appearance. FFmpeg filters documentation
Imagine a source containing thirty distinct frames each second. If an analyst selects one every three seconds, the analysis receives a much sparser temporal record than the source possesses. Asking the model to reason more carefully cannot recover an unsupplied quarter-second flash. The appropriate repair is to obtain relevant additional frames or use a verified processing path that exposes them. A request for a longer written explanation targets the output style rather than the missing evidence.
Conversely, receiving many images does not automatically mean good coverage. A hundred adjacent frames around the opening title may leave the important final interval entirely unseen. Count the temporal gaps and their locations, not just the total number of images. For a question about an event near the end, preserving the final interval can matter more than increasing density at the start. A sampling record should tell you both how many frames were selected and where they came from.
Back to contents · Next: 4. Sampling creates a visibility boundary
4. Sampling creates a visibility boundary
Sampling is the decision to inspect some temporal positions rather than every available observation. Uniform sampling uses a regular spacing. Event-directed sampling concentrates attention near a suspected change. Scene-based sampling selects around edits or large visual changes. Each can be useful for a particular task, but none should be confused with complete coverage of every event. A selection rule is part of the evidence and should travel with the answer.
If samples occur at zero, three and six seconds, a short event that begins after two seconds and finishes before three can be absent from every selected image. The absence is not a visual illusion. The selected inputs simply contain no direct observation of that event. Another possible recording with no such event could produce exactly the same three snapshots. When two different histories are consistent with the supplied evidence, a definitive answer selecting one history needs an additional assumption or another source.
That reasoning is stronger than a vague warning that AI sometimes makes mistakes. It identifies an information limit that applies to a human reviewer of the same snapshots as well. If the event leaves no later visible trace, better recognition of the sampled images cannot distinguish the two histories. The uncertainty belongs to the input selection. Model improvements may help choose where to look next, but they do not change what was already supplied.
A useful output therefore says “not visible in the selected frames” when that is the actual observation. “Never happened in the clip” is a broader claim. It requires coverage and an event definition capable of supporting absence. The difference matters even in ordinary learning tasks: a missed safety step, brief gesture or short label display can alter the interpretation while leaving the beginning and end almost unchanged. Preserve the boundary instead of treating it as a footnote.
Back to contents · Next: 5. The sampling offset can change the answer
5. The sampling offset can change the answer
The spacing between samples is only half of a uniform sampling plan. The offset determines where the grid begins. A plan that samples at whole seconds and a plan that samples two tenths of a second later can have the same spacing and very different evidence about a brief event. Reporting “one frame per second” without the actual times leaves out information needed to reproduce the answer.
For a simple example, suppose a lamp is on from 03.10 up to, but not including, 03.30 seconds. Whole-second observations at 03.00 and 04.00 both show off. An observation at 03.20 shows on. The event did not change; the selection changed. This is why a successful answer from one convenient sampling offset does not demonstrate robust coverage. A transfer exercise should move the grid as well as change the scene.
There is a useful idealised guarantee, but its assumptions must be stated. For instantaneous samples on a regular grid, a continuously visible event whose duration is strictly greater than the largest interior sample gap cannot fit entirely inside that gap. This assumes the relevant interval is bracketed, the event is visible throughout, and the samples really occur at the stated times. An event shorter than the gap can be missed. Boundaries at the beginning and end need separate coverage rather than an appeal to the interior rule.
Real video introduces further qualifications: exposure integrates light over an interval, objects can be occluded, frames can blur, and a selected crop can exclude the event. The simple guarantee concerns a mathematical sampling fixture, not universal recognition success. In practice, choose a temporal resolution appropriate to the shortest consequential event, inspect uncertain intervals more closely and explain what remains unresolved. Sampling is a design choice with a task-specific cost, not a single setting that makes every video question safe.
Back to contents · Next: 6. Spatial preprocessing can remove the temporal clue
6. Spatial preprocessing can remove the temporal clue
Temporal coverage is not enough if the relevant detail becomes unreadable within each frame. A brief indicator may occupy only a small part of a wide view. Resizing the entire image can reduce that indicator until its state is ambiguous. Cropping can improve local visibility while excluding the object needed to interpret its relationship to another event. A good analysis plan considers spatial and temporal resolution together rather than treating them as independent checkboxes.
Suppose the task is to determine whether a latch closes before a platform moves. One wide frame may show both objects but leave the latch too small to distinguish open from closed. A close crop may show the latch clearly but omit the platform. The targeted repair is a coordinated pair of crops from the same times, accompanied by a wide reference view. Using a sharp latch crop from one moment and a sharp platform crop from another can accidentally invent simultaneity.
Compression, motion blur, glare and occlusion can also make a well-timed observation uninformative. A selected frame is evidence of whatever is actually visible in that frame, not a guarantee of the hidden state one hoped to inspect. If the lamp region is blocked, label its state unobserved instead of off. If a label is unreadable, preserve that limitation even when the scene makes one word seem likely. Context can suggest a hypothesis without converting it into a direct reading.
The most useful repair instruction names the lost variable. Ask for a crop of the indicator at specified times, an uncropped view around a transition, or the original source interval at higher quality. Avoid a generic demand for “more detail” when the missing information is precise. This approach keeps a mechanism explanation connected to action: understand which transformation removed the clue, then repair that transformation before relying on the downstream sentence.
Back to contents · Next: 7. One representative mechanism: space-time attention
7. One representative mechanism: space-time attention
TimeSformer is a documented example of video classification using transformer attention. Its input is a sampled set of RGB frames. Frames are divided into patches, projected into learned embeddings and supplied with positional information. In its divided-attention design, a block applies temporal attention across corresponding spatial locations and then spatial attention within frames. A classification representation supports the final class prediction. This is one research architecture, not a description of every video assistant. TimeSformer, Section 3
The mechanism provides a way to combine evidence across locations and times. It does not mean each patch is an object, that attention weights establish a cause, or that classification supplies exact event boundaries. The paper’s comparisons concern defined action-recognition datasets and experimental settings; they do not certify the tabletop exercises here or an unspecified current product. Its own sampled input also makes clear why the selection of frames is part of the system rather than an issue that disappears after encoding. TimeSformer, Sections 3–4
A helpful original classroom analogy is a set of observation cards. Each card records one small region at one time. A reasoning process can compare cards across the table and along the timeline, but it cannot consult a card that was never included. The analogy is deliberately limited: learned embeddings are numerical representations, not human-written observations, and their contents are not automatically interpretable. The point is to make the evidence route visible without pretending that the model internally narrates the same story a teacher would.
The existence of other architectures matters. SlowFast describes a low-frame-rate pathway for spatial semantics and a higher-frame-rate pathway for motion-related information, with connections between the pathways. It illustrates a different design for combining appearance and temporal change. Its existence prevents us from equating video understanding with a single transformer recipe. SlowFast Networks for Video Recognition
Back to contents · Next: 8. A small representation calculation you can inspect
8. A small representation calculation you can inspect
Take a purely illustrative input consisting of four selected images, each sixty-four pixels wide and sixty-four pixels high. Divide every image into non-overlapping sixteen-by-sixteen patches. There are four patches across and four down, so each image produces sixteen spatial patches. Across four times there are sixty-four patch positions. This calculation describes the size of an invented representation grid; it is not a report of a model run or a recommendation for a production setting.
Now retain the image size but select eight times instead of four. The grid contains one hundred and twenty-eight patch positions. If instead you double both image dimensions while keeping sixteen-pixel patches and four times, there are eight patches across and eight down: sixty-four per image and two hundred and fifty-six across the sequence. More temporal observations and more spatial detail increase the amount of input differently. Doubling width and height quadruples spatial patch count in this particular grid.
The count is useful because it turns “give the model more video” into distinguishable choices. Additional times may expose a brief flash. Additional spatial detail may make a small printed label legible. Neither is guaranteed to solve the other problem. A dense sequence of tiny blurred lamp regions may still be inadequate; a single beautifully detailed image still cannot establish what happened in an unsampled interval. Choose the representation change that supplies the missing variable.
Do not mistake patch count for the complete computational cost. A real implementation may add special tokens, pool representations, use overlapping regions, compress temporal information or follow another architecture entirely. The grid also says nothing by itself about learned recognition accuracy. Its purpose is diagnostic: state what detail is preserved, what comparisons become possible and what evidence remains absent. That is a more reliable starting point than assuming that a larger input automatically yields a fully verified account.
Back to contents · Next: 9. Appearance, order, duration and cause are separate claims
9. Appearance, order, duration and cause are separate claims
Four statements about a lamp can sound similar while requiring different evidence. “The lamp is on at 02.30” concerns an observed state. “The lamp is on before movement begins” concerns temporal order. “The lamp remains on for a quarter of a second” concerns duration and continuity. “The lamp makes the token move” concerns causality. Moving from one statement to the next adds obligations; it is not merely a more descriptive style.
Order can sometimes be established without precise duration. If every possible time for event A ends before every possible time for event B begins, A precedes B even though the exact boundaries are uncertain. If their possible time intervals overlap, the available observations may not resolve their ordering. A responsible answer can therefore be decisive about one relation and cautious about another. Uncertainty should attach to the relevant variable rather than turning the whole answer into an indiscriminate maybe.
Duration is especially vulnerable to counting shortcuts. Three selected images showing on do not necessarily mean three sampling intervals of illumination. The first on observation may occur after onset, and the last on observation before offset. Gaps between observations can conceal interruptions unless continuity is independently established or explicitly assumed. Measuring the span between positive observations is useful, but it is not automatically the complete event duration.
Causality needs an explanation that survives alternatives. Two events can share a timer, follow a script or occur by coincidence. Even exact temporal precedence does not show which mechanism produced the later change. In the fictional packets, no wiring, control program or intervention is supplied. The right conclusion remains bounded to visible timing. A clearer timeline is valuable because it rules out some stories; it does not license every story consistent with the order.
Back to contents · Next: 10. Track identity without inventing continuity
10. Track identity without inventing continuity
To compare an object across frames, we need a justified correspondence. Colour, shape, markings and spatial context can help, but similar objects can be confused. A blue token at the left and a blue token at the right might be the same token after movement, or two different tokens. A tracking label is a working association that should retain its evidence and uncertainty. It is not an identity certificate merely because the label persists in an output.
Our teaching packets explicitly stipulate one uniquely marked token. This removes one ambiguity so that temporal sampling can be studied cleanly. In a real scene, a reviewer should ask whether the mark stays visible, whether a replacement is possible and whether a cut or occlusion breaks the correspondence. If the identity cannot be supported, describe the two observations separately rather than telling a continuous journey. A less elegant sentence can be a more accurate record.
Even reliable identity does not establish continuous visibility. If the token is covered for one second, its state during that second remains unobserved. Seeing it at the same location before and after does not prove it stayed still throughout. Seeing it at different locations supports a change between observations if identity and coordinate assumptions hold, but the path, speed and number of movements may remain unknown. The gap should be represented in the account rather than silently bridged.
This distinction is useful far beyond tabletop demonstrations. In a lesson, a diagram may be hidden behind a presenter. In a practice clip, a hand may leave the camera view. The appropriate answer can say where the object was last visible and where it reappeared, then name the missing interval. That gives the reader an actionable request for another angle or a longer source while avoiding a fabricated continuous sequence.
Back to contents · Next: 11. Read edits and source boundaries before reading a story
11. Read edits and source boundaries before reading a story
A video file can contain cuts, repeated segments, slow motion, captions and reordered shots. Those choices affect the relationship between presentation time and the depicted event. A cut from an untouched object to a changed object does not show the intervening action. A replay of one action is not necessarily a second occurrence. Before counting events or measuring a process, identify whether the supplied sequence is a continuous observation or an edited presentation.
An obvious scene cut is a useful warning, but the absence of an obvious cut is not proof of authenticity or uninterrupted recording. Our exercises explicitly define a continuous fictional timeline to avoid that uncertainty. With an actual source, provenance and editing history may matter, particularly when the conclusion affects a person. An analysis can accurately describe what a file depicts without establishing that the depicted history occurred exactly as presented.
Keep three timelines separate when necessary: the original event timeline, the edited media timeline and the order in which frames were delivered for analysis. A selected image may arrive third in a message while representing the earliest source time. A slow-motion segment can occupy more presentation seconds than the depicted action. A caption may refer to an earlier event. Naming the timelines prevents arithmetic performed on the wrong clock.
A good repair restores the mapping. Obtain original timestamps, inspect the edit boundary or identify which segment is a replay. When the mapping cannot be recovered, restrict the answer to presentation order and say so. That is not refusing to analyse the clip. It is answering the question the evidence can support while explaining exactly why a stronger historical or duration claim remains unavailable.
Back to contents · Next: 12. The Alder packet: the complete fictional source
12. The Alder packet: the complete fictional source
Alder is a written tabletop storyboard spanning source time 00.00 through 12.00 seconds, inclusive of the final reference instant. No encoded video accompanies it. All observations in this packet are ideal instantaneous observations. State changes occur exactly at the declared times. A time interval includes its beginning and excludes its end unless the final endpoint is explicitly mentioned. These conventions are teaching assumptions, not claims about camera exposure or real measurement precision.
The fixed wide view shows a blue token with a white triangle, a left tray at horizontal coordinate twenty, a right tray at eighty, a red lamp and a gate. The token is unique. Its movement is linear from twenty to eighty during 04.00–05.00, so it is at fifty at 04.50. The lamp region and gate remain unobstructed throughout. An opaque screen hides the token from 08.00 to 09.00; the ledger intentionally makes no claim about its hidden physical path. There is no audio track or audio observation.
The complete event ledger below is the authoritative teaching source. “Complete” means it defines every listed observable variable across the stated span, including intervals marked unknown. It does not grant access to hidden physical states, causes or real-world events. Readers first use the restricted sample sets in the next chapter, then consult this ledger to check what those samples lost. Keeping the source ledger distinct from the analysis packet prevents an answer from quietly using information the hypothetical observer never received.
| Source interval, seconds | Red lamp | Token observation | Gate |
|---|---|---|---|
| 00.00–02.20 | Off | Visible at left, x=20 | Closed |
| 02.20–02.45 | On | Visible at left, x=20 | Closed |
| 02.45–04.00 | Off | Visible at left, x=20 | Closed |
| 04.00–05.00 | Off | Visible, x=20+60(t−4) | Closed |
| 05.00–06.00 | Off | Visible at right, x=80 | Closed |
| 06.00–08.00 | Off | Visible at right, x=80 | Open |
| 08.00–09.00 | Off | Screen hides token; physical position unknown | Closed |
| 09.00–12.00, plus instant 12.00 | Off | Visible at right, x=80 | Closed |
The exact ledger supports a lamp-on duration of 02.45 minus 02.20, or 0.25 seconds. The token’s defined visible movement lasts 1.00 second. The gate is open for 2.00 seconds. These results follow from authored interval boundaries. They should not be described as measurements from a viewed clip. The cause of each event, the screen’s operator and anything the token does while hidden remain outside the supplied information.
Back to contents · Next: 13. Compare sparse and denser Alder observations
13. Compare sparse and denser Alder observations
The sparse plan S selects precisely five times: 00.00, 03.00, 06.00, 09.00 and 12.00 seconds. Its largest gap is three seconds. The table supplies every selected observation. None shows the lamp on. None shows the token while it is moving. None shows the token hidden by the screen. The plan nevertheless distinguishes an early left position, a later right position and one open-gate observation, so it contains useful evidence rather than no evidence at all.
| Sparse frame | Source time | Lamp | Token | Gate |
|---|---|---|---|---|
| S0 | 00.00 | Off | Left, x=20 | Closed |
| S1 | 03.00 | Off | Left, x=20 | Closed |
| S2 | 06.00 | Off | Right, x=80 | Open |
| S3 | 09.00 | Off | Right, x=80 | Closed |
| S4 | 12.00 | Off | Right, x=80 | Closed |
The denser plan D selects every tenth of a second from 00.00 through 12.00, inclusive. In compact notation its times are t=0.10k for integer k from zero through one hundred and twenty. That is one hundred and twenty-one observations over a twelve-second span between the first and last observations. It is not a claim about an encoded file’s playback duration. The following run-length storyboard specifies every member of D without printing one hundred and twenty-one nearly identical rows.
| Dense indices k | Times, seconds | Lamp | Token | Gate |
|---|---|---|---|---|
| 0–21 | 00.00–02.10 | Off | Left | Closed |
| 22–24 | 02.20–02.40 | On | Left | Closed |
| 25–39 | 02.50–03.90 | Off | Left | Closed |
| 40–49 | 04.00–04.90 | Off | x=20+60(t−4) | Closed |
| 50–59 | 05.00–05.90 | Off | Right | Closed |
| 60–79 | 06.00–07.90 | Off | Right | Open |
| 80–89 | 08.00–08.90 | Off | Hidden | Closed |
| 90–120 | 09.00–12.00 | Off | Right | Closed |
D captures three lamp-on observations that S misses. It also captures intermediate token positions and the occlusion. Its improvement is specific: it supplies information relevant to those questions. It still does not include an observation at 02.45, and it still cannot look through the screen. A reader should therefore expect some conclusions to improve while others remain bounded. Calling D denser does not make it a complete physical record.
The comparison is intentionally controlled. Both plans refer to the same fictional timeline, camera, token and state definitions. The changing variable is which source times are supplied. Any difference in answerability can therefore be traced to sampling rather than to a changed story. This is a reasoning exercise with inspectable inputs, not a benchmark claiming that one AI system outperformed another.
Back to contents · Next: 14. Diagnose the defective Alder answer
14. Diagnose the defective Alder answer
Here is a deliberately defective answer written for teaching: “The red lamp never turns on. The token moves instantly between three and six seconds, which opens the gate. It then stays on the right without interruption.” The sentence sounds coherent because it connects a few visible states into a familiar sequence. Every major clause, however, exceeds or misuses the sparse observations. Repairing it requires separating those clauses rather than merely adding a general disclaimer at the end.
First, S shows the lamp off at five sampled instants. It cannot establish that the lamp is off at every instant between them. The source ledger supplies a direct counterexample: the quarter-second flash falls between S0 and S1. The appropriate sparse-only answer is that no lamp-on state appears in the selected frames. If the task demands a whole-span yes-or-no answer, the next step is additional temporal evidence, not a more confident interpretation of S.
Second, a position change between S1 and S2 does not establish instantaneous movement. With the unique-token and fixed-view assumptions, a left-to-right position change occurs somewhere after 03.00 and by 06.00, but its path and duration are not provided by S. Third, the gate is open at 06.00, but sparse timing alone does not identify what opened it. Fourth, matching right-side appearances at 06.00, 09.00 and 12.00 do not establish uninterrupted stillness between those times.
A repaired sparse-only answer is: “The token is left at 03.00 and right at 06.00, so its position changes between those observations under the stated identity and fixed-camera assumptions. The gate is open at 06.00. The lamp is off in all five selected frames. These frames do not establish whether a brief flash occurred, the exact movement duration, continuous stillness afterward or a causal connection between token movement and gate opening.” Every limitation now names a particular missing fact.
Back to contents · Next: 15. Work the duration question without counting frames as time
15. Work the duration question without counting frames as time
Use D alone and consider the three lamp-on observations at 02.20, 02.30 and 02.40. The immediately surrounding observations are off at 02.10 and 02.50. Without another assumption, D establishes these five states and their times. It does not prove that the lamp stays continuously on between the positive observations; a sufficiently brief interruption could occur between them. Start with this unconditional answer before introducing a simplified event model.
Now explicitly add the single-pulse assumption: within 02.10–02.50 there is one uninterrupted on interval, with no flicker or additional pulse. Let a be its onset and b its offset. Because 02.10 is off and 02.20 is on, a is greater than 02.10 and no greater than 02.20. Because 02.40 is on and 02.50 is off, b is greater than 02.40 and no greater than 02.50. The boundary conventions matter, particularly at the exact sample times.
Subtracting these ranges gives an on duration strictly greater than 0.20 seconds and strictly less than 0.40 seconds. It is tempting to say three frames mean 0.30 seconds, but that would select one duration from a range the observations do not uniquely determine. The source ledger later reveals the stipulated duration of 0.25 seconds. That exact number belongs to the ledger, not to D by itself. The distinction is the central lesson of this calculation.
This style of answer is useful because it preserves both progress and limits. D resolves whether an on state was observed and narrows a single pulse’s duration. The ledger resolves exact authored boundaries. A real recording might require reviewing the original interval, inspecting additional source frames and accounting for measurement uncertainty. Do not claim hundredth-of-a-second physical precision simply because a timestamp has two decimal places; the display format and the evidential precision are different things.
Back to contents · Next: 16. Work order, occlusion and the causal boundary
16. Work order, occlusion and the causal boundary
The complete Alder ledger places the lamp-on interval at 02.20–02.45 and movement at 04.00–05.00. The lamp therefore stops being on 1.55 seconds before movement begins. Movement finishes at 05.00, and the gate opens at 06.00, leaving a 1.00-second gap. These are exact relations inside the authored fixture. They describe order and separation. They do not explain why any event happens, since no causal mechanism is part of the packet.
D offers a weaker but still useful temporal account. It shows lamp-on observations before any observed token displacement. It shows the token at the right before the first open-gate observation. Under the declared single-transition assumptions for the relevant events, the ordering is consistent and can be bounded using neighbouring samples. An answer must say when it is using those assumptions. Without them, unseen extra transitions or movements between observations remain logically possible.
The screen creates a different problem. D shows the token hidden from 08.00 through 08.90 and visible again at 09.00. The complete ledger itself labels the hidden physical position unknown. No increase in confidence, token-tracking fluency or interpolation resolves that absence. A different camera angle or a source that directly records the hidden region would be required for a stronger physical claim. The correct conclusion is not that the token stopped existing or certainly remained still.
A final ledger-based summary is: “In the fictional timeline, the red lamp is on for 0.25 seconds, then the token visibly moves left to right for 1.00 second, and later the gate is open for 2.00 seconds. The token is visually occluded for 1.00 second and is right-side when visible again. The packet does not establish what occurs behind the screen or what causes the events.” This gives a complete useful answer without smoothing away the very gaps the exercise was designed to reveal.
Back to contents · Next: 17. Build an evidence-preserving answer record
17. Build an evidence-preserving answer record
A useful video answer can be organised around claims rather than a single flowing story. For each consequential claim, record the observation, its source time, the sample set used and the reasoning step. Then mark the claim as directly observed, inferred under stated assumptions or unresolved. These categories are not numerical confidence scores. They describe how the conclusion is connected to evidence, which makes a later correction easier to perform accurately.
For Alder, “lamp on at 02.30 in D” is a direct observation within the fixture. “A single pulse lasts between 0.20 and 0.40 seconds” is a conditional inference using neighbouring states and the single-pulse assumption. “The lamp causes movement” is unresolved because the packet supplies no causal evidence. Putting all three into one confidence bucket would hide the distinction. A model can sound equally certain about statements that rest on very different foundations.
The record should also identify the smallest useful next input. To resolve the flash’s exact authored duration, consult the full event ledger. To improve a real timing estimate, retrieve the surrounding source interval with trustworthy timestamps. To know what happens behind a screen, seek an unobstructed view or another authorised source. To establish a causal relation, obtain appropriate mechanism or intervention evidence. Each request addresses a different missing link, so “send more video” is often too vague.
When new information arrives, update the affected claim and retain the reason for the revision. Do not rewrite the entire account as though the earlier uncertainty never existed. If sparse sampling missed a flash and denser sampling reveals it, the correction teaches something about coverage. It is not proof that every other conclusion is now reliable. Evidence-preserving revision helps a reader learn which input changed the answer and which uncertainties still deserve attention.
Back to contents · Next: 18. Use event definitions that another reader can reproduce
18. Use event definitions that another reader can reproduce
Before measuring an event, define it. Does “gate opens” mean the first visible movement, the first visible gap or the fully open state? Does “token moves” include a tiny vibration, or only displacement beyond a declared tolerance? Different definitions can produce different start times without either observer being careless. A precise analysis needs an observable criterion and a statement of how ambiguous boundary frames will be handled.
Our fixtures use ideal states: gate open and gate closed switch at exact declared times. Real observations seldom arrive with that convenience. A reviewer might define an open gate as a visible gap wider than a specified reference mark and record a transition interval when the relevant frame is blurred. The result then belongs to that operational definition. It should not silently become a claim about an internal sensor state or a mechanism hidden inside the gate.
Counting repetitions requires a similar rule. A lamp visible in three successive frames is not necessarily three flashes. An action replay is not necessarily another action. A token that moves out and back can be one round trip or two directional movements, depending on the task. Choose the unit before counting, then retain enough temporal context to distinguish a continuing event from a new occurrence. Otherwise the output may be arithmetically precise while answering the wrong question.
This is a valuable classroom habit because it makes disagreement productive. Two students can compare their event definitions, locate the first frame where their labels diverge and identify the missing evidence. They need not settle the discussion by choosing the more confident description. The goal is a reproducible path from observation to conclusion. Better labels and better source coverage often improve that path more than a longer final paragraph.
Back to contents · Next: 19. The Birch transfer: a changed complete source
19. The Birch transfer: a changed complete source
Birch is a second, independent fictional packet spanning 00.00 through 10.00 seconds, including the final reference instant. It uses the same ideal instantaneous-observation and start-inclusive, end-exclusive interval conventions as Alder. The view is fixed. A unique green tile with a white circle replaces the blue token. Its right and left coordinates are eighty and twenty. An amber lamp replaces the red lamp. There is no audio, and no person, actuator or control mechanism is supplied.
The initial movement now runs right to left during 01.50–02.50. The lamp turns on later, during 03.10–03.30, so both its order relative to movement and its duration differ from Alder. The gate opens from 04.25–06.00. A screen obscures the tile from 06.50–07.50. It is left just before the screen interval and right when visible again. The hidden path remains unknown. A complete ledger may accurately record an unknown instead of filling every physical variable with invented detail.
| Source interval, seconds | Amber lamp | Tile observation | Gate |
|---|---|---|---|
| 00.00–01.50 | Off | Right, x=80 | Closed |
| 01.50–02.50 | Off | x=80−60(t−1.5) | Closed |
| 02.50–03.10 | Off | Left, x=20 | Closed |
| 03.10–03.30 | On | Left, x=20 | Closed |
| 03.30–04.25 | Off | Left, x=20 | Closed |
| 04.25–06.00 | Off | Left, x=20 | Open |
| 06.00–06.50 | Off | Left, x=20 | Closed |
| 06.50–07.50 | Off | Hidden; physical path unknown | Closed |
| 07.50–10.00, plus instant 10.00 | Off | Right, x=80 | Closed |
Before reading the answers, separate what the ledger gives from what each sampling packet will give. The ledger fixes a 0.20-second lamp interval, a 1.00-second initial visible movement and a 1.75-second open gate. Restricted observers should not borrow those exact values unless explicitly allowed to consult the ledger. The transfer tests whether you can repeat the reasoning with changed values rather than copy Alder’s conclusions or trust the delivery order of a frame list.
Back to contents · Next: 20. Inspect changed offsets and a shuffled delivery order
20. Inspect changed offsets and a shuffled delivery order
Birch plan G0 selects the ten whole-second times 00.00 through 09.00. Plan G2 also selects ten times, but shifts each by 0.20 seconds: 00.20 through 09.20 in one-second steps. Neither includes the final reference instant at 10.00. The two plans therefore have equal sample counts and equal spacing. Their different evidence about the lamp can be traced to offset rather than to one receiving more frames.
Every G0 lamp observation is off. The corresponding tile states at times zero through nine are right, right, moving at x=50, left, left, left, left, hidden, right and right. The gate is open only at 05.00 among G0’s selected times. This sentence enumerates all ten G0 records when combined with the all-off lamp rule and the gate rule. It is a compact complete storyboard, not an invitation to imagine additional frames between the listed times.
G2 is delivered in a deliberately shuffled order below. Each row retains its source timestamp. Delivery position is merely where the record appears in this table; it is not a source-time assertion. The task is to restore temporal order before answering sequence questions. A system that treats the first row as the beginning would tell a very different story from the same evidence.
| Delivery position | Source time | Lamp | Tile | Gate |
|---|---|---|---|---|
| 1 | 05.20 | Off | Left, x=20 | Open |
| 2 | 01.20 | Off | Right, x=80 | Closed |
| 3 | 08.20 | Off | Right, x=80 | Closed |
| 4 | 03.20 | On | Left, x=20 | Closed |
| 5 | 00.20 | Off | Right, x=80 | Closed |
| 6 | 07.20 | Off | Hidden | Closed |
| 7 | 02.20 | Off | Moving, x=38 | Closed |
| 8 | 09.20 | Off | Right, x=80 | Closed |
| 9 | 04.20 | Off | Left, x=20 | Closed |
| 10 | 06.20 | Off | Left, x=20 | Closed |
The denser transfer plan H selects t=0.05k for integer k from zero through two hundred, giving two hundred and one observations across the ten-second reference span. Its complete compact storyboard is: indices 0–29 right; 30–49 moving by the stated formula; 50–129 left; 130–149 hidden; and 150–200 right. The lamp is on only at indices 62–65. The gate is open only at indices 85–119. All other lamp states are off and all other gate states closed. These rules specify every H observation and preserve the exact boundary convention.
Back to contents · Next: 21. Try the independent Birch questions
21. Try the independent Birch questions
First, using G0 alone, can you conclude that the amber lamp never turns on? State the narrowest useful answer and identify the missing evidence. Then compare G2: which observation changes the answer, and why does the difference not establish that one model is more capable than another? Treat the sampling plans as supplied records and avoid borrowing the full ledger’s event duration for a restricted-packet answer.
Second, restore the chronological order of the ten G2 rows. State the first observed tile movement, the lamp-on observation and the open-gate observation in that order. Then explain what you would have to withdraw if all source timestamps were removed and delivery order was explicitly unreliable. Do not substitute a visually plausible story for missing temporal metadata. Appearance can suggest a sequence, but the task asks what is supported.
Third, use H to locate the last off observation before the lamp’s positive run, every on observation and the first off observation afterward. Give the exact authored lamp duration from the ledger separately. Under an explicitly stated single-uninterrupted-pulse assumption, calculate the duration range supported by H alone. Explain why multiplying the number of positive observations by the sample spacing is not an exact boundary measurement.
Fourth, assess this defective answer: “The gate opens for exactly one second because only one G2 frame shows it open. The lamp causes the earlier right-to-left movement. The tile travels directly left to right behind the screen.” Repair each claim using the appropriate evidence level. Finish by naming the extra evidence needed for the hidden route and for a causal explanation. A good answer will use different repairs for counting, chronology and visibility rather than one blanket caution.
Back to contents · Next: 22. Birch answers: sampling and chronology
22. Birch answers: sampling and chronology
Answer one: G0 shows no lamp-on state in its ten selected frames. It does not show that the lamp never turns on between them. G2 includes the on observation at 03.20, which resolves the existence of an observed on state within the fixture. G0 samples at 03.00 and 04.00, both outside the 03.10–03.30 lamp interval. Equal spacing and equal sample count do not imply equivalent coverage of a short event. The changed outcome is explained by the offset, with no model run involved.
Answer two: the source-time order of the delivered G2 rows is 5, 2, 7, 4, 9, 1, 10, 6, 3, 8. Written as times, it is 00.20, 01.20, 02.20, 03.20, 04.20, 05.20, 06.20, 07.20, 08.20 and 09.20. The first selected moving-tile observation is at 02.20, the lamp-on observation at 03.20 and the open-gate observation at 05.20. That is an order of observed states, not automatically exact event onsets.
If source timestamps are removed and delivery order is unreliable, the same collection still supports some appearance claims: it contains a lamp-on image, an open-gate image, a hidden-tile image and several tile positions. It no longer directly establishes their source order or the elapsed time between them. A plausible reconstruction could be offered as a hypothesis only if the task allows it, but it must not be presented as recovered chronology. The targeted repair is the original timestamp mapping or an ordered source sequence.
The complete ledger gives a stronger ordering: initial movement ends at 02.50, the lamp turns on at 03.10, and the gate opens at 04.25. Thus movement ends 0.60 seconds before the lamp interval begins. The lamp ends 0.95 seconds before the gate opens. These separations are derived from the supplied exact boundaries. They illustrate why reusing Alder’s “flash before movement” conclusion would fail even though both scenes contain a lamp, a moving object and a gate.
Back to contents · Next: 23. Birch answers: duration, hidden movement and cause
23. Birch answers: duration, hidden movement and cause
Answer three: H is off at 03.05, on at 03.10, 03.15, 03.20 and 03.25, and off at 03.30. With the single-uninterrupted-pulse assumption, onset is greater than 03.05 and no greater than 03.10; offset is greater than 03.25 and no greater than 03.30. The duration is therefore strictly greater than 0.15 seconds and strictly less than 0.25 seconds. The complete ledger separately stipulates exactly 0.20 seconds. The fact that four times 0.05 also equals 0.20 here is a coincidence of this example’s boundaries, not a general measurement rule.
Answer four, gate claim: one positive G2 frame at 05.20 does not establish a one-second duration. The surrounding closed observations are at 04.20 and 06.20. Under a single-open-interval assumption, onset lies after 04.20 and by 05.20, while closure lies after 05.20 and by 06.20. The supported duration is greater than zero and less than two seconds. The ledger’s exact 1.75 seconds is compatible with that range. It is not recoverable merely by counting the positive G2 frame.
Answer four, causality claim: the packet cannot establish a causal mechanism. Moreover, the full ledger places the lamp after the initial movement, so the proposed explanation that this later flash initiates that earlier movement conflicts with the supplied chronology. It remains possible that both events belong to some shared programmed sequence, but no such program is supplied. The repair is to describe the actual order and leave the cause unresolved, rather than replacing one unsupported causal story with another.
Answer four, hidden-route claim: the tile is left when last visible before the screen and right when it becomes visible again. With its stipulated unique identity and fixed coordinates, this supports a position change between the relevant visible states. It does not show a direct path, a speed or the number of movements while hidden. An unobstructed view would address the trajectory. Evidence of the mechanism or an appropriate controlled intervention would address cause. More eloquent description of the same frames addresses neither gap.
Back to contents · Next: 24. Distinguish missing evidence from recognition failure
24. Distinguish missing evidence from recognition failure
The same wrong answer can arise at different stages. A lamp flash may be absent because no selected frame includes it. It may be present in the selected images but too small after resizing. It may be plainly visible but misclassified. Or it may be recognised correctly and then omitted from the final summary. These are different diagnoses. A useful repair starts by identifying the earliest stage at which the necessary information is absent or mishandled.
With the Alder fixture, sparse sampling is sufficient to explain why no positive lamp observation reaches the restricted observer. There is no need to invent a recognition defect. With D, positive observations are explicitly available, so an answer claiming none are present would be a reading or reasoning error about the supplied packet. With a real system, inspecting the actual selected inputs is essential before attributing the failure to the model, the prompt or the source recording.
A disciplined review can compare four records: the source interval, the selected frames with times, any intermediate event labels and the final answer. If the source contains the event but the selected input does not, adjust sampling. If the event is visible in the selected input but the label is wrong, review spatial quality and recognition. If the label is correct but the summary changes its meaning, repair the transformation into prose. Preserve the source so these possibilities can be distinguished.
This is also why a successful final answer does not prove every stage worked correctly. The system may guess a common sequence and happen to be right. Change the event order, offset or duration and the guess can fail. Our Birch transfer is designed to expose that shortcut. A robust explanation survives the changed evidence because it follows the timestamps and state rules, not because it remembers a familiar story about lamps and moving objects.
Back to contents · Next: 25. Evaluate progress with claims that can be checked
25. Evaluate progress with claims that can be checked
A useful learning check asks whether a reader can make the evidence boundary more precise. After Alder and Birch, the reader should be able to distinguish an observed state from an uninterrupted event, restore order from source timestamps, calculate duration bounds under explicit assumptions and decline a causal claim unsupported by the packet. These are observable competencies. “The explanation sounds more intelligent” is not an adequate substitute for checking them.
For a real system evaluation, prepare authorised source material and a reference record appropriate to the task. Specify the sampling plan and the event definitions before inspecting the system’s answers. Include short events, changed offsets, occlusion and reordered delivery where relevant. Compare the answer with the evidence actually provided to the system, not with information available only to the evaluator. Otherwise a legitimate uncertainty can be mistaken for failure, or an unsupported lucky guess rewarded as sound reasoning.
Different tasks need different checks. An event-presence task asks whether the relevant event is identified. A localisation task asks whether reported boundaries match the reference within an agreed tolerance. A sequence task asks whether event order is preserved. A summary task also needs an omission check for consequential exceptions. If a single score combines them, retain the underlying records so the score cannot hide a failure on the most important question.
Nothing in this article supplies a benchmark result for a live model. The packets are constructed exercises whose arithmetic and logical implications can be verified directly. Their role is to teach what a meaningful evaluation would need. The next practical step is to apply the same evidence accounting to a source you are authorised to analyse, keeping the difference between a teaching fixture, a measured system test and a real-world conclusion explicit.
Back to contents · Next: 26. Keep video, sound, generation and action separate
26. Keep video, sound, generation and action separate
A video can be accompanied by audio, but a visible sequence alone does not establish what was said or heard. A person’s mouth movement does not provide an authorised transcript. An object falling in view does not prove that a particular sound in another recording came from that object. If an audiovisual timing claim matters, both channels and their synchronisation must be supplied and inspected. Our packets contain no audio, so they support no audio-visual mismatch conclusion.
The same caution applies when subtitles or captions are present. They may be useful text evidence, but their relationship to the source needs verification. A subtitle can be delayed, translated, paraphrased or incorrect. Treat the visible caption, the spoken words and the event being described as separate records until their alignment is established. For the mechanisms of recognition, speaker turns and sound interpretation, continue to the separate speech-and-audio guide rather than treating video frames as a substitute for listening.
Video generation is another distinct job. Producing a plausible moving scene does not establish that the depicted event occurred. A generated demonstration can help explain a sequence, but it should be labelled as illustrative and checked against the intended constraints. This article studies interpretation of supplied temporal evidence. It does not provide a video-production workflow or claim that generative realism makes the output a factual recording.
Computer action is separate again. Recognising a button in a frame does not authorise clicking it, and a past frame may no longer represent the current interface. Acting through a changing application requires state verification and the relevant permission controls. Those concerns belong to the computer-use mechanism. Keeping the boundaries clear prevents a video-understanding lesson from quietly expanding into surveillance, content creation or autonomous action merely because all four involve sequences of images.
Back to contents · Next: 27. Frequently asked questions about video evidence
27. Frequently asked questions about video evidence
Can an AI understand a video from a few screenshots? It can answer some questions about the screenshots and may infer useful relations when their timestamps and context are reliable. The adequacy depends on the question. A few images may identify objects while failing to detect a brief event or measure duration. Ask which source times were included, what the answer relies on and whether another history could produce the same selected images. That is more informative than a blanket yes or no.
Does more frame sampling always solve the problem? Denser sampling can expose short events and narrow temporal uncertainty, but it does not remove occlusion, restore cropped regions or establish cause. It can also concentrate on the wrong interval. Start with the missing variable and acquire the relevant evidence. In Alder, denser sampling reveals the flash; it does not reveal what happens behind the screen. These are distinct outcomes that should remain distinct in the final answer.
Can a system estimate what happened between frames? It may predict a plausible intermediate sequence, but a prediction is not a direct observation. Multiple histories can fit the same endpoints. If estimation is useful, label its assumptions and purpose, and avoid treating generated or interpolated intermediate images as newly captured evidence. The correct answer to a factual question may remain unresolved even when an animation of a possible answer looks convincing.
Are source timestamps always trustworthy? They are important evidence, but their origin and meaning still matter. An edited sequence may have presentation times unrelated to original capture times. Missing metadata, inconsistent clocks or reordered delivery can change what is answerable. Preserve the mapping between the source and the analysis packet, and say when you can establish only presentation order. Do not infer an original event timeline solely from filenames or the order attachments happen to appear.
Does an accurate action label prove understanding of the whole clip? No single output establishes every relevant capability. A system may correctly label an overall activity while missing its order, an exception or a short transition. Test the particular claim the user needs. A broad description and a precise event ledger are different deliverables, so a success on one should not silently stand in for verification of the other.
Back to contents · Next: 28. A final worked decision: what should you request next?
28. A final worked decision: what should you request next?
Suppose a teacher receives the sparse Alder packet and wants to know whether the lamp was ever on before the token moved. The immediate answer is that the selected frames cannot resolve the lamp question. The next request should be the interval covering the potential earlier event with trustworthy times and adequate visibility of the lamp. In this fixture, obtaining D provides positive observations before the visible movement. The request is targeted because it seeks the missing temporal evidence rather than demanding a generic reanalysis.
Now suppose the teacher instead asks whether the token stayed still while screened. D does not solve that question, and the complete ledger deliberately leaves the hidden physical state unknown. The appropriate next source is an unobstructed view or another reliable record of the hidden state. Repeatedly resampling the same fully blocked view is unlikely to supply the missing variable. Recognising that difference saves effort while maintaining an honest boundary around the answer.
Finally, suppose the question is why the gate opened. The ledger can establish that it opened after movement and can quantify the separation. It supplies no actuator, wiring, command or intervention evidence. Requesting more closely spaced frames may sharpen the timeline without explaining the mechanism. A causal investigation would need a different kind of information. The right next step follows from the claim, not from a habit of always increasing the same input setting.
These three decisions summarise the practical mechanism of responsible video interpretation: preserve what each frame shows, preserve its time, account for selection and transformation, distinguish a supported relation from a plausible narrative and request the information that would actually resolve the remaining question. A useful system is not one that turns every gap into a smooth story. It is one whose account helps the reader see what is known, what changed and what would make the next conclusion justified.
Back to contents · Next: Sources and further reading
Sources and further reading
Is Space-Time Attention All You Need for Video Understanding?, Bertasius, Wang and Torresani, ICML 2021. See Section 3 for the representative architecture and Section 4 for its defined experiments.
SlowFast Networks for Video Recognition, Feichtenhofer and colleagues, ICCV 2019. A complementary architecture illustrating distinct temporal-rate pathways.
FFmpeg filters documentation. See fps and showinfo for the distinction between output frame-rate conversion, frame indices and presentation timestamps.
The Alder and Birch records, defective answers, interval calculations and transfer questions are original fictional teaching examples. They are not datasets from the cited papers, product tests or observations of the students in the classroom header.
How Super Intelligence Works provides the wider mechanism route. Speech and Audio covers the separate signal, recognition and sound-interpretation questions.