Today we landed more media tools on the owned path: acoustic triage, frame sampling, typed multimodal admission, one budgeted re-study pass, on-screen OCR, and anonymous speaker/overlap evidence.
That is easy to overstate.
It does not mean 1321 can see, hear, or understand a video. It means a granted child can run a bounded operation over exact bytes, leave a receipt, and leave a weak range weak instead of turning it into fluent text.
Receipts are not understanding
A producer receipt shows which operation ran over which bytes. It does not show that the producer was right.
An acoustic class is not lyric understanding. A delivered frame is not scene understanding. OCR text is not the identity of what was shown. An anonymous speaker cluster is not a person, and it is not a correct transcript.
A spawn is a different kind of trap. The scheduler can accept a model-authored task without giving that task hearing, vision, research, shell, or computer-use. A role label, a mitosis animation, or a worker count is not a tool. If the matching host operation was never granted and used, the child did not inspect the source.
What the new tools actually cover
The landed rungs share one boring contract: a path-free request, a least-privilege grant, host-owned execution, an immutable observation and receipt, cold audit, and an abstention path when the input is weak or the grant is missing.
Acoustic triage can partition a granted range and keep non-dialogue ranges from acquiring caption text when policy and evidence agree. Frame sampling can deliver authorized PNG bytes to a granted child. Multimodal admission can carry speech, coverage, and cite-only context without upgrading a weak state just because another modality arrived. One attenuated speech re-study can target an exact weak range or one audited speaker-overlap cell. OCR can attach provisional on-screen text as cite-only context. Speaker analysis can preserve anonymous turns and overlap as coverage facts.
None of that authorizes the claim that the system understood the clip. Structural readiness still checks integrity and coverage, not meaning. The first scored Bet G receipt on the hard clip already showed that more machinery can preserve less meaning. That result stays visible.
The Studio trap
Studio has the same honesty problem in a prettier form.
The canvas can show a worker forming and a mitosis wire while the recorded log keeps that worker in spawning. A small projector can say whether a birth had a real lead window or happened in one recorded instant. Useful, and still only a reading of replay facts.
It does not invent a pre-spawn intent the runtime never emitted. It does not prove live swarm cognition. And it does not grant the child a media tool. If the scheduler never issued something like media.frames.sample, media.frames.ocr, or media.speakers.analyze, the UI should not imply the agent looked at the source.
Returning hashes or filenames while the model never receives pixels is the same mistake. Project what the journal recorded. Do not dress it up.
What is still missing
Conditional separation (raw versus stem comparison for exact triggered ranges) is next, not done. Typed schemas and validation are not a separation host. Web research and bounded computer-use are later. We still cannot claim end-to-end understanding of a whole video.
The product is still language intelligence for real-world media. Korean-to-English is the beachhead. Captions and timed text are outputs of a study, not the category name.
We added more honest ways to inspect owned media. We did not earn “the agents understood the video.”