I keep finding the same dead end.

There is a Korean video I genuinely want to watch. Sometimes for enjoyment, sometimes to learn. It was made for native speakers, so there are no English subtitles. People talk over each other. Music sits underneath the dialogue. A name, joke, or cultural reference arrives and disappears at full speed.

At that point I am not missing a convenience. I am locked out of the thing itself.

Translation is only one part of the problem

Written translation already asks for judgment. The same Korean line can land differently in English depending on who is speaking, what was omitted, how formal the relationship is, and what happened ten seconds earlier.

Real media adds a stack of problems before translation even begins:

  1. separate speech from music, effects, and other voices;
  2. decide who spoke and where one utterance ends;
  3. transcribe Korean without turning uncertainty into confident text;
  4. translate meaning and tone into natural English;
  5. time the result closely enough that it still belongs to the scene.

Doing all five while the source is still playing creates a cruel constraint: wait for more context and the captions become late; commit early and they can become wrong. Research on simultaneous speech translation treats this as a real quality–latency tradeoff. Google’s work on live caption translation found that very low latency and strong final quality could still produce unstable text as the system repeatedly revised its guess. The paper is unusually direct about that tension.

The existence of impressive speech and translation models does not make that combined problem solved. Our narrower claim is personal and testable: we have not found a consumer workflow that consistently gives us accurate, well-timed Korean-to-English captions for arbitrary native media with overlapping, noisy audio at a latency that preserves the experience.

We should have to prove that claim against fixed clips. Repeating it as marketing is not enough.

The distance is strange because the connection is enormous

Korean culture is not a niche waiting to be discovered. Netflix says more than 80% of its members have watched Korean content. It reports that Squid Game drew 95% of its early viewership from outside Korea, while KPop Demon Hunters passed 600 million views and became Netflix’s most-watched original title. Netflix’s K-content study, its Squid Game reporting, and its KPop Demon Hunters retrospective make the scale difficult to dismiss.

The connection is economic as well as cultural. The Office of the U.S. Trade Representative estimates that U.S.–Korea goods and services trade totaled $239.6 billion in 2024. The countries are already deeply entangled.

The AI relationship is becoming more explicit too. OpenAI and Kakao have announced a ChatGPT experience for KakaoTalk, while OpenAI’s partnerships with Samsung and SK include Korean AI infrastructure and enterprise deployment. Kakao announcement; Samsung and SK partnership.

None of those numbers validates 1321. They explain why the remaining language barrier feels so disproportionate to the connection on either side of it.

“Just learn Korean” is not an answer

Learning is worth doing. It is also a long road.

The U.S. State Department places Korean in its “super hard” category and assigns an 88-week training objective for professional speaking and reading proficiency. That estimate describes intensive training for English-speaking Foreign Service professionals; it is not a universal countdown, and it should not be applied mechanically to heritage learners or speakers of Chinese and other languages. Still, it is a useful signal of the distance involved. The State Department’s own guidance lists the objective.

For a learner, real native media is exactly where progress should become rewarding. Instead, it is often where dictionary knowledge collapses under speed, sound, slang, omitted context, and multiple voices. Subtitles can be the bridge between studying a language and actually living with it.

So 1321 studies the video first

The answer for this Build Week is not to pretend we can make every live translation perfect. It is to stop forcing the hardest media through a cold, one-pass guess.

1321 is being built so autonomous agents can investigate real media — audio, frames, context — and leave an understanding that can be checked, withheld when weak, and reused. Watch aids (timed Korean and English), explanations, and lines worth learning are Apply outputs of that understanding, not the product itself. When the evidence is weak, the system should expose the gap instead of polishing a fluent mistake.

The Build Week proof bar is deliberately small: one difficult Korean clip, roughly 30 seconds to one minute, handled end to end. That is how we prove the spine this week — not the definition of what 1321 is. We will compare the same source cold and prepared. We will keep failures visible. We will not publish benchmark numbers until the clips, expectations, and scorer exist.

This Journey is the record of that attempt.

Not a retrospective written after everything worked. The decisions, wrong turns, evidence, and unanswered questions while we build.