76.9% preserved. Answered every time, 3 of them wrong.
We test our own translations
We turn Korean video into English. To check that work we locked a set of clips, wrote the right answers down in advance, and had 2 reviewers grade the output without knowing which system produced it.
We publish the result even when it goes against us. On this clip our prepared build scored below a stripped-down version of itself, because it stayed silent on 6 of 13 moments and silence counts as meaning lost.
38.5% preserved. Answered 7 times, 2 of them wrong.
- Score gap
- −5 units
- Coverage
- 1 / 3 clips
- Reviewers
- 2
- Serious errors
- 0
Prepared preserved fewer
2 not yet scored
Blinded human review
None on either side
Coverage
The test set contains 3 clips. One has system outputs, human review, and scores.
The other 2 are held out, frozen and ready but never run, so they cannot flatter or hurt the result. They stay excluded until they complete every step.
- Source
- Original audio pinned
- Key
- Reference answer written before any run
- Output
- A system ran and produced captions
- Review
- People judged it blind
- Score
- Judgements counted into these numbers
- Test clips
- 3
- Scored
- 1
- Pending
- 2
Results
Without preparation the system preserved 10 of 13 moments, against 5 for prepared. Prepared declined to answer 6 times, which is more than the 5 moment gap between them.
Of the moments each system did answer, the cold run was wrong 3 times and the prepared run 2. Every decline counts as meaning not preserved.
No preparation
- Shown & right
- 10
- Shown & wrong
- 3
- Held back
- 0
- Missed
- 0
Prepared
- Shown & right
- 5
- Shown & wrong
- 2
- Held back
- 6
- Missed
- 0
We tried an idea and it failed
We had an idea for translating Korean family terms better. We wrote it down first so we could not move the goalposts, then ran 3 clips with it and without it. It lost every comparison we finished.
We also came up 3 comparisons short of the 9 we promised ourselves, so the test never counted and the idea was never shipped. These clips were graded by us, not by the outside reviewers who graded the test above.
- Runs booked
- 18
- Runs captured
- 15
- Comparisons
- 6 / 9
Method
Two blinded reviewers judged 13 moments that were chosen and frozen before any system ran.
Answering right, answering wrong, holding back, and missing a line stay separate outcomes. No model graded the output.
- Reviewers
- 2
- Review blinded
- Yes
- Model grader
- None
Audit
The page checks itself every time it is built. Each line below is worked out from the saved evidence, not typed in by hand.
If any one of them fails, the page does not build. That is why nothing here is outstanding. What we have not measured yet is on Coverage, not hidden in this list.
See all 15 checks
Before anything ran
- The test set was locked before the recording was made
- The lock time on the score matches the lock time on the test set
- The rule change was registered before its first run started
- The rule change wrote down its pass mark and left the result blank
How it ran
- Both systems were judged on the same clip
- All 18 booked runs left a call receipt, including the ones that failed
- Every booked run made exactly one call, with no retries
How it was graded
- Reviewers did not know which system produced what they graded
- No model graded any output
The numbers add up
- No preparation: 10 right plus 3 wrong plus 0 held back plus 0 missed makes 13
- Prepared: 5 right plus 2 wrong plus 6 held back plus 0 missed makes 13
The files cannot be swapped
- The recording matches the fingerprint the score points at
- The reviewer grades match the fingerprint the score points at
- The locked test set matches the fingerprint the score points at
- The grades were made against that exact recording, not a later one