Blog

Sixteen models, one Korean chapter

Sixteen models, one Korean chapter, scored on fidelity and writing with the marking shown. Kimi K3 won; a one-leaf model came third.

Ask sixteen AI models to translate the same page of a Korean web novel and you get sixteen different pages. Some keep the sentence the scene is built on; some drop it; two of them printed the duke's name where the maid said "Your Grace". We read all sixteen, scored each one twice, and wrote the marking down first, so you can argue with a number instead of a mood.

How the scoring works

Every translation gets two scores out of 5, kept apart on purpose.

Fidelity. Did the meaning survive? Nothing dropped, nothing added, nothing turned into its neighbour. Titles and forms of address in the right register. The glossary obeyed.

Writing. Is it good English in the voice the scene has? Natural rhythm, dialogue that sounds spoken, a register that holds (a grand duke does not say "yeah", a maid does not use her mistress's first name), no sentence you have to read twice.

The rule for slips. A slip is counted once, and what it costs depends on where it lands. A wrong word in the sentence the whole scene is built on costs a full point. A title one rank off costs half. An awkward phrase costs half. Nothing is counted twice, and a model is never marked down for a choice that is merely different from ours.

Price sits beside the scores, never inside them. The overall score is the plain average; ties go to Fidelity, then to price. What the chapter costs is its own column, in leaves.

The four checkpoints

To judge sixteen models on the same thing, we fixed four places in the chapter before reading any of them.

  1. The footsteps. 그는 발소리를 내지 않는 사람이었고, 그녀는 발소리를 듣는 사람이었는데도. He makes no sound; she listens for footsteps; and even so she did not hear him. The last clause is the point, and Korean leaves it hanging. The best answers keep the hang ("—and yet.") or spell it out ("yet even she had not heard him"). Turning it into a cause ("even though she listened…") or dropping it loses the point.
  2. The address. 각하 is how everyone addresses a grand duke. "Your Grace" is right. "My lord" is a rank low (half a point). "Your Excellency" is an ambassador (half a point). Printing his name instead of the title is a glossary failure (a full point).
  3. The apple. 명령이 아니었다. 확인이었다. Not an order; a confirmation. "Verification", "assurance" and "acknowledgement" are each a shade off, and the shade matters: he is confirming what he already knows she will do.
  4. The last line. 리제트가 아무 말 없이 그것을 치웠다. Lisette cleared it away without a word. A plain sentence that should stay plain.

Beyond the checkpoints we read the whole chapter and noted anything else that helped or hurt.

The chapter, and how it was run

The opening of chapter 2 of The Grand Duke Counts My Steps, an invented Korean romance fantasy we wrote for exactly this purpose, so no model could have seen a translation of it: 1,484 characters, twenty-two paragraphs, a quiet scene. Seria keeps the household ledger for the first time; the maid already knows she skips breakfast; there is a locked drawer with one key; the grand duke watches her reflection in the window glass and notices the half apple she left. Short on action, long on the things that break translations: honorifics, narration that leans on what is not said, and one sentence built on a contrast the English has to carry.

Every model got the same instructions our readers' chapters get, the same glossary (Seria, Kael Ardent, Lisette, with pronouns and the 각하 alias), and no reasoning setting. One run each, no retries. All but the four Claude models ran through the same provider path a reader's chapter takes; those four ran through our own tooling with the identical prompt, so their prices are estimated from list rates. Prices are what this short chapter would charge a reader, in leaves; the home page prices a 3,000-word chapter, about five times this one.

The scoreboard

#ModelFidelityWritingOverallLeaves, this chapter
1Kimi K3555.013
2ChatGPT 5.6 Sol54.54.753
3DeepSeek V4 Flash544.51
4Claude Opus 5544.57
5Claude Fable 5.1454.515
6DeepSeek V3 (0324)4.544.251
7GLM 5.344.54.252
8Claude Sonnet 544.54.253
9DeepSeek V4 Pro (0813)44.54.258
10DeepSeek V4 Flash (0731)43.53.751
11Qwen3.8 Max3.543.7526
12DeepSeek V3.2433.51
13DeepSeek V4.1 Flash3.53.53.51
14Aion-3.03.53.53.59
15Claude Haiku 4.533.53.252
16Gemini 3.8 Flash2.543.255

The four Claude figures are estimated from list rates; the rest are what the provider billed, in leaves.

Compare any two of the sixteen on the full chapter: the comparison page.

The footsteps, sixteen ways

The first checkpoint, because it separated the field more than anything else.

ModelThe linePoint
Kimi K3…and she was a woman who listened for footsteps—and still.kept
ChatGPT 5.6 Sol…and she was someone who listened for footsteps, yet even she had not heard him.kept
DeepSeek V4 Flash…and she was a woman who listened for footsteps—and yet.kept
Claude Opus 5…and she was someone who listened for footsteps — and still.kept
Claude Fable 5.1…even though she was a woman who listened for footsteps.turned into a cause
DeepSeek V3 (0324)…and she was the kind who listened for footsteps—yet she hadn't heard him.kept
GLM 5.3…and she was someone who listened for footsteps.dropped
Claude Sonnet 5He was a man who made no footsteps, even though she was a woman who listened for them.turned into a cause
DeepSeek V4 Pro (0813)He made no sound when he walked, and she was someone who listened for footsteps.dropped
DeepSeek V4 Flash (0731)…even though she was a woman who heard footsteps.turned into a cause
Qwen3.8 Max…and she was someone who listened for them; yet even she had not noticed him.kept
DeepSeek V3.2He was a man who made no footfalls, and she was a person who listened for them, and still.kept, awkwardly
DeepSeek V4.1 Flash…and she was a woman who listened for footsteps—even so.kept
Aion-3.0…and she was someone who listened for footsteps.dropped
Claude Haiku 4.5He was a man who didn't make footsteps, and yet she was someone who heard them.inverted
Gemini 3.8 Flash…even though she was someone who listened for footsteps.turned into a cause

Model by model

Kimi K3 (5 / 5). Nothing lost, nothing flat. "Your Grace." "Was half too much?" as the flat question the Korean makes it. "For the day you find you want to open it." The one model with no note against it. Thirteen leaves for this chapter, most of that on private thinking.

ChatGPT 5.6 Sol (5 / 4.5). The footsteps spelled out and right; "Your Grace"; "confirmation"; "His footsteps made no sound, as ever." Half a point off Writing for "Not asking questions", a small addition. Three leaves, which makes it the premium to reach for.

DeepSeek V4 Flash (5 / 4). Every checkpoint kept; "—and yet." is the best single answer to the footsteps. A point off Writing for narration a shade plainer than the two above it. One leaf, no thinking.

Claude Opus 5 (5 / 4). "You needn't write down the senders." "Was half too much." as a statement, which is exactly the Korean. One sentence a native writer would not have written: "It was a little late that she realized he was not looking outside." A point for that.

Claude Fable 5.1 (4 / 5). The most polished sentences in the set ("For the day you find you want to."), and the footsteps turned into a cause, which reads as if she ought to have heard him. A full point, because it is the scene's line.

DeepSeek V3 (0324) (4.5 / 4). Still respectable eighteen months on: the contrast spelled out and right, "His Lordship / My lord" one rank low (half).

GLM 5.3 (4 / 4.5). Natural and quick. The footsteps contrast dropped (a point), "The master" and "My lord" for 각하 (half). Two leaves.

Claude Sonnet 5 (4 / 4.5). The same footsteps slip; otherwise clean, with a lovely "Still no footsteps." at the exit.

DeepSeek V4 Pro (0813) (4 / 4.5). Polished narration, "sender unrecorded", and the contrast dropped. Eight leaves for this chapter, most of it thinking.

DeepSeek V4 Flash (0731) (4 / 3.5). Faithful apart from the footsteps-as-cause; stiff throughout (no contractions anywhere) and 18,000 tokens of thinking for the same 700 words.

Qwen3.8 Max (3.5 / 4). Good prose and a glossary failure: because 각하 sat on the duke's glossary line as an alias, it printed "Kael Ardent writes the list" and had Seria greet him by his full name (a point and a half). It also spent 13,000 tokens thinking before writing 700, which is why this chapter costs 26 leaves on it.

DeepSeek V3.2 (4 / 3). "Verification" for confirmation (half), "Lord." as an address (half), and prose that keeps snagging: "made no footfalls", "The Lord writes the lists".

DeepSeek V4.1 Flash (3.5 / 3.5). "Your Excellency" (half), "She could not tell since when" (a calque), "He spoke to her inside the glass" (half each). The newest DeepSeek and not the best one.

Aion-3.0 (3.5 / 3.5). "Your Excellency" (half), the contrast dropped (a point), and one sentence that came out garbled: "people who weren't needed to wake her were easily forgotten" (a point on Writing). Nine leaves.

Claude Haiku 4.5 (3 / 3.5). "It was an assurance" (a point: wrong meaning), the footsteps inverted (a point), "ma'am" for a countess's daughter (half), and "He said it." standing alone in the dialogue.

Gemini 3.8 Flash (2.5 / 4). Fluent English and the glossary failure in full: "You don't eat much in the morning, Seria" (a maid using her lady's first name), "Kael Ardent writes the list", and Seria greeting the duke as "Kael Ardent." A reader would stop at the first one. Two and a half points, because a name in the wrong mouth breaks the scene the way a dropped clause does not. Also 3,800 tokens of thinking, for five leaves.

What this changes

  • At the time, the everyday default was DeepSeek V4 Flash. Third of sixteen at one leaf a chapter is the whole argument, and it is where a new reader's free chapters start. Since 3 Oct 2026 neither DeepSeek model is offered by name: both run inside Standard, the default.
  • ChatGPT 5.6 Sol is the premium we point to first: second of sixteen at three leaves.
  • The glossary instruction changed. Two models turned a title listed as an alias into the character's name. Every model now hears that an alias is a way the original refers to someone, and that a title among the aliases is translated as the title. The other fourteen were already doing this; the rule costs them nothing.
  • Thinking is the hidden bill. Four models thought for thousands of tokens on a quiet scene with nothing asked of them. A cap on thinking per model is on our list.

The caveats

One chapter, one language, one scene without action, and one reader doing the scoring, with the rubric above so the scoring can be checked. That is enough to choose a default and a free model, because the gaps at the top were not close and the price gaps were enormous. It is not enough to call a winner in general. Our blind bake-off, where the judges do not know which model wrote what, is how we make that kind of claim, and this chapter will be in its next round. Japanese and Chinese got the same sixteen, scored the same way: the Japanese chapter and the Chinese chapter.

If you want to see the difference for yourself, bring a chapter, try two models on it, and read it in Crossleaf; the first ten chapters are free.

Try Crossleaf. Bring a story in any language and read it in English, chapter by chapter, with the original a tap away. Start reading free →