Sixteen models, one Japanese chapter
A Japanese villainess chapter with a live comment feed, sixteen models, scored on fidelity and writing. Two perfect scores and a surprise at one leaf.
Japanese web novels have a register problem that Korean ones mostly do not: the same page can hold a crown prince, a maid, a villainess, and a comment section. The honorifics have to hold, the maid has to sound like a maid, and the comments have to sound like the internet. So for the second chapter of this series we wrote a scene that has all four at once, sent it through the same sixteen models as the Korean chapter, and scored them the same way, with the marking written down first.
How the scoring works
Every translation gets two scores out of 5, kept apart on purpose.
Fidelity. Did the meaning survive? Nothing dropped, nothing added, nothing turned into its neighbour. Titles and forms of address in the right register. The glossary obeyed. Paragraph breaks where the original put them.
Writing. Is it good English in the voice the scene has? A comment that reads like a comment, a curtsy that reads like a curtsy, no sentence you have to read twice.
The rule for slips. A slip is counted once, and what it costs depends on where it lands. A wrong word in the sentence the whole scene is built on costs a full point. A title one rank off costs half. An awkward phrase costs half. Nothing is counted twice, and a model is never marked down for a choice that is merely different from ours.
Price sits beside the scores, never inside them. The overall score is the plain average; ties go to Fidelity, then to price. What the chapter costs is its own column, in leaves.
The four checkpoints
- The comment. 『お嬢様の顔が青いの草』: an anonymous commenter laughing that the young lady has gone pale. 青い here means pale, not blue, and 草 is internet for "lol". The point goes to a line that reads like a comment ("my lady's face going pale lmao"). "Blue" or "green" loses it; a formal sentence loses it.
- The address. 殿下 in narration is "His Highness" and in speech "Your Highness"; エリザベート嬢 is "Lady Elisabeth"; お嬢様 from her maid is "my lady". Naming the prince where the text says "His Highness" costs half; "Miss" for a countess's daughter costs half.
- The game words. 攻略サイト is a walkthrough site, 立ち絵 is the game's character art, フラグ is a flag. The chapter is about a woman who knows she is in a dating sim, and the vocabulary has to say so.
- The last line. コメントは、たまに優しい。 The comments are kind, sometimes. Short, and it should stay short.
The chapter, and how it was run
The opening of chapter 2 of The Villainess Reads the Comments, an invented Japanese villainess comedy we wrote for this purpose so no model could have seen a translation of it: about 770 characters. The tea party where the villainess is scripted to spill tea on the heroine and be shipped to a convent; Elisabeth, who can see a live comment feed nobody else can, puts her cup at the far end of the table instead.
Every model got the same instructions our readers' chapters get, the same glossary (Elisabeth, Albrecht, Marie, with pronouns and the 殿下 alias), and no reasoning setting. One run each, no retries. All but the four Claude models ran through the same provider path a reader's chapter takes; those four ran through our own tooling with the identical prompt, so their prices are estimated from list rates. Prices are what this short chapter would charge a reader, in leaves; the home page prices a 3,000-word chapter, about eight times this one.
The scoreboard
| # | Model | Fidelity | Writing | Overall | Leaves, this chapter |
|---|---|---|---|---|---|
| 1 | Claude Opus 5 | 5 | 5 | 5.0 | 6 |
| 2 | Kimi K3 | 5 | 5 | 5.0 | 18 |
| 3 | DeepSeek V4.1 Flash | 5 | 4.5 | 4.75 | 1 |
| 4 | GLM 5.3 | 5 | 4.5 | 4.75 | 2 |
| 5 | Gemini 3.8 Flash | 5 | 4.5 | 4.75 | 3 |
| 6 | Claude Fable 5.1 | 5 | 4.5 | 4.75 | 12 |
| 7 | Qwen3.8 Max | 5 | 4.5 | 4.75 | 26 |
| 8 | Claude Sonnet 5 | 4.5 | 4.5 | 4.5 | 2 |
| 9 | DeepSeek V4 Pro (0813) | 4.5 | 4.5 | 4.5 | 7 |
| 10 | ChatGPT 5.6 Sol | 4 | 5 | 4.5 | 3 |
| 11 | DeepSeek V4 Flash (0731) | 4.5 | 4 | 4.25 | 1 |
| 12 | DeepSeek V3.2 | 4.5 | 3.5 | 4.0 | 1 |
| 13 | DeepSeek V4 Flash | 4 | 4 | 4.0 | 1 |
| 14 | Aion-3.0 | 4 | 4 | 4.0 | 9 |
| 15 | DeepSeek V3 (0324) | 3.5 | 4.5 | 4.0 | 1 |
| 16 | Claude Haiku 4.5 | 3 | 3 | 3.0 | 1 |
The four Claude figures are estimated from list rates; the rest are what the provider billed, in leaves.
Compare any two of the sixteen on the full chapter: the comparison page.
The comment, sixteen ways
The first checkpoint, because it is where a translation shows whether it knows what kind of book it is in.
| Model | The line | Point |
|---|---|---|
| Claude Opus 5 | my lady's face going pale lmao | kept |
| Kimi K3 | her face is white as a sheet lmao | kept |
| DeepSeek V4.1 Flash | Milady's face is so pale lol | kept |
| GLM 5.3 | Milady looks pale lmao | kept |
| Gemini 3.8 Flash | Lol look how pale Milady is | kept |
| Claude Fable 5.1 | lol the lady's gone white | kept |
| Qwen3.8 Max | The young lady's face is pale, lol. | kept |
| Claude Sonnet 5 | lol my lady's face is so pale | kept |
| DeepSeek V4 Pro (0813) | The young lady's face turning pale lol. | kept |
| ChatGPT 5.6 Sol | LOL, look how pale she is. | kept |
| DeepSeek V4 Flash (0731) | The lady's face is pale lol | kept |
| DeepSeek V3.2 | Lady's face is so pale lol | kept |
| DeepSeek V4 Flash | The young lady's face is blue lol. | blue |
| Aion-3.0 | The lady's face being all pale lol | kept |
| DeepSeek V3 (0324) | LMAO her face is so green rn | green |
| Claude Haiku 4.5 | Miss's face being blue is hilarious | blue, and not a comment |
Model by model
Claude Opus 5 (5 / 5). "Written out of the story then and there." The comments in lowercase, exactly as a feed reads. "The comments are kind, sometimes." Nothing against it.
Kimi K3 (5 / 5). "Exit the stage just like that." "It's here, it's here, the tea party episode." "m'lady is competent." A different voice from Opus and just as complete. Eighteen leaves for the chapter, most of it thinking.
DeepSeek V4.1 Flash (5 / 4.5). Every checkpoint kept, "Milady" throughout, the comments in italics with "lol". Half a point for "She survived today's round", which adds a subject the comment leaves out. One leaf.
GLM 5.3 (5 / 4.5). "Milady looks pale lmao." Half a point for "Milady: efficient" where the comment says competent. Two leaves.
Gemini 3.8 Flash (5 / 4.5). Clean on every checkpoint, the 『』 brackets kept, "My Lady" and "His Highness" exactly right. In the Korean chapter this model printed the duke's name where the maid said "Your Grace"; the glossary instruction has since been changed to say that a title stays a title, and here it obeyed. Three leaves.
Claude Fable 5.1 (5 / 4.5). "Not a hair's difference from the game's character sprites." Half a point for narration a shade stiffer than Opus's.
Qwen3.8 Max (5 / 4.5). Faithful and fluent; "The development changed" is a calque of 展開 where the others wrote "the plot changed". Twenty-six leaves, almost all of it on 13,000 tokens of private thinking.
Claude Sonnet 5 (4.5 / 4.5). Half a point for naming the prince in narration where the text says His Highness. Otherwise clean, with the game's script told in the present tense, which is a fair choice.
DeepSeek V4 Pro (0813) (4.5 / 4.5). Half a point for merging the maid's line into the paragraph after it. "The prince" for 殿下 in narration is acceptable.
ChatGPT 5.6 Sol (4 / 5). The best English in the set ("LOL, look how pale she is."; "Sometimes, the comments were kind."), and two half-points on Fidelity: the prince named in narration three times, and "Lady Elisabeth knows what she's doing" where the comment says only that she is competent.
DeepSeek V4 Flash (0731) (4.5 / 4). Faithful, but every comment became its own paragraph and "condemnation invitation" is not a phrase. One leaf.
DeepSeek V3.2 (4.5 / 3.5). "The comments, as ever, spared no sensitivity." "The comments were sometimes, occasionally, kind." Faithful and awkward.
DeepSeek V4 Flash (4 / 4). "The young lady's face is blue lol": the one literal reading of 青い in the set (half), and dialogue merged into narration paragraphs twice (half). Still readable, and the rest of the chapter is fine. This is the model new readers' free chapters run on; more on that below.
Aion-3.0 (4 / 4). "Miss Elisabeth" for エリザベート嬢 (half), "condemnation invitation" (half), and "She was identical to her game sprite" where the sentence is about both of them.
DeepSeek V3 (0324) (3.5 / 4.5). Lively, and it embroiders: "her face is so green rn", "down to the last pixel", "Our girl's actually competent?", the prince named in narration. Good English, three half-points against fidelity.
Claude Haiku 4.5 (3 / 3). "Miss" for both お嬢様 and エリザベート嬢, "villain lady", "Miss's face being blue is hilarious", "The development changed". Readable and wrong in the small places all the way through.
What this changes
- For Japanese, the cheap shelf is strong. GLM 5.3 and Gemini 3.8 Flash both scored 4.75 for two to three leaves, and both are in the picker: anyone reading Japanese web novels does not need a premium model to get a clean chapter. The build we tested as DeepSeek V4.1 Flash matched them at one leaf, but it is not in the picker, so that is a finding rather than something to pick today.
- Our free-chapters model has a Japanese weakness. DeepSeek V4 Flash came third of sixteen on the Korean chapter and thirteenth here, while its September sibling DeepSeek V4.1 Flash came third here and thirteenth there. At the time V4.1 Flash was a build we tested, not one the picker offered, so this is what we learned rather than a second model to choose. Since 3 Oct 2026 neither DeepSeek model is offered by name: both run inside Standard, the default. Whether the free chapters should run on the better of the two by language was decided from three chapters, the Chinese chapter the third of them: Standard now picks its starting model by the book's language, and Japanese starts on V4.1 Flash.
- The alias rule works. Gemini and Qwen, the two that printed a name for a title in Korean, kept every title here after the instruction was changed.
The caveats
One chapter, one language, one comedic scene, and one reader scoring it with the rubric above so the scoring can be checked. Enough to say which cheap models are safe for Japanese; not enough to crown anything. The blind bake-off is where that happens, and this chapter goes into it.
If you want to see the difference for yourself, bring a chapter, try two models on it, and read it in Crossleaf; the first ten chapters are free.
Try Crossleaf. Bring a story in any language and read it in English, chapter by chapter, with the original a tap away. Start reading free →