Mark RadarMARK RADAR
EN

Research Exposes LLM ‘Language Tax’ as Claude Uses Up to Three Times More Tokens for Chinese, Japanese and Korean

1 reports · First detected 2026-04-30 · Last active 2026-04-30

Large language model API usage is billed by the token, making tokenization efficiency a direct factor in costs and response times. AI researcher Aran Komatsuzaki said Anthropic’s Claude encodes Chinese, Japanese and Korean text less efficiently, creating a “language tax” that leaves non-English users paying more.

As of July 20, 2026, tests showed that Claude’s token consumption for Chinese, Japanese and Korean content was nearly three times that for English at the upper end. Google Gemini and Chinese models Qwen and DeepSeek were more efficient at processing Chinese. The data did not provide actual U.S. dollar amounts, and final costs still depend on each model’s prevailing API price per million tokens.

All Coverage

1 original reports

The Backstory

The history behind this event

No historical echoes for this signal

Mark Radar|MARK RADAR
All times are in Taipei time (GMT+8)