Research Exposes LLM ‘Language Tax’ as Claude Uses Up to Three Times More Tokens for Chinese, Japanese and Korean
Large language model API usage is billed by the token, making tokenization efficiency a direct factor in costs and response times. AI researcher Aran Komatsuzaki said Anthropic’s Claude encodes Chinese, Japanese and Korean text less efficiently, creating a “language tax” that leaves non-English users paying more.
As of July 20, 2026, tests showed that Claude’s token consumption for Chinese, Japanese and Korean content was nearly three times that for English at the upper end. Google Gemini and Chinese models Qwen and DeepSeek were more efficient at processing Chinese. The data did not provide actual U.S. dollar amounts, and final costs still depend on each model’s prevailing API price per million tokens.
All Coverage
1 original reportsThe Backstory
The history behind this eventNo historical echoes for this signal
Subscribe to Mark Radar Weekly
Every Friday, the week's strongest signals in your inbox. Unsubscribe anytime.