🚀NEW LABGetting Started with Claude AgentsStart lab
← All papers  /  Sep 22, 2026
Agents · Retrieval

M-SQE: Multilingual Skill Quality Estimation for Enhancing Language Equality in Agentic Skill Use

First page
M-SQE: Multilingual Skill Quality Estimation for Enhancing Language Equality in Agentic Skill Use
The curator’s take

Yilun Liu, Shimin Tao, Daimeng Wei and colleagues at Huawei audit community agent-skill libraries, find almost no content in low-resource languages, and propose M-SQE, a post-retrieval scorer that picks usable skills from synthesized in-language candidates.

Ask this paper

Key points
01

The audit. Swahili and Hindi have no in-language skill content, so retrieval returns skills written in a different language from the query and task accuracy drops.

02

Why relevance is not enough. Synthesizing in-language skills fills the gap, but their quality varies, so relevance ranking alone often surfaces a related skill that cannot be executed.

03

Two views of quality. M-SQE scores each candidate with a Theory view for intrinsic quality and an Action view for task-grounded utility, then combines them into a domain-conditioned score.

04

Results. Across general, tool-use and cultural tasks and three retrievers, task success beats the average of existing baselines by at least 3.5 points.

05

Largest gains in low-resource languages. Hindi improves by 12.9 points and Swahili by 5.6 points, with strong results across all six culture regions.

Abstract

Agent skills, reusable procedural documents that extend LLM agents beyond their parametric memory, have become an important interface for deploying agents on real-world tasks. Community-maintained skill libraries built around this interface are growing rapidly. However, this ecosystem remains deeply English-centric: our audit finds that low-resource languages such as Swahili and Hindi have no in-language skill content, so retrieval often returns a skill written in a different language than the query, degrading accuracy and recall. A practical solution is to synthesize in-language skills for retrieval but the quality can be unreliable, so relevance in this setting alone often surfaces a related but unusable candidate. To address this, we propose M-SQE, a post-retrieval Multilingual Skill Quality Estimation framework that scores candidates via a Theory view for intrinsic quality and an Action view for task-grounded utility, unified into a domain-conditioned final score. We evaluate M-SQE across three skill-use domains: general, tool-use, and cultural tasks. Empirically, we build three-layer candidate skill pools mirroring today's ecosystem, where M-SQE's task success exceeds existing baseline's average by at least +3.5 points across three different retrievers. Particularly, M-SQE lifts the lowest-resource languages most (+12.9pp on Hindi and +5.6pp on Swahili) and achieves strong performance across all six culture regions, thereby moving agentic skill use toward linguistic and cultural equality.

Every Monday
Get next week’s papers.
Subscribe on Substack