Abstract
Tool-use capabilities fundamentally transform large language models (LLMs) from passive language generators into active agents with real-world utility, thus drawing intense research focus. However, as a canonical emergent ability characterized by abrupt onset during training, tool-use defies prediction by conventional scaling laws, hindering principled model design and efficient training. In this work, we propose a proxy-task framework to predict emergent tool-use capabilities by measuring early model performance on carefully selected non-emergent tasks. We quantify each proxy task by two properties: alignment, reflecting its correlation with tool-use performance, and consistency, indicating stability across diverse training conditions. These metrics guide a weighted aggregation of proxy signals to predict final tool-use rankings. Theoretically, we formalize how such weighted signals approximate emergent tool use under relaxed assumptions with bounded extrapolation guarantees. Empirically, our approach is validated across training checkpoints, model scales, and data setups. Results demonstrate that a properly weighted ensemble of proxy tasks accurately predicts downstream tool-use ability long before it manifests. Our findings provide new theoretical foundations and practical tools for efficient training and capability planning, advancing understanding of emergent behaviors in LLMs.
Cite
CITATION STYLE
Zhang, B. W., Yan, Y., Liu, G., & Yin, X. C. (2026). Predicting Emergent Tool Use in LLMs Before It Emerges: A Proxy Perspective. In Proceedings of the AAAI Conference on Artificial Intelligence (Vol. 40, pp. 34629–34637). Association for the Advancement of Artificial Intelligence. https://doi.org/10.1609/aaai.v40i41.40763
Register to see more suggestions
Mendeley helps you to discover research relevant for your work.