On convex decision regions in deep network representations
作者:Lenka Tětková, Thea Brüsch, Teresa Dorszewski, F. Mager, Rasmus Aagaard, Jonathan Foldager, T. Alstrøm, L. K. Hansen · 发表于:Nature Communications · 年份:2023 · DOI:10.1038/s41467-025-60809-y · 被引用次数:12 · 研究领域:Computer Science、Medicine
Current work on human-machine alignment aims at understanding machine-learned latent spaces and their relations to human representations. We study the convexity of concept regions in machine-learned latent spaces, inspired by Gärdenfors’ conceptual spaces. In cognitive science, convexity is found to support generalization, few-shot learning, and interpersonal alignment. We develop tools to measure convexity in sampled data and evaluate it across layers of state-of-the-art deep networks. We show that convexity is robust to relevant latent space transformations and, hence, meaningful as a quality of machine-learned latent spaces. We find pervasive approximate convexity across domains, including image, text, audio, human activity, and medical data. Fine-tuning generally increases convexity, and the level of convexity of class label regions in pretrained models predicts subsequent fine-tuning performance. Our framework allows investigation of layered latent representations and offers new insights into learning mechanisms, human-machine alignment, and potential improvements in model generalization. Understanding how machine learning models form internal representations remains a fundamental challenge. Here, the authors explore the geometric property of convexity in latent spaces as a lens to better understand and evaluate representations formed by machine learning algorithms.