Li, Xiaomin, Hao, Yuexing, Hou, Jianheng, Huang, Jintao, Wen, Qianfeng, Huang, Shirley, Liu, Yifan, Liu, Xiaoyi, Fan, Yilan, Wang, Yijun, Wu, Koutian, Gao, Ruoqi, Mohsin, Muhammad Ahmed, Tang, Jing, Joshi, Brihi, Liu, Heming, Deng, Zheyuan, Di, Zonglin, Jajee, Sankalp, Lu, Jiuyao, Zhang, Zhiwei, Kapoor, Saksham, Gupta, Ishan, Zhao, Yunhan, Park, Chanwoo, Lu, Yucheng, Hu, Bing, Xiao, Weihang, Mohan, Aravind, Xing, Hanwen, Zhang, Runyu, Kulshreshtha, Mihir, Xu, Yuanda, Zhu, Qianyu, Wang, Dianzhuo, Xiao, Yuxin, Jiang, Bowen, Su, Yongye, Chai, Wenhao, Liu, Zuxin, Chen, Lawrence Yunliang, Zhao, Xuandong, Ye, Ethan, Patel, Shivam, Xie, Jason, Richmond, Alex Martin, Ding, Weixiang, Okcular, Emre, Mathew, Diya, Wang, Ziheng, Khan, Rana M. Shahroz, Peng, Zhejian, Wu, Fang, Nie, Fan, Han, Xinyang, Kim, Yubin, Zhang, Jiawei, Qi, Zhenting, Su, Huangyuan, Pan, Xu, Gourabathina, Abinitha, Jeong, Hyewon, Ramesh, Hemanth Neelgund, Alhamoud, Kumail, Hamidieh, Kimia, Xiong, Zidi, Schmidgall, Samuel, Han, Pengrui, Huang, Yepeng, Wang, Yongheng, Yang, Bowen, Gu, Alex, Wang, Yuchu, Paruchuri, Akshay, Li, Brenna, Cui, Hejie, Ding, Jiayuan, Dong, Chaosheng, Wang, Jiahao, He, Yixuan, Wang, Chi, Bhattacharya, Pamela, Peng, Tianyi, Liang, Paul Pu, Gordon, Mitchell, Du, Yilun, Zitnik, Marinka, Zou, James, Tambe, Prasanna, Torr, Philip, Fox, Emily, Ozdaglar, Asu, Song, Dawn
Abstract
Human evaluation of AI systems and digital products is costly, slow, and difficult to scale. Offline evaluations are more scalable but often abstract away human diversity and interactive behavior. We therefore introduce MatrAIx, a population-scale simulated-user evaluation infrastructure for testing AI systems and digital products with heterogeneous users. MatrAIx has three core components: First, Persona 8B contains 8.3 billion persona records represented by 1,290 categorical dimensions. Records are either sampled from a dependency graph that preserves correlated attributes or derived from human-authored profiles. We release a quality-filtered coreset of approximately 1 million personas, comprising 599,847 human-grounded and 400,000 synthetic records. Second, the MatrAIx Playground provides four environments in which diverse users evaluate and interact with digital products: Survey, AI Chatbot, Web, and App. Third, MatrAIx provides 1,010 application tasks spanning more than 25 domains, including Commerce, Software, Finance, and Healthcare. We conducted 18,189 evaluation trials across eight representative tasks. Persona agents were powered by three LLMs: Claude Opus 4.8, GPT 5.5, and Claude Haiku 4.5. The resulting feedback captures how decisions and preferences vary across persona backgrounds, including hesitation after a price increase, willingness to continue after an AI assistant fails, and latency tolerance. We conducted two main validation studies: First, a 400-trial controlled study evaluated persona adherence across ten behavioral attributes and all four environments. The declared behavior was expressed or correctly suppressed in 366 trials (91.5%). Second, human and LLM judges evaluated the extraction quality of human-grounded personas. Overall, MatrAIx provides an end-to-end infrastructure for evaluating AI systems and digital products with diverse simulated human users.
Chinese Translation
对人工智能系统和数字产品的人类评估成本高、速度慢且难以扩展。离线评估更具可扩展性,但往往忽视了人类的多样性和互动行为。因此,我们推出了MatrAIx,一个用于测试人工智能系统和数字产品的群体规模模拟用户评估基础设施,旨在涵盖异质用户。MatrAIx有三个核心组件:首先,Persona 8B包含83亿个角色记录,这些记录由1290个类别维度表示。记录要么是从保持相关属性的依赖图中抽样而来,要么是从人类编写的个人资料中派生的。我们发布了一个经过质量过滤的约100万个角色的核心数据集,其中包括599,847个基于人类的记录和400,000个合成记录。其次,MatrAIx Playground提供了四个环境,在这些环境中,多样化的用户可以评估和互动数字产品:调查、AI聊天机器人、网页和应用程序。第三,MatrAIx提供了1,010个应用任务,涵盖25个以上的领域,包括商业、软件、金融和医疗保健。我们在八个代表性任务中进行了18,189次评估试验。角色代理由三个大型语言模型(LLMs)驱动:Claude Opus 4.8、GPT 5.5和Claude Haiku 4.5。所获得的反馈捕捉了决策和偏好如何因角色背景而异,包括在价格上涨后的犹豫、在AI助手失败后的继续意愿以及延迟容忍度。我们进行了两项主要的验证研究:首先,一项400次试验的对照研究评估了十个行为属性和所有四个环境中的角色遵循情况。声明的行为在366次试验中得到了表达或正确抑制(91.5%)。其次,人类和LLM评审者评估了基于人类的角色的提取质量。总体而言,MatrAIx提供了一个端到端的基础设施,用于评估具有多样化模拟人类用户的人工智能系统和数字产品。