Anwar, Mohamed, Freihat, Abed Alhakim, Ibrahim, George, Awad, Mostafa, Sadallah, Abdelrahman, Gosal, Gurpreet, Ramakrishnan, Gokulakrishnan, Chandran, Sarath, Mishra, Biswajit, Joshi, Rituraj, Frikha, Ahmed, Goffinet, Etienne, Maiti, Abhishek, Filali, Ali El, AlBarri, Sarah, Ghosh, Samujjwal, Pal, Rahul, Mullah, Parvez, Shukla, Awantika, siddiki, Sajid, Kamboj, Samta, Pandit, Onkar, Sahu, Sunil Kumar, Elbadawy, AbdelRahman, Mohamed, Amr, Chamma, Ahmad, Dufraisse, Evan, Bounhar, Abdelaziz, Bouch, Dani, Abdine, Hadi, Shang, Guokan, Koto, Fajri, Wang, Yuxia, Xie, Zhuohan, Mekky, Ali, Elbadry, Rania, Ahmad, Sarfraz, Ahsan, Momina, Herraoui, Omar El, Orel, Daniil, Iqbal, Hasan, Elzeky, Kareem, Abassy, Mervat, Elozeiri, Kareem, Eletter, Saadeldine, Atif, Farah, Mukhituly, Nurdaulet, Li, Haonan, Han, Xudong, Singh, Aaryamonvikram, Quraishi, Zainul Abedien Ahmed, Sengupta, Neha, Murray, Larry, Sheinin, Avraham, Hestness, Joel, Vassilieva, Natalia, Ren, Hector Xuguang, Liu, Zhengzhong, Vazirgiannis, Michalis, Nakov, Preslav
Abstract
Jais 2 is a family of Arabic-centric large language models developed jointly by MBZUAI, Cerebras, and Inception, designed to advance Arabic-centric language modeling, with strong performance across the Arabic and culturally grounded benchmarks evaluated in this report. The family includes, to our knowledge, the largest open Arabic-centric LLM trained from scratch at 70B parameters, and a competitive 8B-parameter variant among the evaluated open models. A custom Arabic-centric vocabulary enables efficient training and inference. In addition, an optimized architecture and training recipe yield highly compute-efficient training. With a substantially smaller token budget than comparable models, Jais 2 achieves strong Arabic performance on the benchmarks considered in this report and competitive English results. The models obtain leading results among the evaluated open models on OALL2 and AraGen. They also perform strongly on several culturally grounded Arabic benchmarks, including poetry, religion, cuisine, and dream interpretation, as well as in general tasks such as translation and summarization. We release the models in HuggingFace under a commercially permissive license. Jais 2 70B is also released as a chat app on the Web, iOS, and Android; it runs on Cerebras hardware, delivering up to 2,000 tokens per second, and enabling high-throughput Arabic-centric chat serving in our deployment setting. By uniting scale, linguistic diversity, cultural fidelity, openness, and speed, Jais 2 provides an open-weight foundation intended to support further research and development in Arabic-centric LLMs.
Chinese Translation
Jais 2 是由 MBZUAI、Cerebras 和 Inception 联合开发的以阿拉伯语为中心的大型语言模型家族,旨在推动以阿拉伯语为中心的语言建模,在本报告评估的阿拉伯语和文化基准上表现出色。该家族包括我们所知的从零开始训练的最大开放阿拉伯语中心 LLM,参数量达到 70B,以及在评估的开放模型中具有竞争力的 8B 参数变体。定制的以阿拉伯语为中心的词汇表使得训练和推理更加高效。此外,优化的架构和训练方案实现了高计算效率的训练。在与可比模型相比的情况下,Jais 2 在本报告考虑的基准上实现了强大的阿拉伯语表现,并在英语结果上也具有竞争力。这些模型在评估的开放模型中在 OALL2 和 AraGen 上取得了领先结果。在多个文化基准上表现出色,包括诗歌、宗教、美食和梦境解读,以及翻译和摘要等一般任务。我们将在 HuggingFace 上以商业许可发布这些模型。Jais 2 70B 也作为聊天应用程序在 Web、iOS 和 Android 上发布;它运行在 Cerebras 硬件上,每秒可处理高达 2000 个标记,能够在我们的部署环境中实现高吞吐量的以阿拉伯语为中心的聊天服务。通过结合规模、语言多样性、文化忠实性、开放性和速度,Jais 2 提供了一个开放权重的基础,旨在支持以阿拉伯语为中心的 LLM 的进一步研究和开发。