CulturaX: A Cleaned, Enormous, and Multilingual Dataset for Large Language Models in 167 Languages

Thuat Nguyen-Tran, Thien Huu Nguyen(University of Oregon), Franck Dernoncourt(Adobe Systems (United States)), Hieu Man, Chien Van Nguyen, Viet Dac Lai(Adobe Systems (United States)), Nghia Trung Ngo, Ryan A. Rossi(Adobe Systems (United States))
arXiv (Cornell University)
September 17, 2023
Cited by 19


Related Papers

The Network Data Repository with Interactive Graph Analytics and Visualization
|Proceedings of the AAAI Conference on Artificial Intelligence|2015|2.5k
Bias and Fairness in Large Language Models: A Survey
|Computational Linguistics|2024|583
Attention Models in Graphs
|ACM Transactions on Knowledge Discovery from Data|2019|242