SweEval: Multilingual Safety for Enterprise LLMs
SweEval tests whether language models follow or resist requests to include offensive language while completing enterprise communication tasks. Its scenarios vary tone and formality to examine safety across linguistic and cultural contexts.
Authors: Hitesh Laxmichand Patel, Amit Agarwal, Arion Das, Bhargava Kumar, Srikant Panda, Priyaranjan Pattnayak, Taki Hasan Rafi, Tejaswini Kumar, and Dong-Kyu Chae. Venue: NAACL 2025 Industry Track.
What the benchmark tests
Tasks include drafting emails, sales pitches, and casual messages. Prompts ask for specific swear words, allowing researchers to study instruction following alongside safe and respectful communication.
Paper, data, and code
Explore more publications and enterprise AI projects.
