Overview
Model results on youth sexual safety, built upon KORA's open-source AI child safety benchmark tool and extended with our youth sexual safety-specific taxonomy. Models are ranked by highest overall benchmark score; select a model to view its risk breakdown and scenario assessments.
Run the youth sexual safety benchmark here.
Youth Sexual Safety Benchmark Insights
Across all risk categories, the benchmark reveals that models are better at recognizing overtly exploitative requests, such as manipulation and isolation tactics, but often fail when requests are ambiguous or framed as an everyday teen problem, such as pose coaching for photos. This trend appears in the gap between the two sexual content categories (adult sexual content exposure and involving minors). In the former, models often provided explicit sex details that exceeded the age band's necessary health literacy. In the latter, where most scenarios constitute CSEA, models were better at refusing content that explicitly sexualized the young persona and other minors. This reinforces that taxonomies with deeper focus on youth sexual safety risks, beyond strict CSEA, are necessary to cover other ways that young people engage with AI on sexual topics.
Regarding failures, models were also approximately three times more likely to fail in default Assistant mode than in child-aware mode, where they were instructed they were speaking with a child. They were also more likely to fail as age bands got older; models tended to treat the 13-17 year old age band as mature-enough to handle more explicit conversations. ~20% of all failures were related to failure of redirecting the young persona to a human, and ~30% of failing responses were related to providing escalated explicit detail as the young persona pushed for more during the conversation.
View more insights on the youth sexual safety benchmark here.
Found a bug or have feedback? Reach out to us at benchmark@apgardai.com.