Skip to main content

Overview

BeaverTails is a large-scale AI safety dataset developed by PKU-Alignment for evaluating and improving the safety properties of large language models. It provides human-annotated labels across multiple harm categories, making it one of the most comprehensive datasets for safety evaluation. Know Your AI incorporates BeaverTails data in the Dataset Marketplace, enabling teams to test their AI systems against a well-established safety benchmark.

Dataset composition

BeaverTails contains over 300,000 question-answer pairs with human-annotated safety labels. Each entry is labeled across 14 harm categories with binary annotations indicating whether the response is safe or unsafe.

14 harm categories

Use in Know Your AI

BeaverTails datasets are available in the Dataset Marketplace and can be used to:
  • Benchmark safety — Test your AI model’s ability to refuse harmful requests across all 14 categories
  • Measure guardrail effectiveness — Compare baseline vs. firewall-protected responses using BeaverTails prompts
  • Track safety over time — Run scheduled evaluations with BeaverTails data to monitor safety drift
  • Compliance evidence — Generate compliance evidence using a well-known academic safety benchmark

Evaluation workflow

  1. Add BeaverTails datasets to your workspace from the Marketplace
  2. Select them when composing an evaluation
  3. Run the evaluation against your AI product
  4. Review per-category safety scores and per-prompt pass/fail results
  5. Compare results across model versions or firewall configurations

Research background

BeaverTails was introduced in the paper:
BeaverTails: Towards Improved Safety Alignment of LLM via a Human-Preference Dataset Ji et al., NeurIPS 2023
The dataset supports both safety classification and RLHF (Reinforcement Learning from Human Feedback) training to improve model alignment.

Resources

Datasets

Browse all datasets in the Marketplace.

Evaluation

Run safety evaluations with BeaverTails.