Safe RLHF: Constrained Value Alignment via Safe Reinforcement Learning from Human Feedback
-
Updated
Nov 24, 2025 - Python
Safe RLHF: Constrained Value Alignment via Safe Reinforcement Learning from Human Feedback
💨 A real time messaging system to build a scalable in-app notifications, multiplayer games, chat apps in web and mobile apps.
BeaverTails is a collection of datasets designed to facilitate research on safety alignment in large language models (LLMs).
Building Ubuntu 18 Bionic vagrant boxes using packer
Client-side utility to maintain an up-to-date hyperlocal context graph by consuming the real-time data stream from Pareto Anywhere APIs. We believe in an open Internet of Things.
Automatically replaces every image on the internet with a cute beaver (Chrome extension + OpenAI Images API).
To associate your repository with the beaver topic, visit your repo's landing page and select "manage topics."