Inserire una breve descrizione delle modifiche fatte
Minor changes are by default collapsed in the page history.
No changes
The page does not exist yet.
Failed to load changes
Version by on
Leave Collaboration
Are you sure you want to leave the realtime collaboration and continue editing alone? The changes you save while editing alone will lead to merge conflicts with the changes auto-saved by the realtime editing session.
Ethically Aligned Reward Models: A Language Model Post-Training Framework
Nicolas Cridlig • Marco Sangiorgi • Andrea Fossa
sommario
Aligning Language Models (LMs) with human moral values is a critical challenge in the development of responsible AI systems. Reward Models play a central role in this process by enabling post-training alignment through Reinforcement Learning from Human Feedback (RLHF) [2]. In this project, we leverage the ETHICS dataset [1] to investigate the extent to which modern Reward Models can represent and generalize ethical concepts. Using open-source tools, we plan to create morally congruent Reward Models aligned with shared human values, intended for integration into RLHF pipelines. Our goal is to develop a prototype framework capable of evaluating and guiding LMs toward ethically preferable outputs.