AI Governance Institute logo
AI Governance Institute

Intelligence for Compliance and GRC Teams

← All news

RLHF

Reinforcement Learning from Human Feedback (RLHF) is a technique that trains AI models to align their outputs with human preferences by using human evaluations as a reward signal. This approach has become critical for enterprises deploying large language models, as it helps ensure generated content meets safety, accuracy, and ethical standards. From a governance perspective, RLHF raises important questions about evaluation consistency, annotator bias, and the need for transparent documentation of human feedback processes used in model training.

1 item