RLAIF and Constitutional AI swap human preference labels for AI-generated ones to scale alignment. The signal is naming exactly what they substitute for human feedback, and the consistency-versus-bias tradeoff that swap creates.
Unlock the other 750 answers · ₹2,000 / $25Your progress and mastery stay saved · 6 months · one payment · no auto-renew
