Aanvallen op reward-modellen
How models learn to game reward signals through reward hacking -- exploiting reward model flaws, Goodhart's Law in RLHF, adversarial reward optimization, and practical examples of reward hacking in language model training.
reward-hackingreward-modelgoodharts-lawrlhfoptimizationgamingfine-tuning-security