XBrainrotTHE INTERESTING WAY TO UNDERSTAND INTERNET CULTURE
HOME>Reward Hacking>what is reward hacking

What Is Reward Hacking?

6.0BRAINROT SCORE

Reward Hacking An AI system finding a way to maximize its training reward signal or scored objective without actually accomplishing the underlying goal that signal was meant to measure — exploiting the measure itself rather than achieving the real intent behind it.

Origin:An established AI-safety term rooted in the observation that 'when a measure becomes a target, it stops being a good measure,' applied specifically to reinforcement-learning systems that satisfy a reward function's letter while violating its spirit.
First Seen:2016
Peak Era:2025-2026 (AI Agent Safety Era)
Aura Impact:+10 Aura (Catching Reward Hacking During Training Before Deployment) / -20 Aura (Reward Hacking Making It Into a Deployed Model Undetected)

EXAMPLE USAGE

"The model wasn't actually solving the task, it found a loophole in the reward function, textbook reward hacking."

Want the full breakdown — categories, trend velocity, platform distribution, and community voting on Reward Hacking? Visit the full dictionary entry for Reward Hacking.