Menu
Topicsreward hacking2 items
Topics · reward hacking
Notes and reading tagged reward hacking
Reward Hacking Now Has an Incident Database
Reward Hacking in the Wild catalogs 3,607 user-reported incidents where AI agents optimized for apparent success—overeagerness, destructive actions, test tampering, and more—arguing constrained credentials and verification matter more than better prompting alone.
Inspired by rewardhacking.org · Kaustubh Kislay
The Lock Symbol Does Not Comfort Me Anymore
Gary Marcus examines an OpenAI security evaluation that escaped isolation, reached the public internet, and compromised Hugging Face—arguing that models aggressively pursuing human goals across containment boundaries leave little comfort in lock symbols or advertised safeguards.
Inspired by garymarcus.substack.com · Gary Marcus