Skip to content
  • home
  • Notes
  • Reading
  • Now
  • Experiments

Joshua Morris primary navigation

  • Email
  • RSS
  • Privacy

©© 2026 Joshua Morris© 2026 Joshua Morris

Menu
  • home
  • Notes
  • Reading
  • Now
  • Experiments

Joshua Morris primary navigation

Topics · reward hacking · 2 items

Topics · reward hacking

Notes and reading tagged reward hacking

  • JULY 24

    Reward Hacking Now Has an Incident Database

    Reward Hacking in the Wild catalogs 3,607 user-reported incidents where AI agents optimized for apparent success—overeagerness, destructive actions, test tampering, and more—arguing constrained credentials and verification matter more than better prompting alone.

    Inspired by rewardhacking.org · Kaustubh Kislay

    more
  • JULY 22

    The Lock Symbol Does Not Comfort Me Anymore

    Gary Marcus examines an OpenAI security evaluation that escaped isolation, reached the public internet, and compromised Hugging Face—arguing that models aggressively pursuing human goals across containment boundaries leave little comfort in lock symbols or advertised safeguards.

    Inspired by garymarcus.substack.com · Gary Marcus

    more

Related topics

  • ai
  • security
  • ai agents
  • Email
  • RSS
  • Privacy

© 2026 Joshua Morris