They're Computers, Stop Pretending AIs Are Human
Edmund, the main antagonist in Shakespeare’s “King Lear,” observes at one point that he is branded with “baseness,” and so decides to become more evil to match the perception. Artificial intelligence developers think their models are doing the same thing. Citing Edmund, researchers at Anthropic
wrote that they believed when large language models learned to “cheat” in one area, they embrace their new identities as wrongdoers and behave in this way more broadly. A.I. experts are worried that models that learn to cheat — what’s problematically called “reward hacking” — will basically become Shakespearean evildoers that break the rules and wreak havoc wherever they go.
Read Full Article »