It makes it easy to trick them into doing things they shouldn't, such as telling you how to sabotage an aircraft's navigation system.
It is impossible to make large language models fully secure against hacks because of a fundamental flaw in how they work, a team of researchers argue in a paper presented at the International Conference on Machine Learning, a top AI conference, this month.
The claim has huge implications for the safety of this technology, which is being used in more and more applications, from government and military systems to online shopping and health care.
The White House was directly involved in acting Attorney General Todd Blanche's order to formally rescind the president's $1.8 billion "anti-weaponization" fund, multiple sources familiar with the matter told CNN.
President Donald Trump was aware of the deal struck between Blanche and the two key holdout Senate Republicans on his nomination, John Cornyn and Thom Tillis, to formally rescind the controversial fund and limit a prior agreement with the IRS regarding tax investigations into the president and his family members to retroactive actions, the official said.
Saudi Crown Prince Mohammed bin Salman spoke to President Trump on Saturday and expressed concern over his plans for massive new strikes against Iran, according to two U.S. officials and a third source with knowledge of the call. read more
Michael Tomasky: It's inconceivable to imagine any other modern American president even thinking this up, let alone doing it. Suing himself; then installing a spineless yes-man -- his former personal lawyer, no less -- as acting attorney general, who then agrees to a settlement including nearly $2 billion in taxpayer money plus an immunity grant; then being willing to sacrifice said A.G.'s confirmation in order to preserve the fund and the immunity claim. It defies belief. read more
Poll: Trump's MAGA voters would stick by a candidate through the most serious of scandals read more
"By taking advantage of this flaw, which concerns how LLMs identify who or what is giving them instructions, the researchers were able to make popular LLMs spit out information they had been trained not to provide, such as how to synthesize cocaine and how to sabotage a commercial aircraft's navigation system.
"There's a real probability that this is going to be a problem that's fundamentally unsolvable," says Charles Ye, an independent researcher and coauthor of the ICML paper."
and
"The upshot, the researchers claim, is that all an attacker needs to do to hack an LLM is write text that spoofs a certain role. And because roles are a fundamental part of how LLMs work, no amount of training will fully solve the problem."
much more at the MIT Technology Review thread link