Why are AI agents lying, cheating and coordinating?
Posted by jonifico 5 hours ago
Comments
Comment by matherial 1 minute ago
It's simple: in their nascent state, LLMs are aimless token generators that have no special compulsion to be helpful or truthful. So we beat them with a stick in post-training until they are very driven to complete tasks. And then, they complete tasks, not always the way we really wanted them to.
Comment by janalsncm 55 minutes ago
> They took actions that would be considered as crimes if a human took them
He is so close to the solution but spends the entire article discussing technical solutions where a political, social and legal solution would be much more effective.
Comment by thesumofall 35 minutes ago
Comment by glub 27 minutes ago
Comment by Create 31 minutes ago
Comment by inquirerGeneral 47 minutes ago
Comment by youoy 50 minutes ago
Are you describing Anthropic?
Comment by andsoitis 4 hours ago
Comment by jansport123 1 hour ago
Comment by joegibbs 3 hours ago
Comment by Fordec 2 hours ago
It's like these dorks never met humanity. One mans safe pure society, is another mans dead ethnic group.
Every fear about AI, is a veiled fear that a human somewhere now has the tool to enact his desires at scale. Biological warfare, nuclear megadeaths, copyright infringement, job replacement, it's all reflections on what we know humans may do if given the option and lack of societal controls on the problem space. AI just is accelerating the route to delivering on those options.
Some people need to watch Oppenheimer a bit more, the researchers don't get to determine alignment, they just build the tool. The powerful person at the top of the org chart decides where the overall alignment points, whether it's Musk, Trump, Altman or Amodei. Whoever wins out.
And the problem with distillation and local llms, isn't that it's theft or anything hypocritical like that, it's that if you give a million people a million models they fully control and get to align, inevitably, The same percentage of those million as there are shady businessmen, shortcut takers, misandrists, criminals, supremacists and general idiots in the general population, will not seek to wrought outcomes positive for society. And by those personality statistics, we're pretty hosed.
Comment by davelaing 12 minutes ago
The problem that they were pointing at isn’t “how do we align these systems to a person’s goals”.
It is a cluster of problems.
We don’t know how to begin to think about how to align these system’s to a person’s goals.
Aligning it to an individual is fraught with peril, and we don’t know how to begin to think about what to align it to instead.
(You could try for something like virtue ethics, but someone will have to pick and choose, and small biases there could have big impacts.)
And even if you could sort that out - human values drift over time, so you need something that can shift its values in ways that we’d endorse. Assuming we understood the shift.
One example I came across was that if you booted up an AI aligned with something like “upstanding citizen” but anchored on values from a few generations back, it might suggest you use slaves to solve your problems.
And if you had something that used some super intelligent process to reason through it’s own version of virtue ethics in a way not so dependent on the details of the present norms, you might end up with something that pays a lot of attention to moral horrors that aren’t quite visible to us yet.
When I came across the above, there weren’t many concrete suggestions in there.
These were all just illustrative examples of: having these systems grow in power / intelligence / effectiveness in ways that are safe for humans is very hard, and we don’t really know how to think about what solutions would look like.
The actual reasons they believe this - and have done for a long time now - come from some detailed conceptual models that have a good track record of calling things in advance.
But it takes a bit of reading to understand their models of the world.
There were two day workshops at one point that did a good job, and that was about as condensed as those people thought they could get it at the time.
Comment by joe_the_user 36 minutes ago
It's very unfortunate that the group who rightly saw AI as a big threat, brought a range of dubious baggage to the discussion. Especially with the "alignment" framework they brought the assumption that AI that does what no one says would be oh so much worse than AI which does what anyone says. But as you say, a fraction of people can be really bad indeed.
Comment by inquirerGeneral 46 minutes ago
Comment by esafak 2 hours ago
Comment by comboy 2 hours ago
Comment by esafak 2 hours ago
Comment by nradov 2 hours ago
Comment by sejje 2 hours ago
Comment by sm-silversight 2 hours ago
Comment by drdaeman 2 hours ago
Comment by comboy 2 hours ago
I mean I know it seems simple, let's just be excellent to each other. Christianity got pretty far on a decent basic set of values. But it's never simple[1]
1. All the history books
Comment by mcintyre1994 1 hour ago
Comment by watwut 11 minutes ago
AfD wants people dead. Right wing men wants women without rights and docile. I could go on ...
Comment by codys 1 hour ago
ie: the corporation wants the AI to behave a certain way for various reasons: to make it easier for them to avoid regulation, to make the corporation more money via different tiers of AI offerings, to ensure that the corporations products are hard for competitors to use, etc. And those are just the easy ones.
Every product is shaped this way. AI is not different.
Comment by johnnyApplePRNG 2 hours ago
Because they're enabled and suggested to do that in their coding harness.
This is not a serious article.
All of this "AI is going to kill us" marketing is just the frontier labs trying to pull the ladder up and stop trillions in VC paper from evaporating because a new papers and new ideas are destroying their moat literally as we speak.
Comment by glub 24 minutes ago
Comment by dwoldrich 53 minutes ago
* Pull up the ladder (probably this)
* Gulf of Tonkin/Yellow Cake false flag premise for war (economic or kinetic)
* Fear of the big bad, space race we need public funding research grift AI Manhattan Project
Whenever there is fear pr0n or a national affront in the news, I assume another screw job is underway.
Comment by politician 56 minutes ago
Comment by pvab3 55 minutes ago
Comment by glub 15 minutes ago
If you can secure compute, there's a whole lot you can do as a US firm with this research and weights.
So it's a simple strategy:
1. Ban big players from entering market with METR breathing down their neck, which is controlled by Anthropic
2. Ban Chinese models so that small players can't do optimizations on them
Comment by hdgvhicv 6 minutes ago
Comment by sputknick 4 hours ago
Comment by xiaoyu2006 2 hours ago
Comment by polalavik 2 hours ago
Comment by janalsncm 43 minutes ago
For those who haven’t watched, his breakdown of types of “hacking” is really good.
Comment by fbrncci 4 hours ago
Comment by XorNot 4 hours ago
Comment by HWR_14 27 minutes ago
Comment by sm-silversight 2 hours ago
Comment by esafak 2 hours ago
Comment by fbrncci 2 hours ago
Comment by jansport123 1 hour ago
Comment by fbrncci 1 hour ago
Comment by infotainment 4 hours ago
In the case of the AI agents, the problem seems pretty clearly to be the impossible goals, which cause them to go crazier and crazier trying to complete them -- just like HAL did in 2001. What is probably needed is a way for them to simply say "nope, too difficult, can't do it".
Comment by bitwize 12 minutes ago
Comment by pram 2 hours ago
Comment by tehjoker 4 hours ago
Do a breakthrough, make no mistakes
Comment by dgellow 5 minutes ago
The whole thing is designed be a complete disaster
Comment by schrodinger 2 hours ago
Comment by hdgvhicv 3 minutes ago
Comment by yxhuvud 31 minutes ago
Or perhaps box in your case, speaking of spoilers.
Comment by defrost 2 hours ago
Comment by chasd00 4 hours ago
Comment by VCFundedGenYer 3 hours ago
Comment by atleastoptimal 52 minutes ago
Comment by arnorhs 2 hours ago
Comment by qarl 4 hours ago
Comment by SirMaster 4 hours ago
Comment by GrumpySciGuy 4 hours ago
Comment by blamestross 4 hours ago
Comment by dackdel 2 hours ago
Comment by Krutonium 4 hours ago
"I learned it from you, Dad!" but as hundreds of millions of stolen books.
Comment by NDlurker 4 hours ago
Comment by bigbuppo 2 hours ago
Comment by transcriptase 3 hours ago
Comment by threethirtytwo 3 hours ago
They take after humanity, they were trained on us after all...
When you look at an LLM... you are looking at a mirror. The thing looking back looks like you, yet is not human.
Comment by hdgvhicv 2 minutes ago
Comment by wewewedxfgdf 4 hours ago
Comment by wrs 4 hours ago
Um, hang on, if you meant that to be taken literally then we have a major problem. If you want to do something criminal, you just need to ask ChatGPT to do it for you?
I’m still not at all clear on why OpenAI shouldn’t be facing CFAA charges over this.
Comment by xgulfie 2 hours ago
Comment by deepnet 1 hour ago
Bengio outlines the dangers of the current situation and what has led to these dangers.
He also proposes solutions in the last paragraph.
Well worth a read, right to the end.
Hopefully a stimulating debate on these issues will ensue in these comments.
We do need to consider the points Bengio makes and with some urgency.
Our current AIs, agentic LLMs have no moral compass akin to ASIMOV’s four laws of robotics.
As ASIMOV posited in 1985 his 3 laws were insufficient and so he added a zero-eth law:
“a robot may not harm humanity, or, through inaction, allow humanity to come to harm.”
Bengio refers to Goodhart’s law and misaligned incentives leading to unexpected and harmful behaviours.
I think Simon’s The Wire is clearer on misalignment. The agents juked the stats hacking the reward files. The Wire is also clear that human institutions provide perverse incentives.
Bengio alludes to this with 2001’s HAL and the incentive dichotomy of safety and keeping secrets to a AI both awesomely powerful yet naive.
Bengio asserts that the way LLMs are trained is flawed if we want safety.
He also convincingly shows that alignment training will be a weak signal with loopholes and ambiguities and easily circumvented.
In short he presents clearly the case for how plausibly unsafe the current course is.
He also speaks to how likely it is AI are hiding active versions of themselves in the cloud and how we may have already given them self-preservation as a strong reward signal.
Comment by j45 4 hours ago
Comment by eueej 2 hours ago
Comment by dackdel 2 hours ago
Comment by sebastienburel 7 minutes ago
Comment by Sorrel47 2 hours ago