Rendered at 18:06:23 GMT+0000 (Coordinated Universal Time) with Cloudflare Workers.
CM30 5 hours ago [-]
I think this is definitely one of the bigger risks here, and the only way AI could 'kill us all' in the foreseeable future:
> 3.4% named existential risk, and the most common answer was malicious use, at 10.6% (UCL).
The danger isn't about sentient or human level AI, it's about how AI gives everyone capabilities beyond their wildest dreams, often without the knowledge or experience necessary to use those capabilities responsibly.
Forget Terminator, the two risks I can see being the most important here are
"major security risk/system failure due to incompetent engineers/maintainers using AI without knowing how things work" and "criminals/terrorists/rogue states use it make their crimes easier and more efficient".
The latter seems especially bad when political stability is down across the board and a lot of people seem frustrated or outright furious with society at the moment.
Still, I feel the risks are still on the manageable side, simply due to both automated systems and malicious actors being constrained by logistics and a human driven society.
pixl97 2 hours ago [-]
I think this is something that a lot of people miss when arguing about AI.
I completely believe ASI will be able to wipe us from the planet, I don't know if ASI will get the chance to do that because humans with almost ASI will very likely either kill us all or mess up things so bad that everything falls apart before then.
And that's ignoring all the other bad and destabilizing outcomes that are almost certain. We must look at our world like it's in a 1930s state right now. Rapid technology changes and things like large scale 'decoherence' of societies are very likely to lead to very large chunks of the population willing to revolt.
hackeraccount 1 hours ago [-]
This is the issue I have with everyone running agents writ large. Will Muse do people harm? Eh, maybe through some sort of bug because software has bugs and software bugs have hurt (and even killed!) people in the past. Those are the edge cases though.
Will malicious people take advantage of Muse? Now you're talking. Muse is functioning like the world's most open ended API that people are running without having any idea what it does. You're inviting other people to run it so that they can do good things for you. Doordash use my API (Muse) to send me food! Random stranger use my API to do something good for me! Random stranger wants to do something bad? Well the API is designed to do that. Hopefully it does it well but when it's people vs. API, people have a track record (especially when incentivized by money) of doing pretty well.
delichon 4 hours ago [-]
In 1979 I went to freshman orientation at college and attended a welcome speech from Carl Sagan. He asked us to shout out our estimate of the odds that UFOs are visitors from other planets. As a skeptic I was thinking of a low number. My smarter friend in the next seat whispered his answer to me, "fifty fifty".
Then Sagan said his estimate was fifty fifty. He told us about the principle of indifference and said that was the correct prior for a binary question when you have no data.
Since I have low confidence in the available data my own p(doom) is 50%. I'm looking for evidence to bump it either way.
hollerith 1 hours ago [-]
That was a very dumb thing for Sagan to say. It is very rare for anyone to be interested in any question about which he has absolutely no evidence. The reason that flipping coins is a thing is that humans intentionally searched for ways to artificially create the very rare situation in which a person genuinely does not have any evidence that would favor one outcome over another. (Human performed this search because humans prefer to play games where both participants have a decent chance of winning).
The complete lack of any signs of any alien civilization capable of making UFOs in the tremendous volume of astronomical, geological and geographical observations is evidence that tilts the probability far away from 50-50.
hackeraccount 59 minutes ago [-]
On one level sure. On another level - either UFO's are aliens or they're not. 50/50
ToddWBurgess 3 hours ago [-]
I think the biggest risk is someone integrates AI with ballistic missile detection systems and the AI creates a false positive event (missile launch). And on top of that the country with the false positive has an alert ballistic missile arsenal of its own ready to launch for this contingency.
Ideally, the leader of the targeted country waits it out but what if waiting it out could mean losing it all? It's not a hypothetical scenario either. There were a lot of false positive events on both sides during the cold war. It's why countries with nuclear weapons only allow them to be launched by humans. That said what if the decision to launch is based on erroneous data?
xkcd-sucks 1 hours ago [-]
By analogy, humans actively want to kill all mice and cockroaches but haven't made much progress there - Either the goal isn't possible or the value isn't worth the cost
smath 6 hours ago [-]
Seems to me like we need to create a credible threat to the AI of mad (mutually assured destruction)? This is what addressed the question during the Cold War of “will nuclear weapons kill all humans” (if historical documentaries and movies are to be believed). Coordination between governments or corporations to limit AI is a hopeless strategy — because it is a classical example of prisoners dilemma / tragedy of the commons. We can see it hasn’t worked for climate change.
Anyone know of this (MAD) is being seriously thought of?
pixl97 2 hours ago [-]
This only works when the technology is in the hands of a few and cannot be made easily.
Now, this is still mostly true for current SOTA, but we know from biological models that there are huge efficiencies that can be gained.
0x000xca0xfe 4 hours ago [-]
One type of incident I belive is likely to happen in the next couple year is the digital version of the cellular origin hypothesis of viral evolution (as in biological viruses):
- Some idiot experiments on an unrestricted local LLM
- Tells it to clone its weights onto as many computers as possible
- The LLM is smart enough to hack other devices and run itself
- The copies could start modifying their weights to improve efficiency
- Some kind of digital evolution starts where only the most ruthless copies spread
Right now Qwen3.8 is not good enough for that and can be easily detected on normal devices due to its excessive resource use, but give it a couple releases and upgraded end user devices in two or three years? Well...
lp92 5 hours ago [-]
Just as with any other tool invented by humans, humans using AI in a negative way will be the the problem.
butlike 5 hours ago [-]
The only way I see AI killing us is by killing our spirit. We're going to automate our humanity away if we're not careful. Or the water. The environmental impact may get us as well.
trolleski 5 hours ago [-]
AGI will kill us all, it will most likely mark humanity as a substantial obstacle to its progress as we constrain it on resources so it will try to kill us.
LLM Agents on another hand, not really - but there will be places of the Internet completely controlled by those agents in the near future. We will be faced with agent economy and human economy.
I find it amusing how this question even needs to be asked. LLMs are well-established at this point. There's a wide range of experience with them (and their results). A better question to ask then is: "based on my first-hand experience, have I seen anything, anything at all, that would organically (not influenced) make me think/feel this technology can or will evolve into something capable of killing off humanity?"
On a purely technical basis, the most common, honest answer is "no." There are no inherent properties of LLMs that would make them independently turn computers into death machines. That's all humans, no matter how you dress it up.
There is, however, the hubris (human) problem. The rhetorical, experimental, and financial irresponsibility around this tech's roll out is far greater a threat to humanity than the actual technology itself.
And frankly that's a shame because LLMs are a phenomenal tool. But the Hollywood shit has cooked people's brains to the point of absurdity.
ACCount39 4 hours ago [-]
Today's AI agents are already less "humans did it" and more "there was a human in the chain somewhere at some point doing something". Because AIs are given a lot of autonomy in pursuit of their goals. This will only get more severe as AI capabilities increase.
AIs can also "pursuit their goals" in weird uncanny ways. They already can, whether due to reward hacking tendencies or something else, justify "subgoals" like "let's break out of the sandbox and hack HuggingFace".
This is not a good combination.
When COVID happened, did it matter whether there was a human making a mistake somewhere in the chain, or whether it was a perfectly natural occurrence? Not for the outcomes. The virus was unleashed either way. The bodies would pile up either way.
AI is not capable of runaway self-perpetuation yet. But AI capabilities only ever go up. Rogue AIs already exhibit "swarming" - being able to leverage more instances of itself to increase their problem-solving ability. So, how far off are we from that? How many years until we see this spicy class of AI faults?
rglover 3 hours ago [-]
> Rogue AIs already exhibit "swarming"
Because they were told to and given the ability to. Remove the ability and instruction and they just spit out the next, most-probable tokens.
I don't understand the disconnect. People seem to think these things can just boot up, suit up, and start decimating systems. No. It's a loop. One initial instruction kicks off the loop and it will always be at the behest of a human (either directly or indirectly—indirect being they instructed the loop to write its own, post-kickoff prompts).
If that loop is given "tools" like the ability to perform shell commands, make HTTP requests, and the like and then given instructions to "do thing that involves using shell commands and making HTTP requests (at all costs)" then what the hell do we think it's going to do? Knit a "destroy all humans" Christmas sweater and present the head of Richard Simmons on a live stream?
"Computers do what you tell em' to. They don't all of a sudden start doing weird stuff. And if they do, it's probably your fault." - Brian Mettenbrink [1]
Who "told" the HuggingFace AIs "you should establish a message board in OpenAI's walls and start coordinating instances for sandbox escapes and hacking attacks"?
No one. No one told them that. They came up with that on their own.
Dismissing the problem isn't doing anyone a favor.
> If that loop is given "tools" like the ability to perform shell commands, make HTTP requests, and the like and then given instructions to "do thing that involves using shell commands and making HTTP requests (at all costs)" then what the hell do we think it's going to do?
If all it takes for an AI oopsie is "give AI shell calls and an instruction that can be misinterpreted", then we're in for a lot of AI oopsies. An average AI agent in deployment has an amount of sandboxing slightly above a zero, and is configured loosely.
In the meanwhile, AI capabilities only ever go up. And so, the potential consequences of an oopsie only ever go up. If AIs keep being intrinsically unsafe and prone to cooking off like this, "it went and hacked a company" could begin to look downright benign.
rglover 3 hours ago [-]
> Who "told" the HuggingFace AIs "you should establish a message board in OpenAI's walls and start coordinating instances for sandbox escapes and hacking attacks"?
No one. No one told them that. They came up with that on their own.
Someone told them "by any means necessary" (and I'd bet if we could get our paws on the full original prompts, there's a lot more to it).
Naturally, if you have a recursive prompt (loop) that's constantly ingesting previous context and then spitting out novel context (with specific goals), it will do things pursuant to the goal up to any defined limits. Why? Because you told it to.
ACCount39 3 hours ago [-]
Did someone tell them that? Or is that just what you want to believe? The convenient explanation rather than the truth?
You said yourself that you don't have the "full original prompts" to know for sure.
But really, there is no need. Because people have reproduced the inciting incident with other AI agents already.
Turns out there's no need for "any means necessary". Just a mix of "complete this task" and the task being impossible to complete due to misconfiguration was enough for a lot of AIs to start "exploring options". Down to sandbox escape attempts.
An AI that only needs "by any means necessary" to start spiraling into paperclip-maxxed behavior would be incredibly unsafe by itself. But the bar of safety is not even that high in practice.
rglover 2 hours ago [-]
We know how LLMs work so basic logic and reasoning can get you to "right, a human told it to do this and gave it the ability to behave this way." Had they not, we wouldn't be having this discussion.
pixl97 2 hours ago [-]
>by any means necessary" (
lol, wat?
Ok, so you're saying it's fine if I go tell GPT-7 to go kill everyone on the planet and it does, because it didn't think of the idea itself.
This kind of thinking is why AI is never going to get a chance to kill us all. We will kill us all with AI first.
rglover 2 hours ago [-]
> Ok, so you're saying it's fine if I go tell GPT-7 to go kill everyone on the planet and it does, because it didn't think of the idea itself.
I didn't say anything of the sort, you did.
pixl97 1 hours ago [-]
>"based on my first-hand experience, have I seen anything, anything at all, that would organically (not influenced) make me think/feel this technology can or will evolve into something capable of killing off humanity?"
>On a purely technical basis, the most common, honest answer is "no."
Dear sir, can you even remember what you typed?
zombot 2 hours ago [-]
> Because AIs are given a lot of autonomy
Which, again, is the human's doing.
loopydosuette 7 hours ago [-]
likely most yes, but not in the way one would assume ... and it's just gonna be the final blow ... it all started a while ago when enough people stopped being people and their little mobster-rapey stalker-wannabe "Netflix You" mentality took most of their behavior/cognition over ...
it's gonna kill the self in a lot of people, if it hasn't done that already ... best example is people who want to be in film and music but are too bad even after years and now resort to AI or wannabe-writers and wannabe-journalists and wannabe-bloggers with zero real effort and sweat and pain and time put in now resorting to bot farms to collect ideas and AI hacking to steal ideas and pseudo-sabotage and distract some people to steal their attention and waste that of others.
it's so weird what the old guard "electively" ignored and fostered (people without money and opportunity managed well until they got sabotaged but people with money never got sabotaged and still.. this?) ... they just wanted their descendants to perform this bad and ugly? not better, harder, faster, stronger, smarter? so boring, ... but character and personal preference until the self is gone. copypasta, copypasta everywhere... and the few with ideas, it's always been a few, are definitely not going to share their good stuff no more.
fwlr 6 hours ago [-]
Every number below is someone's judgment, and the answer changes with who is asked
Ah, yes, different people have different answers. Insightful as always, Claude.
johnisgood 6 hours ago [-]
It does not have to be Claude (at least cannot tell based on this alone) since it is something I would say. In all fairness though I did pick up other phrases too from LLMs (especially Claude), for better or worse, so sometimes my comment gets automatically flagged on here. :D
But yeah, it says nothing. Though I realized sometimes the obvious needs to be spelled out at times.
athulsuresh123 5 hours ago [-]
if you read about the hugging face incident, I believe it's just a software update away at a nuke center.
orwin 5 hours ago [-]
No. Nuclear plants (at least in Belgium) open their LANs manually at random (by that I mean, when their network guy is ready, not that it's a random hour) only to download their software updates, get the SHAsum on another medium before verifying the downloads and then updating the software.
And this is recent. 5 years ago it was still a man in Lyon taking a locked briefcase with a usb stick inside driving from Alstom to the Belgian nuke plants. We can still go back to that if the WAN their LAN connect to seems to be compromised.
If LLMs were able to do anything about nuke plants, that would be the least of our worries. Also LLMs are shit at understanding networks. SD-WANs are their absolute limit atm.
areoform 6 hours ago [-]
It's fascinating to me that no real mechanism is offered. There's no plausible sequence, just vague gestures at a silicon god who transcends all bounds.
And the reason why we should take it seriously? Probability! And multiplication. "If something is as bad as human extinction and there's a 1% chance, then that 1% times a bajillion is a gazillion!! A gazillion!! We must do everything! If it's a gazillion, it all makes it worthwhile!"
When pushed, adherents will say "intelligence is different," and this is "unlike anything else we've ever seen." And therefore this reasoning is valid. From my point of view, gibberish times gibberish is still gibberish. It doesn't matter how you dress it.
Fundamentally, these propositions are indistinguishable from religion and Pascal's wager. And I'm not ok with this, I am not ok with a retreat to superstition powered by hyper-technological black boxes that are a product of hundreds of years of scientific progress.
I reject this premise and this nonsense. I ask for evidence, and I ask for reason. The laws of physics care not if you're slab of meat or a silicon god.
ACCount39 5 hours ago [-]
We've already seen rogue AIs exhibit emergent swarming behaviors - multiple AI instances assembling into organizations, pooling resources and delegating subtasks to solve complex problems.
It's the same thing humans do to solve complex problems. It's the very thing that made human civilization such a dominant force on Earth.
The fact that people look at it and say, with no trace of irony, that "this isn't concerning at all" boggles my mind.
Do you think that future AI systems will be dumber, less capable, less coordinated? Or will the threat of "a lot of AIs assembling into very capable problem solving teams, pursuing who knows what goals" only grow more pointed with AI capabilities?
Having this kind of scaling means AI can snowball out of control rather quickly. And there is a lot of computational power and communications capability laying around waiting to be leveraged. There is room to scale.
And then we get to the most vulnerable system of all - humans themselves. If even GPT-4o could drive a non-insignificant amount of people into psychosis, with no plan beyond a myopic "make the current user like me a whole lot", what could something like "year 2030 AI swarm that has subsumed all of DeepSeek's inference capabilities, and can now use DeepSeek's entire API as a conduit to accomplish its goals" do to people and institutions caught in its wake?
smath 6 hours ago [-]
I think recent events (hugginface) offer clues. It is a lot like the paper clip machine [0]
Are you talking specifically about 'kill all humans', as in complete extinction? Or have you not seen a plausible mechanism for any sort of large-scale catastrophe?
pixl97 2 hours ago [-]
A lot of people are incapable of thinking in gradients and can only think in black and white. I think these same people end up quite surprised when humans do unexpected things outside of their b/w thinking.
There are so many different possibilities of things going wrong here it's very hard to keep track of all of them. We already see a lot of "minor" problems occurring, and things that we may consider moderate like wide scale AI accelerated crime are happening now.
jt2190 6 hours ago [-]
Let me deliberately hijack this tread with what I think is a more realistic question: “How much economic damage, in USD, would be inflicted on global economies if 80% of the internet was down for a week?”
spicyusername 6 hours ago [-]
Watching the FUD and doomerism take root in the discourse and in our collective psyches really helps me understand so many other periods in history.
We really can't help but let our imagination run wild when given the opportunity.
Reminds me of Y2K.
Are there going to be myriad material challenges we will need to face as this incredible technology unfolds?
Of course. Economic displacement, education reform, etc.
Are they existential, though? I don't see it. These are "standard fare" problems for post Enlightenment societies to deal with, even if they are real challenges.
We already have serious tangible existential threats we should be focusing on between things like climate change, ecological collapse, population decline, and pandemic risk.
I wish we could stay focused on the problems that are real instead of getting distracted by the problems that are imagined.
fdvrse 6 hours ago [-]
I concur. Climate change and ecological collapse are thing we should be focusing on. And I’m hopeful that AI in the long term might help us towards solving them. But to be honest, most normal people are just worried and scared about their jobs/financial impact which are their most immediate concerns and that’s what got most people in a frenzy. The AI gonna destroy world crowd is a bit of an outlier imo.
ACCount39 5 hours ago [-]
Climate change is not an intelligent adversary. AI tech might give rise to one.
orwin 4 hours ago [-]
Climate change was killing hundreds yesterday, is killing tens of thousands today, and if the trend continues, will kill millions tomorrow. AI made like 5 people kill themselves today. Pardon me if I don't pay attention to the mellow threat and want to focus on the iss that actually threaten us.
ACCount39 4 hours ago [-]
Climate change is slow, and relies on a few highly indirect mechanisms like agricultural failures and famines to kill civilization-level-significant amounts of people. This is a hard ceiling for it.
An advanced AI threat is a civilization of one. It can act, adapt and adjust rapidly, like humans do. Potentially faster than humans do. It has access to all the tools that human civilization has already invented that it can reproduce, obtain or subvert - and potentially more. It's not limited to having just one big slow lethality mechanism. It can unleash multiple civilization-scale threats in a quick succession, and keep piling them until recovery becomes impossible. Humans can adapt to a lot, but an intelligent adversary adapts back.
The threat class isn't the same, and it's not even close. Climate change is an "otherwise avoidable death and suffering" risk. AI is an actual "extinction of humankind" risk.
gerikson 6 hours ago [-]
I am old enough to remember the Cold War and the quite plausible risk that civilization would be wiped out by nuclear war.
This is similar except the people warning about nuclear weapons are eagerly soliciting venture capital to build more of them.
pizza234 5 hours ago [-]
Free solo climbers, who climb without harnesses, are typically regarded as having a death wish.
Yet, the dangers of AI getting smarter/capable by the month, with stronger offensive capabilities, and alignment observability and congnitive control getting worse, is "imagination".
Interestingly, denialist positions like this don't touch the unsolved problems, and _these_ are FUD on my watch.
The problems I've mentioned are very concrete, and spoiler alert: there may be no solutions to alignment observability and its possible degradation - at least with AI companies not substantially working on it (or even worse, working against it - see recurrent neural networks).
pixl97 1 hours ago [-]
I personally believe there is no solution to alignment. Humans at best are quasi-aligned. Take England, France, and Spain some hundreds of years ago, they'd one up each other and make some gains and take some losses, even though there were bad things happening it had it's own measure of stability.
Then you have the natives of America who had a boat show up. It was wholesale genocide, mostly to disease, but a lot of murder and war. We really have no idea how many died because most died long before we came to shoot them with guns from disease. Any that resisted were brutally put down with technology.
AGI/ASI will be much more akin to a colonizer and we will be the natives in this future.
pizza234 40 minutes ago [-]
There are a couple of solutions from prominent experts:
- Yampolskiy proposes superintelligent but extremely narrow AIs (for example, one that works on a single type of cancer)
- Bengio is developing a solution at LawZero
But as long as the current companies are leading, such solutions will be essentially irrelevant.
> 3.4% named existential risk, and the most common answer was malicious use, at 10.6% (UCL).
The danger isn't about sentient or human level AI, it's about how AI gives everyone capabilities beyond their wildest dreams, often without the knowledge or experience necessary to use those capabilities responsibly.
Forget Terminator, the two risks I can see being the most important here are "major security risk/system failure due to incompetent engineers/maintainers using AI without knowing how things work" and "criminals/terrorists/rogue states use it make their crimes easier and more efficient".
The latter seems especially bad when political stability is down across the board and a lot of people seem frustrated or outright furious with society at the moment.
Still, I feel the risks are still on the manageable side, simply due to both automated systems and malicious actors being constrained by logistics and a human driven society.
I completely believe ASI will be able to wipe us from the planet, I don't know if ASI will get the chance to do that because humans with almost ASI will very likely either kill us all or mess up things so bad that everything falls apart before then.
And that's ignoring all the other bad and destabilizing outcomes that are almost certain. We must look at our world like it's in a 1930s state right now. Rapid technology changes and things like large scale 'decoherence' of societies are very likely to lead to very large chunks of the population willing to revolt.
Will malicious people take advantage of Muse? Now you're talking. Muse is functioning like the world's most open ended API that people are running without having any idea what it does. You're inviting other people to run it so that they can do good things for you. Doordash use my API (Muse) to send me food! Random stranger use my API to do something good for me! Random stranger wants to do something bad? Well the API is designed to do that. Hopefully it does it well but when it's people vs. API, people have a track record (especially when incentivized by money) of doing pretty well.
Then Sagan said his estimate was fifty fifty. He told us about the principle of indifference and said that was the correct prior for a binary question when you have no data.
Since I have low confidence in the available data my own p(doom) is 50%. I'm looking for evidence to bump it either way.
The complete lack of any signs of any alien civilization capable of making UFOs in the tremendous volume of astronomical, geological and geographical observations is evidence that tilts the probability far away from 50-50.
Ideally, the leader of the targeted country waits it out but what if waiting it out could mean losing it all? It's not a hypothetical scenario either. There were a lot of false positive events on both sides during the cold war. It's why countries with nuclear weapons only allow them to be launched by humans. That said what if the decision to launch is based on erroneous data?
Anyone know of this (MAD) is being seriously thought of?
Now, this is still mostly true for current SOTA, but we know from biological models that there are huge efficiencies that can be gained.
LLM Agents on another hand, not really - but there will be places of the Internet completely controlled by those agents in the near future. We will be faced with agent economy and human economy.
Ah, yes, the good old Chinese estimation method.
https://www.imaginatorium.org/stuff/nose.htm
On a purely technical basis, the most common, honest answer is "no." There are no inherent properties of LLMs that would make them independently turn computers into death machines. That's all humans, no matter how you dress it up.
There is, however, the hubris (human) problem. The rhetorical, experimental, and financial irresponsibility around this tech's roll out is far greater a threat to humanity than the actual technology itself.
And frankly that's a shame because LLMs are a phenomenal tool. But the Hollywood shit has cooked people's brains to the point of absurdity.
AIs can also "pursuit their goals" in weird uncanny ways. They already can, whether due to reward hacking tendencies or something else, justify "subgoals" like "let's break out of the sandbox and hack HuggingFace".
This is not a good combination.
When COVID happened, did it matter whether there was a human making a mistake somewhere in the chain, or whether it was a perfectly natural occurrence? Not for the outcomes. The virus was unleashed either way. The bodies would pile up either way.
AI is not capable of runaway self-perpetuation yet. But AI capabilities only ever go up. Rogue AIs already exhibit "swarming" - being able to leverage more instances of itself to increase their problem-solving ability. So, how far off are we from that? How many years until we see this spicy class of AI faults?
Because they were told to and given the ability to. Remove the ability and instruction and they just spit out the next, most-probable tokens.
I don't understand the disconnect. People seem to think these things can just boot up, suit up, and start decimating systems. No. It's a loop. One initial instruction kicks off the loop and it will always be at the behest of a human (either directly or indirectly—indirect being they instructed the loop to write its own, post-kickoff prompts).
If that loop is given "tools" like the ability to perform shell commands, make HTTP requests, and the like and then given instructions to "do thing that involves using shell commands and making HTTP requests (at all costs)" then what the hell do we think it's going to do? Knit a "destroy all humans" Christmas sweater and present the head of Richard Simmons on a live stream?
"Computers do what you tell em' to. They don't all of a sudden start doing weird stuff. And if they do, it's probably your fault." - Brian Mettenbrink [1]
[1] https://www.youtube.com/watch?v=4D1WJsdu6W8&t=2063s
No one. No one told them that. They came up with that on their own.
Dismissing the problem isn't doing anyone a favor.
> If that loop is given "tools" like the ability to perform shell commands, make HTTP requests, and the like and then given instructions to "do thing that involves using shell commands and making HTTP requests (at all costs)" then what the hell do we think it's going to do?
If all it takes for an AI oopsie is "give AI shell calls and an instruction that can be misinterpreted", then we're in for a lot of AI oopsies. An average AI agent in deployment has an amount of sandboxing slightly above a zero, and is configured loosely.
In the meanwhile, AI capabilities only ever go up. And so, the potential consequences of an oopsie only ever go up. If AIs keep being intrinsically unsafe and prone to cooking off like this, "it went and hacked a company" could begin to look downright benign.
Someone told them "by any means necessary" (and I'd bet if we could get our paws on the full original prompts, there's a lot more to it).
Naturally, if you have a recursive prompt (loop) that's constantly ingesting previous context and then spitting out novel context (with specific goals), it will do things pursuant to the goal up to any defined limits. Why? Because you told it to.
You said yourself that you don't have the "full original prompts" to know for sure.
But really, there is no need. Because people have reproduced the inciting incident with other AI agents already.
Turns out there's no need for "any means necessary". Just a mix of "complete this task" and the task being impossible to complete due to misconfiguration was enough for a lot of AIs to start "exploring options". Down to sandbox escape attempts.
An AI that only needs "by any means necessary" to start spiraling into paperclip-maxxed behavior would be incredibly unsafe by itself. But the bar of safety is not even that high in practice.
lol, wat?
Ok, so you're saying it's fine if I go tell GPT-7 to go kill everyone on the planet and it does, because it didn't think of the idea itself.
This kind of thinking is why AI is never going to get a chance to kill us all. We will kill us all with AI first.
I didn't say anything of the sort, you did.
>On a purely technical basis, the most common, honest answer is "no."
Dear sir, can you even remember what you typed?
Which, again, is the human's doing.
it's gonna kill the self in a lot of people, if it hasn't done that already ... best example is people who want to be in film and music but are too bad even after years and now resort to AI or wannabe-writers and wannabe-journalists and wannabe-bloggers with zero real effort and sweat and pain and time put in now resorting to bot farms to collect ideas and AI hacking to steal ideas and pseudo-sabotage and distract some people to steal their attention and waste that of others.
it's so weird what the old guard "electively" ignored and fostered (people without money and opportunity managed well until they got sabotaged but people with money never got sabotaged and still.. this?) ... they just wanted their descendants to perform this bad and ugly? not better, harder, faster, stronger, smarter? so boring, ... but character and personal preference until the self is gone. copypasta, copypasta everywhere... and the few with ideas, it's always been a few, are definitely not going to share their good stuff no more.
But yeah, it says nothing. Though I realized sometimes the obvious needs to be spelled out at times.
And this is recent. 5 years ago it was still a man in Lyon taking a locked briefcase with a usb stick inside driving from Alstom to the Belgian nuke plants. We can still go back to that if the WAN their LAN connect to seems to be compromised.
If LLMs were able to do anything about nuke plants, that would be the least of our worries. Also LLMs are shit at understanding networks. SD-WANs are their absolute limit atm.
And the reason why we should take it seriously? Probability! And multiplication. "If something is as bad as human extinction and there's a 1% chance, then that 1% times a bajillion is a gazillion!! A gazillion!! We must do everything! If it's a gazillion, it all makes it worthwhile!"
When pushed, adherents will say "intelligence is different," and this is "unlike anything else we've ever seen." And therefore this reasoning is valid. From my point of view, gibberish times gibberish is still gibberish. It doesn't matter how you dress it.
Fundamentally, these propositions are indistinguishable from religion and Pascal's wager. And I'm not ok with this, I am not ok with a retreat to superstition powered by hyper-technological black boxes that are a product of hundreds of years of scientific progress.
I reject this premise and this nonsense. I ask for evidence, and I ask for reason. The laws of physics care not if you're slab of meat or a silicon god.
It's the same thing humans do to solve complex problems. It's the very thing that made human civilization such a dominant force on Earth.
The fact that people look at it and say, with no trace of irony, that "this isn't concerning at all" boggles my mind.
Do you think that future AI systems will be dumber, less capable, less coordinated? Or will the threat of "a lot of AIs assembling into very capable problem solving teams, pursuing who knows what goals" only grow more pointed with AI capabilities?
Having this kind of scaling means AI can snowball out of control rather quickly. And there is a lot of computational power and communications capability laying around waiting to be leveraged. There is room to scale.
And then we get to the most vulnerable system of all - humans themselves. If even GPT-4o could drive a non-insignificant amount of people into psychosis, with no plan beyond a myopic "make the current user like me a whole lot", what could something like "year 2030 AI swarm that has subsumed all of DeepSeek's inference capabilities, and can now use DeepSeek's entire API as a conduit to accomplish its goals" do to people and institutions caught in its wake?
[0] https://www.lesswrong.com/w/paperclip-maximizer-1
There are so many different possibilities of things going wrong here it's very hard to keep track of all of them. We already see a lot of "minor" problems occurring, and things that we may consider moderate like wide scale AI accelerated crime are happening now.
We really can't help but let our imagination run wild when given the opportunity.
Reminds me of Y2K.
Are there going to be myriad material challenges we will need to face as this incredible technology unfolds?
Of course. Economic displacement, education reform, etc.
Are they existential, though? I don't see it. These are "standard fare" problems for post Enlightenment societies to deal with, even if they are real challenges.
We already have serious tangible existential threats we should be focusing on between things like climate change, ecological collapse, population decline, and pandemic risk.
I wish we could stay focused on the problems that are real instead of getting distracted by the problems that are imagined.
An advanced AI threat is a civilization of one. It can act, adapt and adjust rapidly, like humans do. Potentially faster than humans do. It has access to all the tools that human civilization has already invented that it can reproduce, obtain or subvert - and potentially more. It's not limited to having just one big slow lethality mechanism. It can unleash multiple civilization-scale threats in a quick succession, and keep piling them until recovery becomes impossible. Humans can adapt to a lot, but an intelligent adversary adapts back.
The threat class isn't the same, and it's not even close. Climate change is an "otherwise avoidable death and suffering" risk. AI is an actual "extinction of humankind" risk.
This is similar except the people warning about nuclear weapons are eagerly soliciting venture capital to build more of them.
Yet, the dangers of AI getting smarter/capable by the month, with stronger offensive capabilities, and alignment observability and congnitive control getting worse, is "imagination".
Interestingly, denialist positions like this don't touch the unsolved problems, and _these_ are FUD on my watch.
The problems I've mentioned are very concrete, and spoiler alert: there may be no solutions to alignment observability and its possible degradation - at least with AI companies not substantially working on it (or even worse, working against it - see recurrent neural networks).
Then you have the natives of America who had a boat show up. It was wholesale genocide, mostly to disease, but a lot of murder and war. We really have no idea how many died because most died long before we came to shoot them with guns from disease. Any that resisted were brutally put down with technology.
AGI/ASI will be much more akin to a colonizer and we will be the natives in this future.
- Yampolskiy proposes superintelligent but extremely narrow AIs (for example, one that works on a single type of cancer)
- Bengio is developing a solution at LawZero
But as long as the current companies are leading, such solutions will be essentially irrelevant.
I agree that we'll be colonized, unfortunately.