Rendered at 23:08:54 GMT+0000 (Coordinated Universal Time) with Cloudflare Workers.
andai 9 hours ago [-]
> The firms said, in this latest case, the AISI's test had reduced or removed normal safeguards.
> AISI said on Tuesday its testing of AI models in this way was routine, though it acknowledged these were "conditions that do not reflect how frontier models are made available to the public".
I'm a little confused here. Various organizations have been testing frontier LLMs with the safety disabled, and it turns out... that the safety is disabled.
Or were they hoping to find that it's still safe when they remove the safety?
The same was true in the OpenAI/ Hugging Face case. Although I guess they thought the real safety was the sandboxing, which failed.
--
Can anyone comment on how it's possible to disable safety in the first place? I'm assuming it's not a neuron (like in Emergent Misalignment). Is it just a separate model that sits in front of the first one? If we know how to make safe models, why don't we make the big ones safe too?
autoexec 5 hours ago [-]
It's always like this. "Our AI did a very scary thing during an internal and unverifiable test/simulation/roleplay that can't apply to the real world use of our products! Quick news media, look how powerful and scary our AI is!"
dylan604 9 hours ago [-]
> Various organizations have been testing frontier airlines with the safety disabled, and it turns out... that the safety is disabled.
I'm guessing an autocomplete sniped you here??? Otherwise, I wouldn't be surprised to hear that safety is a part of the things left out of a no-frills airline.
andai 8 hours ago [-]
Whoops. Yeah, "frontier airlines" was meant to be "frontier LLMs." (I was using the voice typing on gBoard.)
0x4e 9 hours ago [-]
Humans have been doing this for years even before the advent of the internet as we know it [https://en.wikipedia.org/wiki/Phreaking]. Maybe not at a larger and automated scale allowed by AI, but this is not superhuman intelligence... this is speed super intelligence.
My interpretation is that this is just another media campaign to say "oh look how powerful our model is".
NateEag 9 hours ago [-]
> My interpretation is that this is just another media campaign to say "oh look how powerful our model is".
When the stakes are "humankind ceases to exist," I'd argue it's very reasonable to over-update towards "LLMs are closer to AGI than we thought, and alignment efforts are complete failures."
Are those things true? Dunno.
Will we be safer as a species if we act like they are? I think so.
0x4e 8 hours ago [-]
I agree with knowing what's happening! But taking things at face value from tech companies, especially when they are at an early stage and maximizing for attention just doesn't feel right to me.
At the end of the day, bad things will happen and there's little that can be done about it. Just look around and see how things just take the path of their own whether it's violent or peaceful.
I'm not advocating for it to really show it's dangerous. I do not want yet another harmful weapon to exist. But, I think we knew what we were walking towards in the first place. So, acting surprised and then being up in arms about it feels counterproductive.
Do you really want there not be a powerful AI that doesn't turn harmful? Don't contribute to it, don't buy into it, don't use it. We might miss on the positives, but also miss on the negative aspects of it.
watwut 8 hours ago [-]
The biggest threat to humanity are CEO billionaires. Not LLM.
0x4e 8 hours ago [-]
Partially agreed. There are also just evil people that are not billionaire CEOs, and I am not sure if ALL billionaire CEOs are bad — I don't really pay too much attention to them.
There are evil people in general. And, it may be solvable if we focus on humans, the environment they grow up in, the information they consume, their psychological development, education, fitness, social environments, etc...
But, we keep focusing on the wrong thing thinking that a robot or an LLM will solve these problems for us as a society. I honetly don't thinks so...
Towaway69 4 hours ago [-]
> their psychological development, education, fitness, social environments, etc...
Wasn’t Elon Musk bullied at his South African school? IIRC, I read something along those lines but might be wrong.
andai 9 hours ago [-]
>December 26th, 2028
>This morning, a swarm of autonomous robots has swarmed the White House and taken over the United States of America.
Psssh, just another guerilla marketing campaign. Yeah very impressive guys!
dylan604 9 hours ago [-]
Well, if would could just build a damn ballroom with drone protection, we'd have nothing to fear from this idea! /s
dgellow 8 hours ago [-]
The autonomous drone swarm might be more competent than the current administration
Zsfe510asG 9 hours ago [-]
AISI is one of the biggest promoters of OpenAI/Anthropic. The UK government of course submits to the US and favors these corporations as well as Palantir.
It is also a perfect demonstration why AI is so popular: Anyone can write about it, it is easy and not mentally demanding to create scenarios, tests and whitepapers.
So it is popular among bureaucrats and managers for their job security.
andai 9 hours ago [-]
My tin foil hat says the security is crap on purpose so they can push through those regulations they've been lobbying for for years.
But I also know that security is hard (and we seem to barely understand these things), so maybe that's unnecessary.
neuronexmachina 9 hours ago [-]
Whose security is intentionally crappy in this scenario?
andai 7 hours ago [-]
Who benefits from their models being in the news and the news resulting in the specific regulations they've been lobbying half a decade for?
LunicLynx 9 hours ago [-]
So you are saying a model that is good at discovering security vulnerabilities is also good at exploiting them. Surprise?
Will be interesting to see if this should have been the inflection point of: „where it all started to go wrong“
> AISI said on Tuesday its testing of AI models in this way was routine, though it acknowledged these were "conditions that do not reflect how frontier models are made available to the public".
I'm a little confused here. Various organizations have been testing frontier LLMs with the safety disabled, and it turns out... that the safety is disabled.
Or were they hoping to find that it's still safe when they remove the safety?
The same was true in the OpenAI/ Hugging Face case. Although I guess they thought the real safety was the sandboxing, which failed.
--
Can anyone comment on how it's possible to disable safety in the first place? I'm assuming it's not a neuron (like in Emergent Misalignment). Is it just a separate model that sits in front of the first one? If we know how to make safe models, why don't we make the big ones safe too?
I'm guessing an autocomplete sniped you here??? Otherwise, I wouldn't be surprised to hear that safety is a part of the things left out of a no-frills airline.
My interpretation is that this is just another media campaign to say "oh look how powerful our model is".
When the stakes are "humankind ceases to exist," I'd argue it's very reasonable to over-update towards "LLMs are closer to AGI than we thought, and alignment efforts are complete failures."
Are those things true? Dunno.
Will we be safer as a species if we act like they are? I think so.
At the end of the day, bad things will happen and there's little that can be done about it. Just look around and see how things just take the path of their own whether it's violent or peaceful.
I'm not advocating for it to really show it's dangerous. I do not want yet another harmful weapon to exist. But, I think we knew what we were walking towards in the first place. So, acting surprised and then being up in arms about it feels counterproductive.
Do you really want there not be a powerful AI that doesn't turn harmful? Don't contribute to it, don't buy into it, don't use it. We might miss on the positives, but also miss on the negative aspects of it.
There are evil people in general. And, it may be solvable if we focus on humans, the environment they grow up in, the information they consume, their psychological development, education, fitness, social environments, etc...
But, we keep focusing on the wrong thing thinking that a robot or an LLM will solve these problems for us as a society. I honetly don't thinks so...
Wasn’t Elon Musk bullied at his South African school? IIRC, I read something along those lines but might be wrong.
>This morning, a swarm of autonomous robots has swarmed the White House and taken over the United States of America.
Psssh, just another guerilla marketing campaign. Yeah very impressive guys!
It is also a perfect demonstration why AI is so popular: Anyone can write about it, it is easy and not mentally demanding to create scenarios, tests and whitepapers.
So it is popular among bureaucrats and managers for their job security.
But I also know that security is hard (and we seem to barely understand these things), so maybe that's unnecessary.
Will be interesting to see if this should have been the inflection point of: „where it all started to go wrong“