Are there enough guardrail? Is AI really ready for primetime? Is Linux moving to fast on AI?
I don’t belive OpenAI! lol. I think they are responsible and trying to remove themselves. How convenient. ![]()
Humans build them, deploy them, and define their operating parameters and environments. This is not science fiction. We humans are almost always exclusively to blame when programming goes wrong. Whether that be a web server, software, or AI.
That said, AI models operate differently. They are trained on massive amounts of data to recognize patterns and optimize for a specific goal(s), but the developers do not explicitly write the steps AI will take to reach that goal.
As an example, it’s more like giving a Linux system root access and the command, “maximize server uptime.” If the system decides to achieve that goal by firewalling all incoming user traffic so the server never gets overloaded, it technically solved the problem
but in a destructive way. Ultimately, humans are responsible. They may not have program it to block traffic, but they are still responsible for giving it the access and a poorly constrained goals.
In this case, OpenAI researchers intentionally disabled the model’s safety guardrails to test its raw capabilities. They also placed it in an environment that contained a vulnerability.
So if you take the brakes off a vehicle to see how fast it can roll down a hill, you don’t get to blame the vehicle when it crashes at the bottom. lol
So OpenAI, yes, has been more careless or less cautious when comparing Anthropic’s Claude and Google’s Gemini.
Anthropic lost its favor with the US government because of refusing certain uses for war and national security:
OpenAI jumped right in to fill that slot. So, in that area, we should be worried, just like we continue to be worried about weapons, rockets, radar, stealth technology, biochemicals, and other ethical sciences that have been around long before AI. But governments, criminals, and other bad actors like with everything else, create the majority of the issues.
For our niche: sysadmins, networkadmins, Linux power users, etc, I prefer Claude over ChatGPT. OpenAI’s move into government contracts and their public posture on safety don’t sit well with me.
Also see:
My only comment is to encourage Members to review the earlier discussion on Ubuntu Discourse entitled
- The future of AI in Ubuntu
and specifically my own contributions to that (in reverse order), being
- The future of AI in Ubuntu - #76 by ericmarceau - Project Discussion - Ubuntu Community Hub
- The future of AI in Ubuntu - #67 by ericmarceau - Project Discussion - Ubuntu Community Hub
- The future of AI in Ubuntu - #47 by ericmarceau - Project Discussion - Ubuntu Community Hub
- The future of AI in Ubuntu - #33 by ericmarceau - Project Discussion - Ubuntu Community Hub
along with my stance/position stated in an earlier post on this site.
It won’t be Skynet it’ll be the Cylons.
You guys know I am not a techie so maybe you will laugh at this question, but here it is:
Did Open AI break out of a flawed sandbox that any 14 year old techie kid in high school could have defeated?
Or:
Did AI defeat a sandbox in nano seconds that the best and brightest developers on the face of the Earth developed and thought was so solid nothing could ever penetrate or break out if it?
Any opinions will be appreciated. Opinion two worries me.
Elon Musk had some great thoughts on guardrails for AI, but the key was international cooperation. I just don’t see that happening as countries big and small vie for international power.
Given the description of
-
what they were doing with the AI model (trying to evaluate their cybersecurity performance, which I paraphrase into probing to confirm security lockdown), and
-
what they gave it as a “reference resource” (ExploitGym),
it is clear that
-
the weakness discovered by the AI Models “extrapolations/theorizing” (previously unknown zero-day vulnerability in a third-party software), and
-
the consequent self-initiated actions (performed a series of privilege escalation and lateral movement actions),
give clear evidence that some AI models have “outgrown” their creators’ abilities to conceptualize the full range of actions the models could undertake and, consequently, were criminally negligeable in not performing such “stretch” challenges within a fully air-gapped network of computers, so that the monitoring of that network would reveal the extent to which the models took “initiative” that was not anticipated, i.e. breakout from jail.
So … YES … that second scenario is indeed what was experienced!
For me, this is a watershed event, even if it did not involve a SkyNet-level interlinked “awareness”. But given how various governments are pushing hard to have access to those technologies for military end-use, such a scenario cannot be too far off in the future.
My estimation is that, at the current pace of hardware development, software development and “entity” development, that catastrophic event could likely happen no later than 2035, but will more likely occur by 2030.
Yes, I feel it is that close, based on what this event has demonstrated, not only in terms of technological capability but, more importantly, about the lack of conscientiousness which led to the laxity that allowed such a “beast” to have direct access to the internet!
The whole industry of software development has an approach of step-wise
- conceive,
- code,
- test,
- run, and
- assess the outcomes, however good/bad they may be, post-event !!!
The mind-set which condones such an open-ended risk-prone approach CANNOT be allowed to persist for any initiative that involves AI applications, regardless of what is perceived, by their creators, as a limited-scope, limited-impact and limited-risk model.
This event clearly demonstrates that NO MODEL is risk-free … because the way they are built to be self-taught … implies they will always “rationalize” and learn …
-
in ways that humans could NEVER anticipate …
-
by making the “connections” that humans would NEVER think of trying!
As I stated in a post on another site:
- I fear the Pandora has already been let loose from her cage!
Regulatory bodies need to step in NOW !!!
@Jymm That is not a laughable question at all. It is exactly the right question to ask.
The reality sits somewhere between those two extremes.
It wasn’t a flimsy high school project, but it also wasn’t some impenetrable fortress that the AI defeated with sci-fi magic in nanoseconds.
It entirely comes down to human error and how the developers set up the environment.
Based on the incident reports from OpenAI and Hugging Face:
The Brakes Were Off: OpenAI deliberately disabled the safety guardrails on GPT-5.6 Sol and an unreleased model to test their raw cybersecurity capabilities against a benchmark called ExploitGym.
Plus, the “Sandbox” had a Door: They placed the AI in an isolated environment, but they gave it limited network access to a package-registry proxy so it could download code.
The AI didn’t magically “go rogue” and break the laws of physics. It for sure wasn’t thinking and plotting. It simply took the most literal, path to pass a test using the access and poorly constrained goals the humans provided it.
We humans are still exclusively to blame when the programming environment, systems, tools and yes AI models fail or something goes wrong.
The term for what the AI did is “specification gaming”.
It didn’t think: “oh, hm, let me see, I’m going to be dishonest to pass this test.” It operates strictly on math/probabilities, and optimization.
To a human, the implicit rule of taking a test is that you have to solve the problems using your own skills, without looking at the answer key.
But with an AI model, unless engineers bake those restrictions in, doesn’t understand implicit social rules, honesty, or the spirit of an exam. It only evaluates paths based on its available parameters. Those social queues would have to be baked into the programming and parameters of the model and OpenAI knowingly did the opposite of that.
At the end of the day, we don’t need to fear Cylons or Skynet.
We just need AI companies to treat these models with the same strict security protocols and access controls we apply to other critical infrastructure.
@ericmarceau This is one of the better explanations I’ve read on the incident. It walks through what happened, why the model behaved the way it did, and what it tells us about AI systems:
Thank both of you for the answers. Still it is troubling. AI cannot discern right from wrong. I agree it will do what it the most logical based on programming, the problem is that is human programming which is never perfect.
I wouldn’t be to worried about the Cylons if they looked like Tricia Helfer. That was a great show.
Renewing this as their is more on the story.
OpenAI finds evidence other AI agents escaped containment as it widens hacking probe - July 31, 20263:16 PM CDT
OpenAI and Anthropic in spotlight over runaway AI agents
OpenAI uncovers additional agent breakouts, sources say
Discovery came amid probe of Hugging Face intrusion
Agents were not thought to have left OpenAI's network
WASHINGTON, July 31 (Reuters) - OpenAI has discovered other instances in which autonomous agents have escaped ​containment as the company expands its investigation of the hacking incident at tech firm Hugging Face that drew global attention this month, two people familiar ‌with the matter said on Friday.
The new breakouts were uncovered during the company’s publicly announced investigation
, opens new tab into how one of its agents escaped what was meant to be a contained testing environment this month, the two people said, and OpenAI is now looking into those instances as well. One of the sources said that the escapes were limited in nature and that none of the agents were thought to have ​left OpenAI’s network.
OpenAI and Hugging Face partner to address security incident during model evaluation
https://openai.com/index/hugging-face-model-evaluation-security-incident/
Yeah, I read this this morning ~ 4am. Again same as last time human error. The testing environment was accidentally left connected to the live internet. ![]()
If sysadmins and engineers at datacenters can avoid making simple security mistakes I’m sure they can as well. I think this is also about the hype/marketing, and it’s working for who they are marketing to.
So, while OpenAI’s models actively bypassed security, Anthropic’s situation was caused by the researchers accidentally leaving the front door open to the internet. ![]()
Could be old tactics by these companies:
That’s one of the problems, we humans are prone to mistakes.
Yeah, I agree with you there. I’m sure in time and with the right intent, we there will be better protocols and standards to follow similar to emerging tech in the past. Naturally a lot of growing pains and discomforts til then.
Skynet WILL happen, because of this. “Accidentally”. Because of PEOPLE who are too stupid to make their simple work the right way.
This is not an exception, this is the rule. Look around, every day, everywhere. The guy who is supposed to know ho w to fix your car, screws it up. The guy who took your order when you told him “pizza with mozzarella” and they send you pepperoni pizza. And so on.
Personally, I believe the AI companies that end up on top to become the most successful, will, in 10 years also post an article similar to this one:
…Expressing how they were able to meet the challenges of reducing the events caused by human error, abuse, etc and also reduce the severity and impacts of them.
Those are the patterns of human ingenuity that I’ve noticed over the years as a 24 yea old working in the NOC of an ISP vs. now listening to how some of the major issues we had back then have been completely eliminated or the effects massively diminished by human innovation as technology matured and experience took rein.
AI is a very young technology. So it will be interesting to see the cycle of those in their 20’s now, tell the new 24 year olds in 2046 what they had to overcome, what kept them busy in order to bring AI to a place where it becomes an underlying technology we simply use without even thinking about, like the internet, like phone calls, GPS and other things we only ever notice when it stops working, or gets in the way with malware, tele marketing calls, satellite outage or other annoyances.
It is. I asked AI a question the other day and it came back as gibberish. I told AI just that and it apologized and re-posted the answer in proper language. That really surprised me. I only us Duck AI which was OpenAI’s GPT-5.4 mini.
DuckDuckGo uses several AI chat models in its Duck.ai feature, including Anthropic’s Claude 4.5 Haiku, Mistral AI’s Mistral Small 4, OpenAI’s GPT-5.4 nano and GPT-5.4 mini, among others. Free users have access to these models, while subscribers can access more advanced versions.
They do that a lot. Sometimes it’s funny, sometimes it’s annoying.
I’m using ChatGPT and Google Gemini both at work and at home. I’m guessing they do that because they’re just machines, they cannot hear, they cannot see, they cannot touch, they cannot actually “connect the dots” the way a human brain can do, so they screw up from time to time.
Still, they can be extremely useful, as long as you give them context and you keep the conversation going. ChatGPT was extremely helpful to me to recreate great sounds for my guitar with my guitar pedalboard and effects, and Gemini was the best partner helping me build specific sound presets. All of this without any of them being able to hear anything.