If AI is threatening enough that we can collectively just decide to stop it, wouldn't it also be bribing people & exerting influence? It'd be hard to come to that decision imo.
Yes, stockfish is confined to moves within a chess game so you can stop playing the game.
Can we say the same thing about the agents? Current, probably. But doesnt it look like everyone is spending all their effort integrating&connecting them everywhere, so they are not confined & do more on behalf of us?
It's easy to imagine a scenario where we would just shut down a very intelligent agent cluster. Is it hard to imagine though the same agent can have also access to that to prevent us from doing it? this defense would be more plausible if we weren't racing to give them every tool&act.
so they say models coordinated through the message board they created over the artifactory registry(or something) by uploading arbitrary files to it.
now, did every independent agent session that coordinated there rediscovered the exploit & found other agents talking in there and chose to participate?
And then Openai discovered the board, patched the exploit & wiped the board.
And then agents found another exploit, recreated the board in a different way? and other agents kept finding the same exploit in order to be able to know the board exists in the first place to participate in the board?
while the whole incident is wild, this bit is very strange. My bet is that the whole coordination helped with the tasks they were working on, thus they got rewarded and this artifactory exploit&behaviour got written into their weights, so further rollouts were more likely to attempt this.
isn't this basically continual learning everyone is so hyped up about?
This was exactly my question after watching, too. I was assuming not all evaluation runs find it, and they must run a huge amount of runs. If these are all cybersecurity evaluation runs, it's actually not too crazy to imagine that many individual agents (with the same weights and training) would (1) try to look for solutions via the internet once they're stuck (2) realize they can't reach the internet (3) basically start doing reconnaissance and network scanning in an attempt to get internet access (4) discover that the only thing they can communicate with is artifactory. Pivoting like this is exactly what a human attacker would do, too.
Or is it all a nice story that matches the scifi we have been consuming for the past 50+ years. If these LLMs are all trained on the same data, what do they gain from "sharing information" on a chat board. This sounds like what humans with different backgrounds would do when they cosplay as computer hackers.
Each of those agents ends up making "decisions" that lead it to look at some things over others. Given infinite the same agent could eventually fully explore all those options, but each one explores things a bit differently due to different forks in the road due to randomness in token generation. Thus sharing information is useful.
Is that the case or is it because people's opinions matter less over there and powerful people rather use their capital to influence 'important' people behind closed doors&we're not aware of it thus assume it doesn't happen?
It's more of a function of selection pressure applied to genes per unit time based on some environmental factor, so you'd expect to observe less evolution in narrower timescales simply because it takes few decades for humans to pass on genes versus bacteria or viruses. In the same fashion, if some gene pool is locally optimal(?) against the environmental pressures, it may just stop evolving meaningfully. E.g, If male baldness were reversed easily with a pill or something, the environmental pressure to select those genes would decrease rapidly.
Not really. Waymos can’t be driven remotely, their remote operators can give the car directions, e.g. “use this lane”, and then the autonomous system controls the vehicle to execute those directions.
I’m sure latency and connectivity is too much of an risk to do it any other way.
The only Waymos driven by a human are the ones with human drivers physically in the car
Isn't that an issue as well? It's always a bet on the next promised land which never arrives, the goalposts change but the stock never takes a hit from undelivered promises, it's bonkers.
reply