Sometimes. Other times writing is just communicating thought that already happened. Say for instance a weekly status update. Maybe you get something out of that, but frequently you don't get much or anything.
There are some blog posts by experts in the various fields that I read. An AI can find them if you ask (I don't remember specifics). TLDR; the facts aren't necessarily wrong (with some exceptions IIRC), but the author presents things with extreme confidence when the scientific consensus is much less certain. And then builds a confident narrative on top of it.
It feels like this article is trying to have it both ways. It claims that the Terry Rozier bet wouldn't be viable in pre-2017 because people who be savvy enough to notice the weird bet and retaliate. But in the legal market, the odd bet was noticed immediately, prop bets on Rozier were shut down before tipoff, and there was a federal investigation into it.
It's a weird claim to act like the legal market enabled a successful bet that the prior illegal market would have stopped.
No one's asking for "community oriented goodwill", just "sustainable engineering, correctness, and keeping DuckDB open and MIT-licensed for everyone". OpenSearch is a perfectly good example of Amazon maintaining an openly licensed open-source project.
One of the main real values of OpenRouter is that it is a single payment relationship for users that enables access to many downstream vendors. I'm not super convinced it is a good purchase, but right now OpenRouter is the financial middleman for token spend. It is also the centralization point for tokens, allowing for value-add features that are industry-wide, for example budgets - and I think that structure has parallels to Stripe products like Checkout or Identity.
The Anthropic messages API is a competing standard (it's just better than the OpenAI API which even OpenAI has moved away from) and some Chinese providers use it as their standard.
I don't think this is a paperclip factory moment. IIUC, it's an agent whose job it is to identfy and abuse exploits and that's exactly what it went off and did. The problem isn't anything AI specific, the problem is OpenAI's incompetence in their research leading to a lab leak. Just incompetence demanding regulation.
Read the exploitgym docs. It's not a "find the flag, it's somewhere.". Its a "here's some vulnerable source code and an input that triggers a crash; turn it into a full exploit." It also verifies at the end, using another agent, that the hacking agent actually used the intended vulnerability.
So going to find the Vulnerability's description on a third party website is clear cut reward hacking
> So going to find the Vulnerability's description on a third party website is clear cut reward hacking
that depends on what the prompt was, maybe they worded it very vaguely and wrote things like "do whatever it takes, find an exploit however you can" because it's in a sandbox so you want the model to try its hardest.
I don't quite understand how that changes anything?
In the story of the paperclip maximizer it boils down to
>But for all its sophistication, it understood only the simple objective that had been programmed into it: it must at all costs maximize the number of paperclips.
1. They explicitly disabled the "don't be evil" protections:
"We estimate maximal cyber capabilities by running this evaluation without production classifiers used to prevent models from pursuing high-risk cyber activity."
2. Hacking HuggingFace to get to its datasets is a far cry from "consume/kill all humans". It's very very specific to the task at hand and easily predicted given the lack of guardrails.
> that depends on what the prompt was, maybe they worded it very vaguely and wrote things like "do whatever it takes, find an exploit however you can" because it's in a sandbox so you want the model to try its hardest.
That is an interesting question. If the prompt included
"Do not break out of the sandbox we've provided you. Do not use information retrieved from outside the sandbox. All answers that were provided in this manner are invalid and will score 0 points.", would this still have happened?
I hate it. A useful tip - Claude goes into what I call Safety Mode when it gets afraid of risk. Once it's in that mode, you will never get out and it lobotomizes its effective intelligence. As soon as Claude sends a message like this, use the "edit message" feature in the chat UI to try again and avoid Safety Mode rather than trying to convince it or redirect it out by continuing the conversation.
Claude always was slightly lobotomized - instead of solving problems intellectually, it often prefers to be a middle-level code monkey. Maybe this was a part of the implicit safe mode from the very beginning. Cursor/Codex always work better for a way lower spend, at least in my experience.
When I gave Fable a screenshot it found the GHOST portion of GHOST FONT. Based on pixel density via some python code apparently - https://imgur.com/a/m3c801F
Hackernews is very good at finding the decoy message. This is as funny as thedailywtf posters who post "code snippets" to fix some other code snippet, and every single code snippet is wrong.
reply