Hacker Newsnew | past | comments | ask | show | jobs | submit | TimPC's commentslogin

It seems to me that recent changes have dramatically exacerbated this problem. Pangram is still claiming their extremely low false positive rate from before adding humanization detection and now very frequently detects actual human works as humanized AI.


There seems to be a huge backlash to existing works where AI was detected already at least in my art (creative writing). So I think we already have evidence consumers care. I don't think blockchaining is the solution for the general public, but I'd probably do exactly that if I was taking a university course these days since they tend to be hyper sensitive about AI and err on the side of the accuser so I want strong evidence my writing is mine. Which of course is ridiculously insufferable because I used to do a good portion of my homework in paper notebooks on the subway without stable internet.


LLMs still write middling quality prose. Aspects of it are good but they write in a very tropey manufactured style that isn't something we are fatigued of, it's something that's inherently flawed and disliked.


A lot of writing publications are so overwhelmed by submissions that they run an AI tool on everything and don't even read the ones that fail (for whatever threshold they think is best). Plus look at the recent backlash against the book where the author according to detectors used about three quarters AI gen. If you do your art for anyone but yourself you care a great deal about your work showing as human and being treated as having soul rather than testing as machine and being treated as witch hunt grade slop.


Pangram’s market claims are very different from independent validations of the platform. The 99.98% rate is on a toy dataset not resembling reality.

Even if 2/10000 is true that’s nowhere near accurate enough to make aggressive accusations that create anxiety at levels people need to medicate with potentially fatal consequences.


Someone also input lots of pre-GPT texts to Pangram and they were all clear. The point is Kramnik’s claims are much weaker and not comparable.


> Pangram’s market claims are very different from independent validations of the platform.

Do you have a link to independent validations? The ones I found, like an upcoming nber paper confirm the companies claims.


https://github.com/deepanwadhwa/ai_detector_fails/blob/main/...

the text in above is all ai generated. i uploaded a picture of two kittens to gemini and asked it to guess their age. then asked gemini to change the language to of someone who is born in deep Appalachia but has Japanese parents. then asked it to change the bullets to paragraphs written by someone who is talking fro life experience rather than grounded in published science. pangram gave it 100% human generated. go ahead, try it. it boggles me how people believe these ai detectors.

if it makes you feel good the original response from gemini gets 100% AI detection and for that I think pangram should get the credit but in practical terms the empty pots in my yard are more useful than pangram at this point.


You are talking about false negatives, but this thread started with a discussion around the potential for false positives causing writers anxiety

> Even if 2/10000 is true that’s nowhere near accurate enough to make aggressive accusations that create anxiety at levels people need to medicate with potentially fatal consequences.

False positives and false negatives are different problems that have different impacts.

Does your single example prove that Pangram doesn't work, or that it doesn't work on short snippets of text? Try to get a detection error on a longer run of text. I'd be curious to see the results, particularly if you can get a false positive.


Well, false positives must anyway be zero for a product which is built around detecting AI. let me explain - the value proposition of this product is that we detect AI. Its not we don't call human writing AI generated. The latter is implied and must not happen ever. The value is that pangram detects AI generated content which it did not in the above example. if this explanation doesn't make sense then think of pangram as a pregnancy test where presence of AI == presence of baby. If a pregnancy test keeps saying you are not pregnant when you actually are then that's a problem, right? - you can't argue that at least its not saying you are pregnant when you are not.

But here's a bigger problem - Imagine a teacher grading 50 students - 40 of them use AI to write their answers and 10 write honestly. All 40 of them use the hack that I used and get 100% human from pangram and the other 10 also get 100 human from pangram. what is the product really adding to the workflow? Nothing, zero or zilch!!! - but then you come along and say hey at least those 10 honest students also got 100% from pangram. This whole argument of false positives being very very low works if your false negatives are tight which they are not.


Neutrality is impossible. Left, right and centre are all relative terms that vary dramatically across cultures. A model that is slightly left in the U.S. is right wing in most of the world.


I've heard people say this before, but is it true? Aren't large chunks of Asia, Africa and South America further right than the US often times?


I think pop musicians are capable of doing greater works later, but the perception of pop works are so heavily influenced by the image/presentation of the artist that we view the works as lesser. I don't think there is something fundamentally different about pop music that leads to best works being earlier relative to other genres of music beyond that.


A great deal of pop music, performed by teens-20yos, is written and produced by seasoned professionals who are in their 30s-40s-50s.

The exceptions to that pattern are remarkable.


If we limit the definition of pop music to what charts I think it makes all the sense in the world that it is a young person’s game. So much of what drives chart success is what is in fashion at the time. Trend setting will always be the domain of youngsters.

If we expand the definition of pop music to all music that isn’t classical/jazz/experimental, etc. then older, more experienced musicians should be able to do quite well. Frank Sinatra honed his craft over the decades. I think the stuff he did in his 40s and 50s is probably his best.


> So much of what drives chart success is what is in fashion at the time. Trend setting will always be the domain of youngsters.

I would suggest it's more the demands of poverty that make it a young person's game. So, so, so many pop musicians were "I was living in squalor for a decade plus was extremely depressed and was about to hang it up when <thing happened> and we got popular." Huey Lewis, Annie Lennox, ... I can go on and on.

There was a metal artist that was being interviewed about when they were going to tour again and was "Yeah, we'll consider it. But I've got a lot of work at my tattoo business right now." There was another guy that was like "Yeah, had this fame hit in our 20s this would be nice but in our late 30s it isn't really useful. We figured out how to do life by now, and we're not going to disrupt that."


Elderly retirees seldom move. They’ve long paid off their mortgage and we keep property taxes extremely low so they are pretty much immune to the cost of housing. Once you put down roots in a community it’s hard to leave.


They also usually don't want to leave their established doctors. This is actually one of the reasons why we need mixed-housing options within neighborhoods so that elderly people can downsize into more manageable one and two bedroom apartments, condos, or duplexes etc without having to leave the neighborhood. Downsize into housing stock without stairs, without a large yard to upkeep, downsize into something smaller that would be more affordable to adapt for someone aging in place.


It's easier to define safeguards and the definition of inside information for stock markets than for prediction markets though. There is plenty of information that should ban someone from prediction markets that also wouldn't meet the definition of material non-public information.


OpenAI is a bet on LLMs replacing a large chunk of the labour force in whatever sector it’s best at replacing. It’s essentially looking to get companies to pay $5k-$10k a month to have coding agents replace the output of a single software engineer.

If the S-curve levels off below that level OpenAI will be an unsuccessful company.


Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: