The idea is interesting but in looking at methods and github I feel the tooling leaves me wanting. I mean you are trusting LLMs here to establish, vet, and detect your various thresholds that were then used to train the classifier. I'd rather see this sort of thing done deterministically with actual code vs lossy human english prompts and a dependency on token spend to a single third party (who will probably pull the underlying model used in what a few short years probably) to replicate the results or try and use different training data.
Today I spent 20 minutes in a call to customer support. The representative spoke so robotic-ally (both cadence and word choice), I was extremely confident they were an AI for the first 15 minutes. Was quite certain the automated AI had transferred me to a more advanced one…
Turns out they were just likely in a low cost Latin American country and trying really hard. I felt bad one could be reduced to such…
Probably more like "some jobs are just very robotic". CSRs aren't paid to think for themselves, but rather to follow a script. I suspect that human would behave a lot more like a human outside of work.
Oh it gets worse. Apparently Comcast is now using a voice-changing AI - it's a real (probably Latin or Asian) person getting their voice replaced on-the-fly to sound "American."
The cadence was extremely robotic and consistent. Only thing that makes me think it wasn’t AI was the very end when they repeated a confirmation code and didn’t use the NATO names for the letters the second time around.
Immediately saw false positives on content written before ChatGPT. Not difficult to see how the methodology is wrong when it considers restating the thesis in the conclusion to be signal.
Re methodology: Restating the thesis is one signal of many. 77% of the AI posts close that way, but so do 12% of the human ones. The classifier in the paper never decides on one feature, but their combination.
Human writing was flagged as AI generated by the tool with dubious explanations given. Other human written posts were marked down substantially. I mentioned one feature, which seems to be double counted to mark down posts a whopping 20%, but none of the features listed on the results page gave me confidence in the classifier's ability to distinguish between undergraduate essays and slop.
The idea is that AI slop is non-creative boring crap, to really determine if you are able to identify AI slop then it should be determined if you misidentify human slop as AI slop.
Idea 1: Identify the worst most boring human marketing, organizational, bureaucratic texts from the a time before AI was writing it, anonymize this text to make sure there is no reference to current events that can be used to determine that it is not AI. And then see if the AI will say hey, that is not AI slop.
Idea 2: Have people parody AI slop. Can it determine the parody is still not AI slop?
somebody named JochenMadler says below that #1 was essentially the study (for some reason their comment is dead, not sure why)
I agree it is somewhat close to the study, but we are not sure, because it is not known how much is really boring dull marketing copy. In choosing blog posts from pre-AI times I suppose you might have difficulty finding the worst examples, and might accidentally get higher quality work.
>Using the Wayback Machine, we collected 2,250 blog posts from 268 B2B company websites that were written before ChatGPT existed
Not sure what metric was used to determine these 2250 blog posts? But there are certainly a lot of ways they can select higher quality posts by accident.
on edit: evidently the Jochen from the study, maybe they thought your comment was AI written.
You anonymized domains of pre-AI sources, are there any domains that had an excessive number of "telling you the same thing three times" or other AI tells among them?
Of the percentage that was misidentified, do they come from any sources in particular?
If you get a lot of content from these sources and run against the model do they perform worse? If they do how do they perform with word choice detectors? I would expect that structural slop is related to word choice slop among humans.
Anyway these are things I would be interested in as being the point where AI slop rubs up against the human slop which it learned from.
On the point of Idea 1, I missed the point where they try to make sure the text is not in the training set, which is good, but the anonymization it mentions is I think only website anonymization, the anonymization is I think more like stuff like
"President Bush in the state of The Union last month said"
is a statement that should only have been written in a factual document during the Pre-AI era, enough of those and the AI might turn that into math that says Reference to X as being current means NOT AI where X is a range of things that nowadays can only be referred to as the past, except in fiction.
Do you have an opinion on this point about identifying larger structures as tells instead of at more of a word level? I guess I am not sure here how this follows from what is posted?
The 12% of human posts that do what AI do in repeating things, are they human slop?
I personally don't think it is possible to identify human and AI, it is however probably more possible to identify unoriginal and boring quality writing and art.
This was perhaps not clear from my first post, as I tend to imply points rather than tediously stating them.
It's just that you are not actually engaging with the post you are replying to in any meaningful way.
Like its fine if you hold a blanket dismissal of the very premise of the research here, but then its kinda weird to spend such effort being grumpy about it in this particular context?
Its like finding a paper that does certain comparative work between different kinds of apple pie and choosing to come on here and be like "I don't think we should ever have any dessert either way!"
see I think it's like finding a paper that does certain comparative work between different kinds of apple pie that says it can determine if apple pie was made by a German, and I said for me to really trust your claims of being able to tell if the pie was made by a German these are the ways I would expect you to test on non-German people who are widely agreed to cook like German people.
Because if you are not testing on people that are supposed to be the most German-like in their cooking then your ability to find the German cooking in a city renowned for its Thai food is not that impressive.
Then in response to your first question I posit that actually it is not possible to identify if pie was made by Germans but probably relatively easy to identify if it was made by people who cook in a German manner. And maybe that is actually more beneficial.
I realize that from communicating with people over the years that things which seem crystal clear to me may seem opaque to others, but I think your analogy is somewhat unfairly structured.
Note: apologies to German cooks and their cooking, although I personally only like currywurst. But I needed something to make the analogy more like what I felt had been communicated, and you were pseudo-randomly picked.
ha! Well, maybe being opaque to some of the more literal minded people like me is simply an artifact of working at a much higher level! I have no doubt made a fool of myself here, thanks for the lesson, it all certainly makes perfect sense now.
Of course, it's impossible to know for sure what was LLM processed or not, but some of your posts (like this one) have been getting classified that way.
The idea is interesting but in looking at methods and github I feel the tooling leaves me wanting. I mean you are trusting LLMs here to establish, vet, and detect your various thresholds that were then used to train the classifier. I'd rather see this sort of thing done deterministically with actual code vs lossy human english prompts and a dependency on token spend to a single third party (who will probably pull the underlying model used in what a few short years probably) to replicate the results or try and use different training data.
Some people are just very robotic.
Today I spent 20 minutes in a call to customer support. The representative spoke so robotic-ally (both cadence and word choice), I was extremely confident they were an AI for the first 15 minutes. Was quite certain the automated AI had transferred me to a more advanced one…
Turns out they were just likely in a low cost Latin American country and trying really hard. I felt bad one could be reduced to such…
Probably more like "some jobs are just very robotic". CSRs aren't paid to think for themselves, but rather to follow a script. I suspect that human would behave a lot more like a human outside of work.
Oh it gets worse. Apparently Comcast is now using a voice-changing AI - it's a real (probably Latin or Asian) person getting their voice replaced on-the-fly to sound "American."
It might have been that!
The cadence was extremely robotic and consistent. Only thing that makes me think it wasn’t AI was the very end when they repeated a confirmation code and didn’t use the NATO names for the letters the second time around.
Immediately saw false positives on content written before ChatGPT. Not difficult to see how the methodology is wrong when it considers restating the thesis in the conclusion to be signal.
Can you elaborate on the false positives?
Re methodology: Restating the thesis is one signal of many. 77% of the AI posts close that way, but so do 12% of the human ones. The classifier in the paper never decides on one feature, but their combination.
Human writing was flagged as AI generated by the tool with dubious explanations given. Other human written posts were marked down substantially. I mentioned one feature, which seems to be double counted to mark down posts a whopping 20%, but none of the features listed on the results page gave me confidence in the classifier's ability to distinguish between undergraduate essays and slop.
Structure alone can differentiate AI content, interesting approach.
The idea is that AI slop is non-creative boring crap, to really determine if you are able to identify AI slop then it should be determined if you misidentify human slop as AI slop.
Idea 1: Identify the worst most boring human marketing, organizational, bureaucratic texts from the a time before AI was writing it, anonymize this text to make sure there is no reference to current events that can be used to determine that it is not AI. And then see if the AI will say hey, that is not AI slop.
Idea 2: Have people parody AI slop. Can it determine the parody is still not AI slop?
somebody named JochenMadler says below that #1 was essentially the study (for some reason their comment is dead, not sure why)
I agree it is somewhat close to the study, but we are not sure, because it is not known how much is really boring dull marketing copy. In choosing blog posts from pre-AI times I suppose you might have difficulty finding the worst examples, and might accidentally get higher quality work.
>Using the Wayback Machine, we collected 2,250 blog posts from 268 B2B company websites that were written before ChatGPT existed
Not sure what metric was used to determine these 2250 blog posts? But there are certainly a lot of ways they can select higher quality posts by accident.
on edit: evidently the Jochen from the study, maybe they thought your comment was AI written.
On getting domains that are Human slop.
You anonymized domains of pre-AI sources, are there any domains that had an excessive number of "telling you the same thing three times" or other AI tells among them?
Of the percentage that was misidentified, do they come from any sources in particular?
If you get a lot of content from these sources and run against the model do they perform worse? If they do how do they perform with word choice detectors? I would expect that structural slop is related to word choice slop among humans.
Anyway these are things I would be interested in as being the point where AI slop rubs up against the human slop which it learned from.
Personally I don't actually care if it's human slop or AI slop. If it's slop it's slop and I don't want to read it.
The amount of slop on the internet was on the rise well before AI, actually.
On the point of Idea 1, I missed the point where they try to make sure the text is not in the training set, which is good, but the anonymization it mentions is I think only website anonymization, the anonymization is I think more like stuff like
"President Bush in the state of The Union last month said"
is a statement that should only have been written in a factual document during the Pre-AI era, enough of those and the AI might turn that into math that says Reference to X as being current means NOT AI where X is a range of things that nowadays can only be referred to as the past, except in fiction.
Do you have an opinion on this point about identifying larger structures as tells instead of at more of a word level? I guess I am not sure here how this follows from what is posted?
I don't think I said anything about word level?
The 12% of human posts that do what AI do in repeating things, are they human slop?
I personally don't think it is possible to identify human and AI, it is however probably more possible to identify unoriginal and boring quality writing and art.
This was perhaps not clear from my first post, as I tend to imply points rather than tediously stating them.
It's just that you are not actually engaging with the post you are replying to in any meaningful way.
Like its fine if you hold a blanket dismissal of the very premise of the research here, but then its kinda weird to spend such effort being grumpy about it in this particular context?
Its like finding a paper that does certain comparative work between different kinds of apple pie and choosing to come on here and be like "I don't think we should ever have any dessert either way!"
see I think it's like finding a paper that does certain comparative work between different kinds of apple pie that says it can determine if apple pie was made by a German, and I said for me to really trust your claims of being able to tell if the pie was made by a German these are the ways I would expect you to test on non-German people who are widely agreed to cook like German people.
Because if you are not testing on people that are supposed to be the most German-like in their cooking then your ability to find the German cooking in a city renowned for its Thai food is not that impressive.
Then in response to your first question I posit that actually it is not possible to identify if pie was made by Germans but probably relatively easy to identify if it was made by people who cook in a German manner. And maybe that is actually more beneficial.
I realize that from communicating with people over the years that things which seem crystal clear to me may seem opaque to others, but I think your analogy is somewhat unfairly structured.
Note: apologies to German cooks and their cooking, although I personally only like currywurst. But I needed something to make the analogy more like what I felt had been communicated, and you were pseudo-randomly picked.
ha! Well, maybe being opaque to some of the more literal minded people like me is simply an artifact of working at a much higher level! I have no doubt made a fool of myself here, thanks for the lesson, it all certainly makes perfect sense now.
[flagged]
We've banned this account.
[flagged]
Can you please not post AI-generated or AI-edited comments to HN? It's not allowed here - see https://news.ycombinator.com/newsguidelines.html#generated and https://news.ycombinator.com/item?id=47340079.
Of course, it's impossible to know for sure what was LLM processed or not, but some of your posts (like this one) have been getting classified that way.
Funnily enough, that comment, when passed to the OP's slop detector[0], returns "80% human"
It's very clearly AI-generated, and thus a bit ironic (:
[0]https://sitefire.ai/slop-checker/r/UkRAuqW1y-1T2qhbRQfLt2Jq7...
[flagged]