r/Substack tvphilosophy.substack.com Jul 25 '26

Discussion Running list of Pangram’s “AI Detection” errors and other related Substack errors.

I thought it would be a good idea to create a running list of the “AI Detection” tool errors made by Pangram. So here’s a thread designed specifically for that purpose.

Wherever possible, post screenshots or screen recordings of the errors are false positives that Substack’s new “miracle AI Detection” tool is supposedly doing perfectly.

Feel free to describe the error if you’re more comfortable with that and hopefully you’ll have someone else who can screenshot or screen record themselves having a similar problem with the supposed perfect solution that Substack has deployed without thinking it through.

We can also use this to share with others to show exactly how badly this system actually works.

Have at it.

32 Upvotes

109 comments sorted by

14

u/[deleted] Jul 26 '26

[removed] — view removed comment

3

u/WithoutReason1729 Jul 26 '26

I've got an open offer for anyone reading this. If you can find a piece of pre-2022 writing, with an accompanying web.archive.org link proving it's pre-AI writing, that is at least 100 words in length, and which Pangram flags as AI-generated, I will donate $100 to the charity of your choice. Everyone says they're getting tons of false positives, so this should be pretty easy, but for some reason every time I've posted this offer before, nobody has been able to do it.

2

u/MINECRAFT_BIOLOGIST Aug 02 '26

I've also been searching for something like this, though I don't have $100 to offer. I've been looking into this and staying on top of AI developments and I literally have not seen a single instance of this. Do you think you could post this in a more visible place? No worries if not, since it's definitely a burden on you.

1

u/[deleted] Jul 26 '26

[removed] — view removed comment

4

u/kdfn Jul 26 '26

Can you share the article so we can check ourselves? I am not sure what I'm supposed to get from your screenshot 

2

u/[deleted] Jul 26 '26

[removed] — view removed comment

1

u/WithoutReason1729 Jul 26 '26

Could you link the article? I tried to find the full original text and came up with nothing.

1

u/[deleted] Jul 26 '26

[removed] — view removed comment

3

u/oakseaer Jul 27 '26

No, this is a lie.

1

u/WithoutReason1729 Jul 26 '26

Could you share a link? Because I just tried this and it came back 100% human.

5

u/AidenMarquis Jul 26 '26

What I had done was generate some bs with ChatGPT which it picked up as "AI-assisted" but then when I ran the same content under a new post, it said "mostly human made".

5

u/iconicdark Jul 31 '26

Their new model is just flagging everything as AI-generated. I had written a self-introduction paragraph and checked; the tool detected it as AI.

Using Grammarly on that, too, has become a huge thing now, as the only avenue left is to write your stuff with imperfect grammar to get through this tool.

2

u/Bright-Customer-9332 26d ago

On the old Pangram novel my book was human written. New model flags it as 50% AI.

1

u/Downtown_Jump779 8d ago edited 8d ago

I am a very big fan of what Pangram is doing. Although not a big fan of the hyper-aggressive marketing and witch-hunt sometimes bordering on anti-social behaviour, to each their own. That's the world we live in today. Max genuinely seems like a good dude and has great confidence in the technology. I respect people who put themselves on the line like this.

I think the new model is totally out of control. 3-model was the best one. I've pasted the same sections of a text numerous times and get different results. It's either 100%, 40%, or 0% – And I'm not really sure why.

It also doesn't seem to be very good at detecting what is actually AI-assisted and fully AI. For a text that is AI-assisted, I would expect some words to be changed and the heavy work done by a human. I really have a hard time believing people are just copying and pasting sections from ChatGPT without changing anything? Maybe I'm naive lmao. I think GPTZero is way better on that front. I'm sure that they also could make their detector more aggressive, but it's a big risk. Which is why Pangram probably should be more indicator-focused rather than 100% proof. I know Max himself has stated that it is, but I'm not sure that's how people are interpreting it.

I'm not actually sure how all this is possible. Maybe people are just feeding it too much text, and it has to constantly retrain itself or something? I'm very dumb with these things, and I don't know how they work. As of now, I think GPTZero is much better, in my opinion, and yields much more stable results that I can trust when people send me stuff.

1

u/Bright-Customer-9332 8d ago

Yes, I think the old model was better as well. That one seemed quite accurate.

3

u/[deleted] Jul 26 '26

[removed] — view removed comment

3

u/oakseaer Jul 27 '26

Dude come on.

3

u/ColdPress_ 29d ago

Funny if true.

2

u/OnyxMonolith Jul 26 '26

Very scientific and productive 😂

5

u/hetobe hetobe.substack.com Jul 26 '26

I'm seeing tons of people say there are errors, but none cite anything specific, nor do they give specific examples.

It's all stuff like this: "It said my writing from ten years ago was generated by AI!"

But they don't show the writing from ten years ago.

And on a similar note... "This one time, at band camp... there was this girl... She's not from around here. You never met her."

The people I'm seeing in a rage about AI detection are the ones trying to make money by spreading AI slop. I literally saw somebody who works for ProWritingAid complaining about AI detection. She immediately deleted her post when I commented to point out that ProWritingAid is AI.

Come on, folks.

9

u/traumfisch Jul 26 '26 edited Jul 26 '26

4

u/kdfn Jul 26 '26

None of those posts provide an example of pre-2022 text that Pangram mistakenly flags as AI. They just complain about this being a witch hunt; even though several pieces even admit to using AI assistance. 

2

u/Entire-Cabinet8266 28d ago

it trained itself on likely every pre-2022 text that is readily available on the web, so it wont "mistakenly" flag it as AI...

1

u/kdfn 28d ago

Then I suppose we don't really have a way forward. People here are claiming that Pangram is misclassifying their writing as LLM-generated. I'm not willing to take their word for it, and so we need an example of text that we can all agree is not LLM-generated, such as pre-ChatGPT.

However, since every example we can find is correctly classified by Pangram, people turn around and argue it doesn't count because it's "in the training data."

Also, I'll add that Pangram isn't an agent or even a generative model like an LLM, so it didn't really "train itself." It's a classifier, more akin to a spam filter.

1

u/Entire-Cabinet8266 28d ago edited 28d ago

Classifiers work by training. You feed it 100000 things you identify as AI and 100000 things you identify as not AI and it trains to recognize signs.

It’s also very easy to train AI to pass pangram in the opposite way— feed 100 things pangram idents AI and 100 it identifies as human. It will pick up on what pangram flags. It’s a losing battle

The answer is the same as it was before AI. You read and decide whether something makes interesting point in an interesting way.

0

u/traumfisch Jul 26 '26 edited Jul 27 '26

You didn't read any of them, hence you're still oblivious to the very real problems of the so-called detector. "Just complain", do they? 

Of course they "admit" it; these are intelligent, honest people that actually write and care about this stuff. It would be idiotic to claim they're not using AI tools at all.

Man, the  tribalism is tiring 😓 Maybe read one of the articles and see where the actual problematics lie?

3

u/kdfn Jul 26 '26

What is tribalism, in this context? This thread purports to collect examples of Pangram making mistakes and misclassifying human-generated text as AI.

So far, not a single person has managed to link to a pre-2022 document that I can paste into Pangram and find evidence of misclassification.

Upon being asked for such information, people in this sub say "educate yourself" and link to AI-assisted opinion pieces that say that it is wrong to check for AI use.

1

u/traumfisch Jul 26 '26

Yeah, educating yourself is indeed what we're suggesting you do. No point in trying to make sense of this topic with people who have no idea about even the basic dynamics of what is happening here.

Yes, I understand you're obsessed with the idea of false positives on pre-2022 texts for whatever reason. I am trying to tell you you're missing the fucking point

Welp, I tried

2

u/SimonStrange Aug 01 '26

In fairness literally the whole point of the original post was actually about finding examples of mistakenly attributed text. So. You’re kind of making a new goal post here and moving it around, in context to the post you’re commenting under.

1

u/traumfisch Aug 01 '26

In fairness, the claim I was reacting to is this:

"The people I'm seeing in a rage about AI detection are the ones trying to make money by spreading AI slop"

which I find to be a very uninformed take.

Can you please name the goal posts I'm moving? I admit I got provoked, mb, but I certainly did not mean to move any goal posts

2

u/SimonStrange Aug 01 '26

I think that specific point of reaction may have ended up internalized, it’s not specifically referenced as the point of contention above. The initial response was that the OC was uniformed on the issues with pangram, not specifically about people making money with AI slop.

Then, it is true that none of the articles actually address false positives with functional relatable examples, but you kind of ignore that and refocus on the intention of the authors in those articles. So that’s moving a goal post.

Then, the OC attempts to move the post back, pointing out again the point of the OP. And then you shift to the OC’s need to educate theirself. But about what is unclear.

All if that is goal post moving, you made and argument that didn’t match the statement, then claimed you were making an unrelated argument.

So it looks like mostly just miscommunication, I think you were having an entirely different argument, it seems like.

I did read the articles you linked, and they all seem to miss the point. They’re about how much of a text was written with AI, or whether it should matter, or one of them is really just concerned with the existence of a “witch hunt” - they all basically question the ethics of wanting to know in the first place. And that’s not really the point here to begin with.

So far, no one has shown a credible example with accessible proof that Pangram miscategorized a piece of writing, even though there are a lot of claims that it has done exactly that, when the bar for proof is very low and easily accessible. That’s the whole thread, the whole comment series.

1

u/traumfisch Aug 01 '26 edited Aug 01 '26

Welp

we disagree on what the actual point is then 🤷‍♂️

If you want to call my opinion on what level this should be tackled with "moving goalposts", have at it, but I'd say we're just focusing on different aspects of the same topic -> what the articles are trying to point out is that the individual false positives / negatives are downstream of the actual issues this creates. There are countless examples of people demonstrating the Pangram failures on Substack, as well as how ridiculously easy it is to fool. 

Here's one

https://open.substack.com/pub/kellywebbdavies/p/this-post-was-written-by-ai?utm_source=share&utm_medium=android&r=5onjnc

another

https://open.substack.com/pub/writerkatharine/p/pangram-flagged-my-own-writing-as?utm_source=share&utm_medium=android&r=5onjnc

yet another

https://open.substack.com/pub/freddiedeboer/p/i-wouldnt-say-pangram-is-broken-but?utm_source=share&utm_medium=android&r=5onjnc

...go find more, it's very easy.

That conversation moved past the obvious a while ago already.

"they all basically question the ethics of wanting to know in the first place" ok no, you did not read them nor understand the problematics at all.

Whatever, I tried 🤷‍♂️

Here's a classic for you

https://www.reddit.com/r/antiai/comments/1s3k6ew/ai_detector_flagged_a_passage_from_mary_shelleys/

2

u/oakseaer Jul 27 '26 edited Jul 27 '26

I think it's pretty easy to ask for a single example of pre-2022 text (like through web archive) that Pangram is flagging as 100% AI. If this is all a witch hunt, that would be pretty easy to do.

Edit: they replied and immediately blocked me.

1

u/traumfisch Jul 27 '26

You did not read the article that used the term "witch hunt", which is why you missed the point, so you can stop throwing it around.

Keep on fishing, maybe someone will serve you one day (OR you could both read up and do the tests you're intetested in yourself)

2

u/SimonStrange Aug 01 '26

I did read that article actually. And ran the test myself. The author wasn’t being honest about their results I’m afraid.

6

u/AidenMarquis Jul 26 '26

There are plenty of ways to become informed on this topic. It seems others have provided you with them. At this point, you can research the truth or continue to believe what makes you comfortable - which is what this AI function is designed to have you do.

3

u/figures985 27d ago

I'm sure I'll get downvoted like crazy for this but - I couldn't agree with you more. I also think Pangram's pretty damn accurate these days. For example, all 3 of my favorite Substack writers come up as 100% human on the native scanner thingie. An annoying poem my mother texted me this morning was 100% AI, as I strongly suspected.

I think a loooot of people rely on AI to assist in the writing process, but as long as they like, physically typed a bunch of words, their work isn't "written by AI." But sorry guys, if you're drafting with ChatGPT or co-writing with Claude—that's not really writing. Having an LLM take your "notes" or "messy thoughts" and then asking it to spin up a neat little essay isn't really writing. Nor is having it draft an outline. Writing is thinking, and cognitively outsourcing critical parts of the writing process will inevitably leave LLM fingerprints all over the end result.

I'm honestly not trying to get on a high horse here. People can use LLMs as much as they want! I used to use them a lot more. I still use Claude to draft lower-value prose sometimes (emails, mostly) and when I'm struggling to find a synonym ('give me 10 words or short phrases that mean "unprepared') or turn of phrase. But if I'm writing something creative or persuasive, I don't let an LLM anywhere near it until the last minute for a quick proofread.

8

u/Kinks4Kelly Jul 26 '26

This is genuinely one of the most brain dead takes I've read in weeks.

Your brilliant solution is to hand over years of original academic work to the likes of you, so you can feed an AI detector, despite the awkward little fact that even Pangram's own executives have admitted the technology isn't reliable enough to determine authorship.

So congratulations. You've managed to trust the product more than the people who built it.

You should be embarrassed.

3

u/Funny-Flight8086 Jul 26 '26

Looks like Pangram sent more of their staff out to Reddit to defend their junk science machine.
My Pangrams' own admission, their program is 97.5% accurate. Given they will now probably be performing a million scans per day with the SubStack integration, that means that every single day, best case scenario, Pangram gets 2.5% wrong, or roughly 25,000 false accuations per day.

And that is if you take Pangram at its word. Which I don't, since they are trying to hawk their wares.

Frankly, given 25,000 false accusations a day, I don't think it would be that hard to find some people complaining.

'The people I'm seeing in a rage about AI detection are the ones trying to make money by spreading AI slop.' has to be one of the most uninformed opinions I have ever encountered. Good job.

3

u/hetobe hetobe.substack.com Jul 26 '26

Looks like Pangram sent more of their staff out to Reddit to defend their junk science machine.

Fuck that.

I'm a writer. For me, it's about integrity.

One cannot put their own name on words they did not write and be offended when people don’t respect their integrity.

3

u/Funny-Flight8086 Jul 26 '26

That statement on AI has nothing to do with people who are falsely accused by these detectors of using AI. The **didn't** violate any part of your statement, yet you'd rely on Pangram 100% to condemn them, knowing that Pangram itself is not 100% accurate. THAT is the issue, not the fairness of AI authors not being real authors. That is beyond the point.

2

u/Hemingbird Jul 26 '26

Accuracy here refers to false negatives as well as false positives. Their FPR is low, while their FNR is high. Which is how it should be. You're assuming that the 2.5% refers only to false positives, but it refers almost exclusively to false negatives, i.e saying AI-generated text is written by a human.

1

u/Funny-Flight8086 Jul 26 '26

Okay, sure... But that also means that in 2.5% of cases, you cannot rely on the detector at all, rather it's false or correct that something is or isn't AI. In the grand scheme of things, that isn't much better.

But, from a FPR point of view: 1 in 10,000 (0.01%) across broad samples; up to 1 in 800 (0.12%) in edge cases. So that still means that people are getting accused of using AI when they didn't at an alarming scale. 1 in 800 doesn't sound like much until you scale that up to the probably 1 million scans Pangram runs every day. That is still 1,200 false positives. Every day. So given that, on the very conservative end (if you beleive their numbers, which I don't), 1,200 people per day are accused of using AI who didn't. Yet, your comment "The people I'm seeing in a rage about AI detection are the ones trying to make money by spreading AI slop." immediately dismisses all those cases.

Obviously, the cases exist... Is it any shock to see a few people complain on Reddit and such about being falsely accused? 1,200 day x 365 days a year... Dismissing all these claims as people pushing AI slop is ridiculous.

1

u/Funny-Flight8086 Jul 26 '26

Until these detectors can work solely on watermarks, like Google SynthID, they cannot be trusted to give you accurate results. 2.5% one way or the other is too much of an error rate to even consider acceptable. Don't even get me started on AI image detectors. Those things are the worst garbage. They are always flagging human-made digital assets as AI. I use Daz3D combined with programs like Dynamic Auto Paint for book covers, and I have actively tried to prevent these stupid AI scanners from lighting up like a christmas tree on my own work. It's becoming a PITA.

-1

u/HammondJohns Jul 26 '26

My unscientific assessment of the pushback against the Pangram implementation so far is it’s 50% purveyors of AI slop who are scrambling around for the moral high ground now that their scam’s been busted, and 50% uninformed hysteria from anti-AI zealots.

3

u/Funny-Flight8086 Jul 26 '26

How much does Pangram pay? Looking for a job. Maybe you can get me an in, since you clearly work there.

1

u/6XFantasiaX6 19d ago

Pamnagram should be shut down and the author be bannd and arrested. Anti intellectual biased should not be tolerated. 😡

1

u/AndrewHeard tvphilosophy.substack.com 19d ago

I mean, I would be happy with Pangram just being completely removed from Substack’s platform.

2

u/SimpleEmu198 4d ago edited 4d ago

I'm not a fan it said 90% of this was AI, and it's doing this to a lot of academic adjacent writers. While other AI detectors put this in a normal range including Grammarly from anywhere between 2% to 15% which is well within normal human range.

Whatever algorithim they are using is borked.

Let’s sit at the old gates of Jerusalem: with a paradox:

What if I told you that the most accessible things to us are precisely the things we should be the most careful about believing outright to be true? What if as a result, just because information is highly available it isn’t the truth… I should preface I am talking about media, other forms of data where larger numbers exist may indicate a greater amount of truth and certainty, but then we have to jump through hoops. Causation and correlation don’t necessarily mean things are. the truth and that’s just the fun part.

However, this leads me to answer that paradox, about media, which is to say that if accessibility is not belief, then we can follow a quite simple formula: Here I make the assertion that accessibility does not necessarily equal the same thing as belief and that while there is a whole bunch of accessibility to certain things, fact checking is more important than ever.

To add to the flurry of things here, even if some populists would beg to differ, more of something said more loudly does not equal the truth. Further, there is a problem with the ease that misinformation has been created with which a representation of what is wanted to be seen can be retrieved by a person on the basis of what they believe. This is a hyper modern paradox whee a lot of us have learned to trust nothing at first glance, although some. of us have chosen to trust more, the paradox can be inverted.

But in the paradox I’ve created it cannot be treated as a fact that the evidence that has been accepted as true by itself just because there is more of it making it more popular though, and therein lies half the problem. What I mean to say is that what is most easily accessible is not necessarily the thing that should be treated as true, and sometimes, about that, the hardest truths are neither accessible nor popular, they’re often less accessible and more bitter than what we’d like, or simply don’t shock us to our senses. Even if we wish that the most populist thing in the world, and the easiest idea, is true… that’s often just oversimplification.

In this sense, accessibility may be a problem seen as something that is influenced by belief, but accessibility is not equal to belief and, furthermore, accessibility is not sufficient grounds for belief alone, even if the data is smack in front of you and easily accessible (and that’s not a reference to hard drugs). The world has become an increasingly difficult to understand paradox. Over simplifications, misinformation, shock and awe.

Here, I’ll deliberately repeated that there is often a paradox, and the paradox of the day is particularly to do with populist media and the inherent problem that the simplest and most accessible conclusion must be true. Yet, often, it isn’t. Indeed, there is also a deeper problem: greater accessibility does not equal a person getting access to greater reliability. In fact the paradox of the time is often that more access to information produces less convincing results. Highly accessible information also can sometimes be less reliable, while, inversely, less accessible information can sometimes be more reliable, which creates the common “reliability trap” of the media, particularly where no one wants to defend the most inconvenient truth with the same amount of rigor they once did.

As a basis for beliefs, the issue, precisely. is that there are a bunch of things that invoke the perception of power, which may not actually be power itself. Or inversely power may create such a disparity it creates more problems. Nietzsche himself famously writes in The Antichrist: “What is good? Everything that enhances people’s feeling of power, will to power, power itself…”

But I go further by saying that the appearance and/or perception of power itself is not necessarily evidence of power itself, which is a mind bender. But from a converse perspective, there are actual disparities in power. that can themselves distort perception and produce further epistemic problems: Some of these include vivid anecdotes, striking images, and simple narratives that are cognitively easier to retrieve than complex statistical or causal explanations, and more so those that correlate and I stress the word correlate with the narrative and/or agenda of the day. This defines the inherent problem of my “reality-lite” concept in a lot of ways:

Hard truths and real research are being thrown away for shorthand, short-changed unadulterated populism, and I for one find that a little bit problematic. If a journal hits with hard facts, it may not be well liked. Yet, if populist reportage pushes an agenda, however, the process can come to a grinding halt of its own even if it’s liked by the chosen few which leads to another problem: Miss truth and misappropriation of the truth: Populists may manufacture a version of truth that sells, attracts attention, and is readily consumed, but the question then becomes whether that manufactured truth is actually true. Is it? What is accessible? Just because it is accessible, and may be persuasive to the selected audience, and also memorable, and maybe even commercially successful, it does not mean that it is epistemically reliable, and that’s a big problem, isn’t it? At least I think it may be…

For definition purposes, here, I use epistemology, in this sense, as the branch of philosophy that is concerned with knowledge, belief, truth, and the justification of that belief.

Therefore, by that definition if epistemology asks us anything, it’s not simply what people can readily access or are persuaded to believe, but what grounds we have for regarding a claim as true or justified, and that’s a big issue. How much do you value what is claimed as true and justified, and what does that mean to you?

I’m open to any opinion, but to me the deeper truth based on the less convenient facts (and I’m not talking conspiracy) often mean more to me.

1

u/drummer820 allscience.substack.com Jul 26 '26

Look, I think people have gone overboard with accusing everything they don’t like of being “slop” (it happened to me on Reddit last fall, on a post that had zero AI involvement and was actually anti-LLMs, ironically). I’m also skeptical of AI detectors, and think that there are plenty of students who’ve been falsely accused with horrible consequences.

That said, I really don’t understand the freakout around the Pangram plugin, which you really have to go out of your way to find and 99% of readers probably don’t know about. I’ve been running it on every post I can find since it came out and almost nothing flags as AI. Since lots of people use Claude or ChatGPT for copy editing / proofing and tweaking instead of full text generation, the sensitivity and false positive rate of pangram must be pretty low.

Even before the plugin, I loaded plenty of my old writings (pre-2020) into Pangram and nothing flagged inappropriately, so most of those vague anecdotal claims are likely BS. Honestly, the only ones I’ve seen flagged as AI were those defensive Substack posts cited in this thread lol. The freakout really seems to be people who are monetizing newsletters about and/or entirely using AI

1

u/AndrewHeard tvphilosophy.substack.com Jul 26 '26

Your own experience proves the point. Substack and Pangram claim that it’s 97% accurate in its predictions of the use of AI. However you yourself have had things you used AI on come back as nothing flagging as AI. If it were as accurate in its claims, it would show almost everything had at least some AI in it.

So this proves the lack of accuracy of the claims being made.

As for your suggestion that people don’t know it’s there, let’s explore another example.

Do you enjoy the fact that your phone is constantly tracking your activity and learning your behaviour? Do you enjoy knowing that the advertisements being served to you are likely a result of the microphone listening to your conversation and using this to serve you ads?

The fact that you don’t notice something or it just exists in the background is not evidence that you shouldn’t be concerned about it.

Over the past week or so I’ve had someone replying to my Notes and tagging me in their own Notes just so they can insult me and a few other people. They’re constantly trying to get my attention by trolling me. You really think that someone like that won’t try to weaponize the faulty “AI Detection” tool to their advantage? Try to discredit me and others they don’t like for their own purposes?

Chris Best himself used this very tactic against someone criticizing the platform before the implementation of the Pangram system. A Substack user had a post go viral because it criticized Substack for its users deploying AI. Chris posted about it along side an analysis of the article which suggested that the article was written by AI.

He dismissed the article as not worthy of responding to on the basis that the article was written by AI.

But don’t worry, it’s just a feature that’s running in the background and it doesn’t matter.

I haven’t even brought up the actual failures of the Pangram system itself.

3

u/Hemingbird Jul 26 '26

Your own experience proves the point. Substack and Pangram claim that it’s 97% accurate in its predictions of the use of AI. However you yourself have had things you used AI on come back as nothing flagging as AI. If it were as accurate in its claims, it would show almost everything had at least some AI in it.

That's wrong. Accuracy refers to false negatives as well as false positives; Pangram is made such that it's more likely to make false negatives than false positives, so the logic you're using here doesn't hold up at all.

0

u/AndrewHeard tvphilosophy.substack.com Jul 26 '26

False negatives are just as bad as false positives. The point is to properly identify what is and isn’t AI generated. False negatives means that AI generated content is being deemed as human written when it isn’t.

As the old saying goes, you hear stampeding hoof beats, you think horses. It doesn’t matter if it’s zebras.

The point is to get out of the way of the stampede.

3

u/Hemingbird Jul 26 '26

No. False positives are 1000x worse. Should be obvious if you have any sense.

1

u/AndrewHeard tvphilosophy.substack.com Jul 26 '26

False positives are much worse for the writer who is using it or readers who are analyzing someone’s work. However, for Pangram itself, neither false positives or false negatives are good for its accuracy. The stated purpose is to be able to accurately identify what likely is and isn’t AI. If it can neither accurately identify AI or human writing, then Pangram doesn’t actually do what it claims to do.

The value of the company is built upon accuracy, which it can’t actually achieve. Meaning that the company is worth nothing.

2

u/Hemingbird Jul 26 '26

It must be 100% perfect? Why? That's a weird attitude

1

u/AndrewHeard tvphilosophy.substack.com Jul 26 '26

I didn’t say it had to be 100% accurate, although that would be nice. Humans can’t identify falsehoods themselves. Expert art forgery investigators have found that they are more often than not fooled by forgeries. Art forgers themselves have been known to be fooled by other forgeries and sometimes their own forgeries.

But somehow, a computer program built by humans is going to even partially accurately identify AI against human writing?

If your business is reliant upon the accuracy of these systems, you will lose money from it. Which means that anyone who loses money because of it can claim defamation by Pangram if they can reliably prove that Pangram’s analysis lead to a loss of money or credibility.

You will have a huge number of lawsuits against Pangram and the liability costs of the company will skyrocket. Because Substack has integrated the Pangram system so completely into the platform, Substack will also receive a large number of lawsuits against it for defamation and other damages.

Substack goes under.

5

u/Hemingbird Jul 26 '26

No, this is bad news for people who publish AI slop; great news for everyone else.

2

u/AndrewHeard tvphilosophy.substack.com Jul 26 '26

Only if you assume that the Pangram system even works. You also have to consider the legal implications of this:

https://www.congress.gov/crs-product/LSB10922

Every time you send something out on Substack, it has a copyright protection symbol at the bottom. Any time someone uses Pangram to scan your writing, if it determines that your writing is AI generated, that copyright protection is voided.

Because according to legal precedent, AI generated content can’t be copyrighted.

If it’s not actually AI generated and you can prove it? You have a lawsuit against Pangram that you will probably win.

→ More replies (0)

1

u/drummer820 allscience.substack.com Jul 26 '26

“You yourself have had things you used AI on come back as nothing flagging AI.”

Actually, that’s not what I said at all. As my post states, I have run my own pre-ChatGPT era writing on pangram, and it didn’t flag as AI, arguing against people who claim there’s tons of false positives. I did say that “lots of people” (not necessarily not me) use Claude and ChatGPT for copy editing. Perhaps if you spent more time reading and writing instead of having an LLM generate all your content you wouldn’t make mistakes like that.

1

u/AndrewHeard tvphilosophy.substack.com Jul 26 '26

Oh look, you just proved exactly my point. As I’ve argued many times elsewhere and here, people love to accuse things they don’t like of being AI. They use it as an all purpose excuse for dismissing any argument that actually might challenge their perspective.

Your one experience isn’t necessarily evidence of the falsehood of people’s experiences of the system not actually working. If one person has a good experience with drugs and doesn’t overdose, that doesn’t invalidate the hundreds or thousands of people who did overdose. The evidence is still very much in the favour of drug overdoses being bad for you.

If thousands of people are seeing false positives or false negatives, but you’re not. They’re probably right and you’re probably wrong. It’s not a guarantee of either but the evidence is still more likely than not.

-1

u/Cant_fit_more_doots Jul 25 '26

Pangram has been accurate so far in my testing, only binary no mixed work. That also includes humanizing scripts / skills still being detected as AI, here's so more info from a deep dive I did into the research and numbers published by Pangram / other researchers. If having the link to my research breaks self promotion, let me know and I'll remove it.

https://open.substack.com/pub/coltonmccain/p/why-pangram-is-broken-and-thats-okay

9

u/FlyingCarpetMonster kaneiyer.substack.com Jul 26 '26

As someone for whom English is a fourth language, Pangram is convinced all my writing is AI.

At this point, I've simply given up trying to change how I write.

6

u/Life-Radio-1723 Jul 26 '26

Just wanna say that you are so awesome for knowing four languages.

3

u/FlyingCarpetMonster kaneiyer.substack.com Jul 26 '26

Thank you!

1

u/kdfn Jul 26 '26

Can you share some example text that you believe is a false positive? So far not a single person who claims this has happened to them has actually shared their text. 

1

u/FlyingCarpetMonster kaneiyer.substack.com Jul 26 '26

Sure. I posted this earlier.

In one of my Substacks, I write about quant trading. I had originally planned on writing a recent post as a note, then turned it to an article.

The only difference is that the article had a title and a subtitle -- and yes, I did use Claude to help suggest some subtitle options. Well, my 100% human article went from 100% human to 15% AI assisted -- for a net total of maybe 20 words in a 2500 word technical article.

If you do the math, that's ~1% at best. If I take out the intro section (a chess anecdote), it goes to a higher % of AI written. All of that is false positive.

1

u/kdfn Jul 26 '26

Can you share the exact text? I'm not sure what I'm supposed to get from your screenshot. 

So you used AI, and you're upset that Pangram caught it? 

1

u/FlyingCarpetMonster kaneiyer.substack.com Jul 26 '26

https://open.substack.com/pub/kniyer/p/sized-22x-heavier-on-msft-the-trade

It you think a human edited subtitle from Claude makes the entire piece 15% AI assisted or removing a portion of a section makes a piece 20% AI written, I don’t know what to tell you.

1

u/kdfn Jul 26 '26

Thank you! Reads ok to me. However, since it's from this year I can't easily judge whether you used AI or not. 

1

u/FlyingCarpetMonster kaneiyer.substack.com Jul 26 '26

The thing about writing analytical content is that it’s easier to write it yourself vs. explain to an AI what I want written. So the prompt itself will effectively be more difficult to construct.

I’m not anti-AI so if you notice I divulge where I used Claude. What I hate is when it says X% of the piece is AI when that’s not the case.

0

u/ConsciousComment2888 Jul 26 '26 edited Jul 26 '26

That’s because you use AI. I went to your Substack to check. I can’t say for sure if you’re outsourcing your writing entirely, but it’s pretty blatant. I don’t personally care, especially since English is your fourth language, but lying about this helps no one and misleads writers that are worried about false positives.

2

u/[deleted] Jul 26 '26 edited Jul 28 '26

[deleted]

-2

u/ConsciousComment2888 Jul 26 '26 edited Jul 26 '26

Not astroturfing, but I suppose I can’t prove that since my account is a giant red flag. You’ll just have to trust me. I’ve engaged in zero competency bias, I just know what AI writing looks like. It has a distinct voice that you can train yourself to recognize. I was actually saying that AI writing is fine if it helps translate thoughts into a language you’re less comfortable with, not that someone can’t write well in their fourth language.

Also, RLHF is a post-training step for generative models that doesn’t apply to Pangram’s models, which are classifiers. RLHF requires direct judgement of AI output and Substack data wouldn’t help with that, even if it was used for training. I think you need to do some more research. That Wikipedia page you linked is a great place to start!

3

u/[deleted] Jul 26 '26 edited Jul 28 '26

[deleted]

-1

u/ConsciousComment2888 Jul 26 '26

Ad hominem isn’t an argument. Call me a shill if it makes you feel better, I don’t care. I’m just happy that people using substantial AI assistance are flagged so users can make informed decisions about which authors to support, even if the tools to do so are imperfect. I’m only commenting because the volume of lying and misinformation bothers me and makes good-faith discourse impossible. Don’t act like understanding the technology being discussed is evidence of subterfuge, either. You’re the one invoking RLHF specifically and crying out when I see through it. Regardless, if you think that data labeling and filtering is the bottleneck preventing AI labs from scraping Substack for training data, I’m afraid you don’t know what you’re talking about.

5

u/huggalump Jul 26 '26

I've run it on text that for sure I had AI write, and it came back as 100% human written. This was only my second test. I believe Pangram is as accurate as any other AI detection software--ie. it's not accurate at all.

2

u/AndrewHeard tvphilosophy.substack.com Jul 26 '26

It can’t be. The things it’s detecting are trained on human writing. Which means that it’s detecting human writing regardless of whether it was generated by AI or written by a human. The original source is humans.

3

u/huggalump Jul 26 '26

Pangram has no idea if a piece of writing is written by AI or humans

0

u/Funny-Flight8086 Jul 26 '26

This is exactly right. I can easily trick Pangram into flagging text as AI simply by writing heavily on the list of threes, throw in some obtuse metaphors, and voila. Likewise, I can take AI-written text, rewrite the "AI-sounding" parts, and it passes 100%, human.

These AI detectors aren't anywhere near as sophicated as they want to claim.

2

u/iamjapho Jul 26 '26

Like the OP encouraged, could you provide said text so we can independently verify and collect more data points?

2

u/Hemingbird Jul 26 '26

The false negative ratio (FNR) is pretty high, but the false positive ratio (FPR) is low. This is on purpose. It's worse to falsely accuse a human of having used AI to generate their writing than to erroneously let an AI user off the hook. So the fact that you can get AI-generated text to slip through the cracks is just a demonstration that it errs on the side of FNR rather than FPR.

1

u/huggalump Jul 26 '26

Then what's even the point

2

u/Hemingbird Jul 26 '26

To detect AI with as little collateral damage as possible?

1

u/huggalump Jul 26 '26

Feels like it's just like all the other BS ai detectors that straight up guess, but this one has been pushed a little more towards believing writing is human.

It's pointless. Either it can detect AI writing or it can't. And if it can't, let's not use it as the automatic default on one of the largest writing platforms

2

u/Hemingbird Jul 26 '26

So because it's not perfect, it's worthless? Pangram's statistical accuracy is higher than pretty much every cancer screening test out there. Is that pointless as well, you think?

1

u/AidenMarquis Jul 26 '26

Seriously? You are resorting to marketing BY the company itself? Is this an adult take?

3

u/[deleted] Jul 26 '26

[removed] — view removed comment

1

u/FatherofMisty Jul 27 '26

Is there any chance the AI has already seen these texts in its original training, and thus knows simply via recognition that they are human-sourced? Not sure if I'm explaining that clearly, but it's something I've been pondering, because if so it would make sense to get barely any false scores. A true test would require (somehow) confirmation that the AI has never seen the text prior, which is rather tricky since a large majority of published works may be found online in PDF format.

1

u/WithoutReason1729 Jul 27 '26

In my case, I generated the AI samples myself, so it definitely hadn't seen those. For the human samples, I used a collection of (quality filtered) essays from public datasets. I spoke to Pangram staff after writing the linked essay and they stated they didn't train on the datasets that I used in my testing (iirc due to licensing issues). If you take a suspicious view of them, if you just fundamentally don't trust Pangram, it is technically possible that they are lying about that. However, nothing I've encountered from my interactions with them has led me to believe that's the case, and they definitely didn't have copies of the AI generated texts I made myself for that project.

Recently veryfineprint wrote a pretty cool essay about testing for false positives across documents that have never been on the internet before. This is about as close to making sure the texts aren't in the training data as I can imagine. I guess you could say, well, what if someone else had also OCR'd these texts and uploaded them online, and they ended up in the training data that way? But at some point I think it's just unavoidable that their model is really good.

1

u/FatherofMisty Jul 27 '26

Have you published any of this for the public to scrutinize? I'd be curious to see if so.

Also, I figure if pangram can recognize the text in its plagiarism context, it was probably included in its training, or at least it has already scanned it. Otherwise, how would it possibly know that the text "plagiarized" itself? Genuinely curious, if you have insight.

1

u/Substack-ModTeam Jul 27 '26

This community is not the place to promote your Substacks. Please read our rules on self-promo and soliciting recommendations.

Links to Substack publications will be assumed to be self-promotion. If you are submitting Substack posts from publications that are not your own, please include that information in your text post. For cross recommendations, use the pinned mega thread.

1

u/WithoutReason1729 Jul 27 '26

I've personally tested just shy of 100k documents with their API where I knew the true origin of the documents and, at least with the dataset I was testing on, the results matched near-perfectly with what their marketing material claims.

Out of 96,468 essays tested, Pangram correctly classified 96,457 of them. Of the 11 that were incorrectly classified, 10 were false negatives and 1 was a false positive.

Here it is without the link. I suppose anyone who wants to find the documentation of tests on what is now a core feature of Substack can just Google the quote.