I wouldn't say Pangram is broken, but I would say that it's brittle
Posted by antigizmo 4 days ago
Comments
Comment by timpera 1 day ago
It's especially bad that they keep insisting that it works very well, because thousands of people will probably end up falsely accused of AI usage as a result.
Comment by skippyfish 1 day ago
I'm not saying that's you, but since a proof is easy to produce, it would be nice if you could share.
Comment by amrit3128 1 day ago
Comment by aleph_minus_one 1 day ago
> I'm Kenyan. I Don't Write Like ChatGPT. ChatGPT Writes Like Me.
> https://marcusolang.substack.com/p/im-kenyan-i-dont-write-li...
HN discussion:
Comment by coldtea 1 day ago
It's not command of the english language it flags (for native vs non native to matter), it's stylistic mannerisms, some of which happen to also exist in some direct translation of some foreign patterns for some languages.
But it detects them in a crude manner, missing obvious AI slop, and misflagging human texts.
Comment by ano-ther 1 day ago
But I don’t have a license to confirm.
Comment by coldtea 1 day ago
How would you know if you just take whatever slop verdict is serves as correct?
Comment by phoghed 1 day ago
Comment by skippyfish 1 day ago
Comment by embedding-shape 1 day ago
Comment by Ariarule 1 day ago
Do people not recognize names nor search for them before making insinuations? It's not like OP is a brand-new pseudonymous essay on a default-template blog with one or two other posts in the history: https://en.wikipedia.org/wiki/Fredrik_deBoer
Comment by skippyfish 1 day ago
Comment by zahlman 1 day ago
> Be kind. Don't be snarky. Converse curiously; don't cross-examine. Edit out swipes.
> When disagreeing, please reply to the argument instead of calling names. "That is idiotic; 1 + 1 is 2, not 3" can be shortened to "1 + 1 is 2, not 3."
> Please don't comment on whether someone read an article. "Did you even read the article? It mentions that" can be shortened to "The article mentions that".
Comment by phoghed 9 hours ago
Comment by Wowfunhappy 1 day ago
Pangram should refuse to work with small amounts of text. (Well, I know it already does, but the threshold should clearly be significantly higher.)
Comment by ThrowawayTestr 1 day ago
Comment by johnfn 1 day ago
Comment by btown 1 day ago
And even this isn't perfect, nor is it guaranteed that enterprise and individual accounts are tuned the same way. So students also need to proactively use audit/keystroke logging systems to protect themselves against accusations, which creates a type of "panopticon" on one's early/ephemeral drafts, including language of frustration (who among us hasn't typed curses into an unsaved draft at some point?), that can massively stifle creative thought. And if an institution provides such a tool, their centralized access simply worsens the "panopticon" characteristics.
There's no easy solution, here, sadly.
Comment by runako 1 day ago
LLM-generated text does not carry a watermark or other identifying marks. The "theory" is that an LLM trained on human writing, to mimic human writing, can be distinguished from actual human writing in under 100 words.
Notably the first diagram on the research overview page (https://www.pangram.com/research/how-it-works) shows feedback for "misclassified human examples." This is a category error; Pangram will not find out when it has misclassified text in the wild, except in rare cases. Only the "licensed human-written text" in its training data can be used as feedback.
Scams like Pangram also cause real harms, mostly because laypeople do not understand that what is being offered is not possible. Pangram advertises 99.98% accuracy, and they pitch it as a tool for teachers and universities. Translated: if a college like University of Alabama rolled this out, you could expect ~40 students to have their lives upended by this snake oil, every year. (And how can one even prove that an allegation is false, that they did write a given text?) And this is the best case, using the number on Pangram's homepage.
Comment by Smaug123 15 hours ago
Comment by runako 7 hours ago
It cannot. this is a misconception. A human being is fully capable of writing 200 words that Claude will identify as not being written by them, because human beings are far more complex than Claude. A human can even choose to deliberately write in the style of a different human, even one who does not exist.
Sometimes people are writing instruction manuals; those are not written like their professional emails, which are not written like their personal emails. It is normal for people to be able to write in different voices/styles/etc. People code switch, people write for different audiences, people change over time, people are hurried or tired or sick, etc.
So no, an LLM cannot identify you uniquely in 200 words. But more to the point, most human communication is not in training sets. And Pangram has no way of course-correcting on the vast amount of data that is not in its training sets.
By comparison: the autonomous vehicle companies actually do need their products to verifiably work. So they also feed back human-analyzed data from real trips into their models. They can tell the model where it was right or wrong in the real world. This is the part Pangram cannot do! Pangram deployed at a university may be used to accuse a student of cheating, but then Pangram will never know for sure whether the text in question was written by a human or machine. The feedback loop is missing a critical step!
Comment by Catloafdev 1 day ago
Comment by supermatt 12 hours ago
By analysing a specific piece of prose you probably arent using the same chunk as would otherwise be analysed and can end up with a different result.
Comment by achileas 1 day ago
Comment by pixl97 1 day ago
Comment by sscaryterry 1 day ago
Comment by embedding-shape 1 day ago
Holy strawman-batman, not only does the founder of Pangram not have a proper response to the actual criticisms, he feels the need to completely make up very different situations to try to illustrate some completely different point... I guess good job of the founder to engage at all, as deBoer does bring up a lot of valid points and criticisms of why people really shouldn't rely on "tools" like Pangram, too bad the founder failed completely at addressing the more serious points, and instead just chose to say "Well, there will be false-positives, what can you do?".
Comment by cgio 1 day ago
Comment by ameliaquining 1 day ago
Comment by phoghed 1 day ago
Comment by johnfn 1 day ago
Comment by coldtea 1 day ago
And the obvious absolute test that would prove its shit - millions of books and posts written pre-AI would unfortunately be in its training already with some date associated, and thus it would not really detect them, it would just know "x text, written pre 2020".
Comment by 650 1 day ago
Comment by phoghed 1 day ago
Comment by vidarh 1 day ago
Comment by daflkfdslkfds 1 day ago