I also uploaded my PhD dissertation from 2012 and got back 40% machine, which is just barely below the 42% threshold.
I don't publish anymore... but does this mean I wrote like a LLM or did LLMs learn from me? :p
Update: I also uploaded a IEEE CLUSTERS paper I wrote in 2015 and it came back 74% machine written :|
The biggest results: in Jan of 2026 about 39% of papers got flagged as AI written. In computer science speicifcally the peak was at 65%. Mathematics barely moved away from 0.7%, though the proof heavy math texts might just not get picked up by the detector properly.
All this is a detector estimate of a statistical signal and not a proof any given author used AI. Machine written can also mean heavy AI-assisted editing.
I've also uploaded text samples from my own (unreleased) research from pre-LLM era, and it's seemingly scoring pretty high on the LLM-detection scores. On other papers, nearly every sentence is highlighted as red "machine-leaning," but that does not impact the score? Additionally, there are dramatic differences between the scores for identical text with and without LaTeX formatting, despite the fact that it should not matter.
The takeaway from this should be "it is difficult to detect generated text and we should be careful about accepting results simply because they confirm a hypothesis."
--
Relatedly, the text above scores as highly machine-written, despite the fact that I just wrote it with my human hands, I promise :)
When 65% of the papers you read have the characteristics of being AI written, whether or not you use AI to write, your writing will be influenced by the AI style. I imagine this must be particularly the case for newbie researchers who are still developing their writing style
It is not because we ask LLM to write us a paper about X. It is because we are bad at writing in a scientific style, and american editors expect us to do it. With LLMs, we can write in basic sentences and tell the LLM the idea and it converts that to nice writing.
If you write each paragraph and have an LLM make that paragraph more scientific, it is entirely your paper, but it is flagged as LLM generated. If you have an LLM write the paper but speak good enough English, you can make it look human even though it is not human.
Perhaps that’s what’s happening.
For example, LLMs love to talk about LLMs (and the people who write with LLMs love to write about LLMs). Could "large language model" itself therefore be flagged as an AI-like phrase by this approach? It didn't exist much in the literature before 2022, does now, and certainly does more in AI-generated text: but, it is not actually a great way to distinguish modern AI generated text from human written text.
A helpful control would be to show that on some cohort of papers that can be declared reasonably clean of LLM generated text post 2023 there are very low rates compared to the arxiv.
For example, while papers in the journals Nature and Science are unlikely to be entirely LLM free at this point, if those were tested through 2026, we should see a line significantly lower than the arxiv's growth.
The question is just how to organize these outputs and conclusions in a way that is consistently reproducible and also how to correct errors or remove LLM nonsense where it refuses to take a position on something.
Before it made sense to do this in papers but it feels like we need something like a paper format.. that is fully reproducible ideally and optimized for aggregating knowledge in a better way. I.e. before a person spent months on one of these and there was just more filtering, and the output itself was a clear signal of time spent and effort that no longer exists.
but this post makes me wonder, if more papers' are written with AI, or the shape of knowledge of converging?
I don't want to read slop generated by AI. AI written articles are generally low effort.
So what?
Genuinely. I get that there may be some visceral reaction against this, but when I break it down, I mostly fail to see the problem. Seems like what is actually important is:
Compared to before, when a human reads it, do they -- or society -- get something good out of it? Is it worth it to add this to the "pantheon?"
If that's not what's happening enough, and if this doesn't describe the process -- then the problem lies elsewhere, no?
There really is no point, as long as you verify the content matches your intent and edit out anything poorly written.
Frankly, I've read plenty of papers by native English-speakers over the years that'd strongly benefit from being rewritten by an LLM too...
If we imagine a set of all human ideas that these models have access to, then the set of possible discoveries would be something like the superset of all possible combinations of those ideas. I think all LLM discoveries are bounded by that space.
Looking at the recent OpenAI math discoveries, that seems to be pretty much what happened. Existing ideas were used as building blocks, the model found a valuable combination, and the result was something new that had real value.
If fake papers weren't already a big problem before AI and the fields had already been policing themselves adequately, if this was already a functioning high-trust domain, maybe we could ignore this a bit more, but the fields already manifestly had problems. People taking advantage of that are reasonably more likely to use AI. The pressures to publish or perish provide the voltage and the AIs are a rather convenient path-to-ground.
I agree in some sense that if a truth is published, it doesn't matter if the AI or a human published it. However there are perfectly reasonable reasons to be concerned that AI usage is correlated to not publishing truths, especially in a world where merely being human-generated was already not an adequate check against that.
In other words, it "sounds smart" without necessarily having anything to back it up. In even more critical terms, it's very good at bullshitting.
Unfortunately for us, the scientific community current relies on a certain amount of trust. (To do otherwise is very expensive! see: bitcoin). When you introduce a known-bullshitter to write your papers, every human in the loop effectively has to defend against an adversarial attack. Not just the readers at home, or the peer reviewers, but even the author needs to be wary that the facts and arguments coming out of the LLM are true and meaningful.
Personally, I've been a minor contributor to several high-profile papers. I don't know how every field does it, but in my experience, the corresponding author (generally the PI or other senior scientist), is responsible for the accuracy of the paper. They ultimately have to trust the people who did the work that the facts are true. Introducing LLMs into the mix make it more difficult for them to identify and review sections they are unsure of. (An honest person will typically write at a confidence level reflecting their certainty. LLMs do not do this in any reliable way.)
I've also found that LLMs frequently use metaphors that are unhelpful, or used out of context in a field that isn't familiar with them. This makes understanding the text more work, for no good reason. Introducing terms or definitions with low relevance reads as impressive at first glance, but avoiding the standard terminology in the field just adds confusion. (As an analogy, imagine if you were reading a CS paper that, for no particular reason, devoted a section to a new data structure called an "akimbo tree," which after much untangling, you realized was a reinvention of a randomized splay tree.)
What about the papers that graduate to proper publication?
Arxiv is full of pre-prints that anyone can upload.
To me it sounds like 1. Either your tool is just not that good and reliable as you thought, 2. AI is trained on human written articles, so some of that human written content informed the now established “AI slop”.
There are people who shipped “slop” before AI.
Have you tried doing that or even read the article?
The article says that their detector flags 0.4% of pre-AI papers as AI-written.
If I paste the first page from this paper (https://www.fourmilab.ch/etexts/einstein/specrel/specrel.pdf) in https://unslop.run/app, I get a 0% chance that it was AI-written.
A possible conclusion for this could be: If the majority of CS papers is AI written, let's just accept this reality universally and stop worrying about it altogether.
Pre-LLMs, a paper with no spelling or grammar errors showed that somebody had put effort into writing and editing it. If they cared about the presentation, they probably also cared about the content. LLMs routinely produce nonsense that looks superficially like high-quality work.
There are far too many papers to read all of them. LLM slop is evidence that something is probably low quality. As the saying goes, "if you can't be bothered writing it, I can't be bothered reading." The rare outliers will get enough citations and recommendations to overcome this filter.
My problem therefore is: we are seeing more and more papers written with tools that are known to make up facts, citations, and even entire papers. And the number of papers has increased, too. I therefore see it less as "people are being more productive" and more "people are releasing bad science much faster than we can keep up with".
Sure, some people are artists - but most aren't.
The difficulty with this is then: How do you get a clean post 2023 dataset? I have no straightforward idea for this. You can't use other AI detectors to build it because then you'd never outperform them.
We were discussing research in general and I asked them: "Do you prefer the writing of the papers or the research?"
They, almost unanimously, agreed that they preferred the research. This makes sense as if they preferred writing they probably would have chosen another profession.
I say this b/c having LLMs available to turn research diagrams, code etc into a paper (or at least the starting point of a paper) will probably lead to MORE quality research papers. This is b/c I'm sure there was some friction in a researcher's mind of "I would love to do the research on this but don't want the trouble of writing the paper".
Put another way: on a 2D plot with one axis being the skills as a researcher and the other being hatred of writing, LLMs may "unlock" the people high on both axes to get more papers out.
Post Script: I agree that this could also lead to more BAD papers but the net may turn out to be positive in the long run.
FYI this is all relatively new so there might be lots of issues and iterations coming.
The funny thing is that "slop" was defined by the writing habits of AI model, which we have learned to pick upon and recognize.
The "It's not X, it's Y", the rhetorical questions and other patterns would have been the tools of a skilled writer, and those people writing "like AI" before AI most likely would have been recognized as such.
I don't think the problem is as bad as a naive reading of this article suggests. I'm highly skeptical that anywhere near 65% of recent CS papers that I've read (mostly systems papers) are substantially AI-written. I threw some recent papers I've read into the system and they come back as 0-7%.
Pot meet kettle?
One thing I see a lot is papers flagged as AI because they include llm rollouts in the paper as examples.
Its not clear from your writeup what threshold needs to be reached to be classified as "machine written". A preprint where half the text is human and half is 100% AI should be a different category than a preprint where 100% of the text is AI-assisted.
Also its cool that you're making the detector available. When you say "cheap to run", do you know how this compares to pricing for a commercial detector pangram or GPTZero?
I asked Codex to generate an article with a high score and then asked Codex to (reverse?) hill climb that score. The original generation scored 97% and then the optimized one scored 1%. Both are pretty bad and read like slop.
https://gist.github.com/wbew/8a2bd6686bf875210f2244ac8ea65bf...
consider the fraudster that went around suing people on the basis of his absurd claims of being bitcoin's creator. He's now transitioned to using AI to gather graduate degrees and is obtaining masters and doctoral degrees at a regular place and writing multiple 'papers' per day that are all quite obviously AI slop.
People report his cheating and publications and simply no one cares... and this is someone court adjudicated to have fabricated evidence in court on a massive scale, including through the use of AI.
But when it comes to the degrees and publication everyone involved that wanted paid got paid, and apparently that's all that matters.
Gotta be trolling :-D
This is a sin.
"More scientific" is not some merely stylistic thing that faithfully preserves the original meaning of what you wrote. The precise details of each paragraph matters a lot in terms of what and how it communicates. The fact that these details do matter means that, according to my accounting, it is not entirely your paper.
Also, I find striving for "scientific" to be a pretty undesirable thing. Why should papers read like that? What is the benefit? The best papers (in terms of their writing and communication) are unpretentious and conversational. I'm pretty sure I'd prefer your "basic sentences", especially if they were wholly yours. (I understand that there are also external forces at play here as you mentioned.)
It's hard to say if this code is structurally better or worse than before, but it's certainly voluminous and as far as leadership could ever tell, with their flawed metrics, that is all that matters. It will be years before we figure out if this is a good idea and worth the cognitive atrophy.
Anyone not using LLMs all day is just not going to be as prolific. I can't imagine that the same factors aren't at play in the scientific research community where it's all about how much you can publish.
A more interesting question, to me, is looking at a corpus of essays and analyzing how writing has changed with the introduction of LLMs. We can look at changes in vocabulary, linguistic features, style embeddings, regular embeddings, typos, errors, and references over time. When looked at in this way it is clear that academic writing has changed at the population level but what has led the change is harder to track down.
The article doesn't seem to mention consideration of AI for polishing human work.
I see plenty of anecdotal evidence that models have been trained fantastically well—and getting better—at writing to trigger the right neurons in the human population to produce “This is interesting/informative/correct” responses in bulk.
Could their ability to produce those responses run far ahead of their ability to actually achieve the last in reality? Sure seems plausible, and then where are we?
That’s a big if. We all know that’s not what’s happening.
Reviewers were also totally unengaged. Of 20 reviews I read (from my reviewers or from reviewers on the same papers), maybe 2 were mediocre, and the rest were crap (though likely not AI).
The notion that science will somehow benefit from this is about as stupid an idea as you can have. Science relies on skepticism. AIs are not skeptical, and many folks are submitting papers because they stand to gain something, not because they are motivated to do good research or develop new understanding. Fields are being inundated with garbage that is maximally indistinguishable from real work (that's the training objective for LLMs). This in turn maximizes the cost of identifying bad work.
This is the same enshittification process that we see everywhere else. You get spam phone calls because there is no reason for a spammer not to call you. "Researchers" are submitting spam papers because there is no cost to doing so with some possible gain. Absent intervention, this eventually drives the community value of the network to zero (or potentially negative, if friction costs to switching are high).
if 50% of the work is nonsense, then there's a serious concern that we can't move forward at all.
Perhaps the most darkly amusing consequence of this particular mania is that by poisoning the majority of our information environment with hallucinated slop, we have likely crippled the next several generations of machine-learning techniques before they're even invented! Small, locally-hostable LLMs will rattle along spewing spam long after the broader "genai bubble" pops, and building clean training datasets will permanently be more difficult and expensive.
I am almost certain he was "hallucinating" the results. This was in the 2010s
There are well known issues in academic publishing, though I imagine it has become much noisier like open source
That's a big if. ArXiv is not peer reviewed and LLMs basically interpolate and extrapolate text, which makes them essentially fluff generators. Even in the most charitable interpretation, LLMs enable those with nothing to say to say nothing while meeting surface-level style guides.
Other problems include: Signal to Noise Ratio going through the roof.
AI hallucinates and makes up stuff 100% percent of the time. Never been a fan of that word for this.
Again, I fail to see the problem here that isn't solved by careful reading WHICH IS WHAT PEOPLE SHOULD BE DOING ANYWAY. I would like to see room for AI disclosure, maybe a statement of "this is how much AI I used."
But this blanket X% of this looks like AI? Again, so what?
Yes, I feel like there's room to improve things, I just strongly doubt that "using AI to detect AI" is a particularly useful thing to do here.
Trust is very important to human progress.
Pangram agrees: https://www.pangram.com/history/3de33376-94e3-404d-bbb0-751a...
> The article doesn't seem to mention consideration of AI for polishing human work.
Because it isn't a consideration. You are what they are looking for.
This was my big fear before we saw price increases. Now I'm pinning all my hopes on AI being too expensive to justify further big corporate pushes. (Sigh.) I love having new tools, but I hate being pushed to use ______ tool to meet some managerial metric.
I can't speak for the entire scientific research community, but I can say that for basic biomedical research (not clinical research), LLMs are mostly ignored. They simply don't have the ability to parse most raw wet lab data.
This obviously has not stopped the creation of a legion of startups, new Asst Profs, etc, claiming that they are using "AI" to crack new frontiers. In practice, the most effective of these use machine learning, rather than LLMs.
Maybe for very short phrases, but otherwise I disagree. Phrasing very quickly runs into a combinatorial explosion. In the words of Noam Chomsky, "Virtually every sentence that a person utters or understands is a brand-new combination of words, appearing for the first time in the history of the universe."
In my opinion, the difficulty in LLM/human text discrimination isn't that a person might coincidentally write exactly the same text as an LLM would, but rather that 1) LLMs aren't hard locked to a single phrasing (so this is a tougher problem than matching to a single static document, e.g. plagiarism detection) and 2) text has relatively low information density, so you need quite a bit of it to gather enough data to run a statistical test with a reasonably narrow confidence interval.
That's a huge assumption, and one that goes against the whole notion of using LLMs to generate text. AI slop is by far the norm.
But I suspect that a lot of academic's feelings about it are informed by what others have told them and how they've been trained, rather than by what's actually permissible in the publishing system.
Huh?
But that's easy to tell?
It's better then incompetents code, worse then a motivated average dev... But good enough hence the real question is value aka time& money invested/quality.
That's much harder to tell, and I currently think opus/fable generated code is decent enough to be safely in favor - at least on subscription
Can you unpack that a bit? It produces measurably, meaningfully inferior code everywhere I see it in use.
> "leadership encourages this because from what they can tell, there is no downside"
As said 'leadership' I find this a bit puzzling. I'm seeing strong, quantifiable evidence of increasing churn, increasing incident count, and length of downtime from the date of our biggest push into GenAI, and I'm organizing efforts on my teams to mitigate those issues and actively reduce GenAI adoption.
If you mean my c-suite, you're mostly correct although they are already rumbling about seeing zero or negative ROI on GenAI investments.
> Anyone not using LLMs all day is just not going to be as prolific
Agreed, but prolific != productive.
Hah. I've been working with ChatGPT-5.6 Sol and a lot of guidance to port a particular geometry from 2800 lines of Python-scripted Rhino3D (basically a custom parametric CAD kernel in there, plus use of it) to Python-scripted OCCT/FreeCAD (an existing CAD kernel) and it's up to 36,000 lines and only about half way there. And I've been setting goals and targets for duplication minimization, code size reduction, etc. The results are fine for my purposes, but if I put a positive value on "voluminous" I'd be super impressed; with my actual negative value assignment on lines of code, this is more "absolute crap but still useful to me."
I think "absolute crap but still useful to me" is a pretty high value and not worth neglecting, but I'm definitely scared by leadership who sees the toilet overflowng and assumes that means that value has been maximized.
We scored the full text of 12,750 arXiv papers and found that about a third of new ones read as machine-written. Here is the method, the results, and an honest account of the limitations.

Share of new arXiv papers flagged as machine-written, 2021 to 2026, at a threshold calibrated so pre-ChatGPT papers flag at 0.4%. The eight slate points are the pre-LLM control months; the band is a bootstrap 95% interval.

Share of papers flagged as machine-written by field, over the 12 months to July 2026 (about 300 papers per field).
There is a genre of headline that says "N% of X is now AI," and most are not worth reading, because the detector behind the number also flags some share of genuine human writing. If a tool marks 40% of new papers as machine-written but also marks 20% of papers written before ChatGPT existed, the real story is the 20% nobody mentioned.
So we built the study around that objection. Our detector, described here, is calibrated for academic writing; at a 0.4% false-positive rate it clears 99.6% of genuine pre-LLM scientific text and recovers 85% of AI academic text. We made that false-positive rate the anchor. We took papers submitted in 2021 and 2022, before ChatGPT, treated them as ground-truth human, and set the flag threshold so that exactly 0.4% of them trip it. That line is the floor. Every number we report is a share of papers above a threshold where genuine pre-LLM writing sits, by construction, at 0.4%. The pre-ChatGPT years then act as a built-in control: if the rise were an artifact of the detector, 2021 and 2022 would flag as high as 2026. The first figure shows they do not.
We sampled ten field groups, roughly 25 papers per field per month, from January 2023 to July 2026, plus eight control months across 2021 and 2022, for 12,750 papers in total. For each one we pulled the version-1 PDF, so a paper revised in 2026 cannot leak modern text back into its 2023 slot. We scored the full body text instead of the abstract, because abstracts understate the signal: we have seen the same paper score under 20% on its abstract and over 70% on its body. Every reported figure carries a bootstrap 95% confidence interval.
The flagged share is flat at 0.4% through 2021 and 2022, lifts off within months of ChatGPT, and climbs in two waves to about 32% over the most recent complete quarter, peaking near 39% in early 2026. The spread across fields is large, and it is the table and the second figure that carry it. The values below are each field's flagged share over the 12 months to July 2026, alongside its pre-LLM control level.
| Field group | Pre-LLM control | Recent flagged share | 95% CI |
|---|---|---|---|
| Computer science | 0.2% | 65.0% | [59.3, 70.3] |
| Quantitative biology | 3.5% | 56.3% | [51.0, 61.7] |
| Electrical eng. & systems | 1.7% | 51.3% | [46.0, 57.0] |
| Economics & finance | 2.5% | 47.0% | [41.3, 52.7] |
| Applied physics | 1.3% | 34.0% | [29.0, 39.7] |
| Statistics | 1.8% | 31.3% | [26.0, 36.7] |
| Condensed matter | 0.0% | 24.0% | [19.3, 29.0] |
| High-energy physics | 0.5% | 14.0% | [10.0, 18.0] |
| Astrophysics | 0.0% | 10.7% | [7.3, 14.3] |
| Mathematics | 0.0% | 0.7% | [0.0, 1.7] |
Computer science leads at about 65%. Mathematics is lowest, near 0.7%, and the limitations section explains why its low value is hard to interpret. The control column is each field's 2021 to 2022 flag rate averaged over three sensitivity settings; the fields that rise most are not the ones with the highest pre-LLM control level, so an elevated starting point does not explain the rise.
Control sample size. Each field's pre-ChatGPT control is 200 papers. At a 0.4% flag rate only eight papers flag across the entire 2,000-paper control, spread thinly over ten fields, so a single-threshold per-field control rate is coarse. The pooled floor is well estimated and is what the study is anchored to, but the per-field control levels are only approximate, and a larger control would not fix this: pinning a fraction-of-a-percent rate per field would require thousands of control papers per field that pre-2023 arXiv volume does not contain.
A low score can indicate low adoption or a detector blind spot. Mathematics is the clearest case. Mathematics papers are dominated by notation and theorem-proof structure, and once equations and references are removed the remaining prose is sparse and unlike the scientific English the detector was trained on. A mathematics paper drafted with heavy model assistance may score low because its prose is out of distribution for the detector, so a low score in mathematics is weak evidence that a human wrote the paper. The result is consistent with two very different explanations, lower adoption or reduced detector sensitivity in that register, and this data cannot separate them. The fields with the strongest in-distribution assumption, the prose-heavy ones, are also the ones that rise most, so this confound does not account for the aggregate trend. But in the low-scoring fields the ranking should be read as a lower bound on adoption.
Detector coverage. The detector is more sensitive to some generators than others, and we cannot evaluate it against the exact, private mixture of models and prompts that authors actually use. Incomplete coverage lowers the flag rate, so the reported prevalence is a lower bound: the true share is at least what we measured. The detector write-up reports the per-generator performance.
A flag is not authorship. The detector estimates whether text reads as machine-written, at a calibrated probability with a known error rate. It cannot separate a lightly-edited document from a wholly-generated one, and a single score is never grounds to accuse a specific person. We report the prevalence of machine-like writing, which includes heavy AI-assisted editing.
The detector is cheap to run and we make no money from it. You can try it for free on any arXiv paper here, and on your own text here.
There. AI-polished sentence.
Not sure if its just me, lately I have started feeling pretty offensive about the increased usage of the word. Its management not leadership by any means.
> "leadership encourages this because from what they can tell, there is no downside"
For most people in management its easier to pick the current set of slangs/abbreviation's, general trend and go with it. Understanding the details would take time, raise questions and no one in management has time or political capital to spend on it.
That's well put.
> But good enough hence the real question is value aka time& money invested/quality.
There's time invested SO FAR and time that will have to be invested to maintain it. In my experience, even with Fable, it's not there yet. It's the reason why it's easy to vibe code an app from scratch, but at some point when complexity significantly increases, the codebase becomes a mess.
Humans write slop too, you know. Just saying.
If I extrapolate this example to my professional life, this code now manages millions of dollars, a single mistake can wipe it all out, it has to be maintained by 5 other engineers and understood by 5 other domain experts.