A llm emulating a person is why many of my used sites banned llms due to scraping bandwith costs
Is this about access or accessibility to claude (for example)
I am sure the entirely of human computing knowledge is not that big.
I think the right choice is pretty clear...
I would be surprised if a majority of ACM members were to say yes should we ask them (but ACM is not known for such democracy). Along with book authors, we are one of the many people that provide the knowledge and expertise on which large tech firms train their models, and get nothing in return. Actually, life is getting worse for us: extra workload in universities with students' AI use, a completely broken peer review system, etc. Hence the irony of ACM thinking about licensing, and only licensing, at a time where this is the least of our priorities.
Does the ACM really think LLMs haven't already consumed 90% of the content through other sources?
Either way, I'll pirate.
It's almost like, "don't hurt yourself unc, we will just search arxiv".
now it will be fodder for the slop machine
(i think that LLMs are going to wreck the peer review system for all but hard-experimental papers)
There's something hilarious about that, but also, snake eating its own tail.
You don't own the knowledge you put out there unless you have a limited time valid patent. The rest is absurdity. If you want to keep your findings to yourself, keep them secret.
I get the copyright aspect of this and I'm not arguing that here. I'm more asking about the moral / ethical idea of choosing who can benefit from your science.
Obviously there are the moral / ethical arguments about AI in general here to weigh against - those have been hashed out significantly elsewhere, and I'm not interested in debating them. What I'm asking about here is the impact on science by sharing it with tooling that distributes it in ways not generally considered when originally written.
A quick check of your post history suggests the frame that you work in strongly is privacy related research (observation - may be wrong). I'm curious how that impacts what you wrote here generally.
(Just to be perfectly clear, I'm not arguing your points here, trying to understand them better)
Yeah, too bad
AIs, at least in their current form, make you more who ever you were. If you want snap, glib answers of dubious accuracy, they'll give them to you, more easily than ever before. If you want to dig back into primary sources and get the original content, they'll do that for you, more easily than ever before.
Can't speak to how the science infrastructure is going to handle them, but if it takes down the peer review system, which I think has been worthless for probably going on two decades and has just given the entire enterprise a false sense of assurance, it'll probably be a net gain in the end. Peer review is a source of more problems then it is solving right now.
I think I may be shadowbanned, but at what point do we start viewing LLMs as a national security threat?
I suppose they could be sued.
The issue is not who owns knowledge, it's how it benefits humanity.
The entire point of copyright law is so that people can make their writing public and still be able to control the right to make copies (for example, into your dataset for training an LLM).
That said, what matters here is the social contract, what do I bring to society and what do we get from tech companies. For most people around the world, access to the typical leading models is out of reach. Not many on this planet can pay the subscriptions (or even API keys) that offer access to the best models. So I'm not buying the argument that tech companies are broadening access. What we're creating is a increasingly discriminatory society where the few get access to information, and the many don't.
Sorry to ruin your day, but if the people with money could have introduce equal society, you would have noticed their attempts by now.
https://dl.acm.org/openaccess ---> So how to access the content? Do I have to register or what? It the "open access" only for academic org's people or for everyone in the world?
https://dl.acm.org/ ---> Okay, nice simple search field without loggin in, but when you try to search something, you get thousands of results, even if you search specific author and the exact name of paper, you will get hundreds of results and the thing you want is buried somewhere on page 247. Filters of authors, years etc. are for Premium subscription. But just googling the thing finds the link to ACM... And want to get the actual PDF? Hope it says "free access"...
https://authors.acm.org/open-access/acm-open-for-authors-hom...
> Beginning January 2026, all ACM publications and related artifacts in the ACM Digital Library will be made open access.
If someone reads a philosophical essay, has a heureka moment from it, applies the principle to their work, and makes bank (commercial profit), they never have to pay a percentage to the author of the essay.
A non profit could well pay, and there are plenty of reasons frontier models should be managed by non profit. After all, why allow a for profit to benefit from free contributions?
Google scholar has become the way to search papers (which is somewhat worrisome). What people want from there is a download for all the stuff that does not list an author copy. Open access IMHO is just a reaction to the fact that mostly you would not need a subscription anyhow. Now the authors are paying upfront or universities are paying flat for all their researchers. The problem is now the incentives are not increasing the number of subscriptions but increasing the number of papers published.
Currently patents are predominantly used to prevent interoperability and impose costs.
The output of it is also a derivative work, and derivative works also infringe copyright. Its only not a problem if you ignore copyright entirely
Humans are the only entities that get to enjoy special idea-learning-exemptions, not AI
It's a racket.
A short session is just retrieval, a long session is always unique. The more the user writes the more it diverges from any content in the dataset.
As someone with lots of open source code out there that has likely been used as LLM training data, I'm very sympathetic to this point of view, but that doesn't seem to be the legal reality. Much of this has not been fully tested in court, but it seems likely that LLM training is not copyright infringement, as long as the training material itself was acquired legally.
Derivative work or transformative? It's not the same.
There's also been court cases where material has been found to be infringingly used, eg song lyrics, so the case where copyright ceases to exist doesn't seem to be coming through yet, thankfully. It'd be the most staggering upheaval of copyright of all time if this doesn't turn out to be true
AI works also clearly aren't transformative in many cases. If you ask it a question about a paper, it'll quote bits of the paper at you. That serves as an exact substitute of the original work. If you ask it for song lyrics, or information about the news, its content is a direct substitute for the original source it was trained on. This clearly does not fall under a transformative use case
You could argue that some uses of it are transformative, but even then - its easy to find some piece of training data in the source code that the output work supersedes. By its very nature it does not have the capacity to genuinely invent under the law (as it is not human), and a prompt isn't a significant enough part of the processing to count here
Everyone treats the human user of the AI as furniture, but they steer the whole process into unique directions.
> and yes that's 100% copyright infringement
Says what court of law?
I'm kinda getting tired of this stuff. I'm someone who has been, and still to some extent is, uncomfortable with the possibility of copyright/license laundering in LLMs, but they way you are making your argument is incredibly off-putting and not sympathetic. You're throwing out wild assertions about the law that are not supported by... anything, really.
You can use all the ideas you want, but AI cannot because its not a person, and does not enjoy the same protection under the law. The copyright holders by and large did not agree to you using their content like this
If we enable this, people won't create anything because all their work will immediately be stolen by the AI models. Copyright partially exists to promote the creation of new content, because theft disincentivises novel creation
>Says what court of law?
If you turn a png into a jpeg, and distribute it, that's copyright infringement. There isn't a court in the land that wouldn't find you guilty of that
Why ACM believes the rewards outweigh the risks in opening up the Digital Library to large language models.
Posted Jul 15 2026

Credit: Jackie Niam / Shutterstock
As AI systems increasingly become a primary interface through which knowledge is discovered and used, the Association for Computing Machinery (ACM) must decide how to engage with these technologies in a manner consistent with its mission and the long-term interests of the computing community. ACM has intentionally taken a cautious approach. Rather than rushing to monetize content, ACM prioritized protecting the integrity of the ACM Digital Library and understanding the views of authors and volunteer leaders. Discussions across ACM governance have shown that while skepticism remains, many now recognize that AI systems are becoming part of the infrastructure through which scientific knowledge is discovered, synthesized, and applied; continuing to exclude ACM Publications from the reach of large language models (LLMs), AI-related research-focused pilots, and a growing list of discovery services embedding AI into their services will put ACM’s mission and community of authors and researchers at risk.
ACM’s peer-reviewed corpus represents one of the world’s most authoritative sources of computing research. If trusted scholarship is excluded while lower-quality material is incorporated into AI systems, future knowledge tools may underrepresent ACM research.
Responsible inclusion of ACM Publications can help improve the quality and factual grounding of AI systems while making advanced computing knowledge more accessible across languages, regions, and levels of expertise. Researchers publish not only to archive discoveries but to influence future work. As students, engineers, policymakers, and practitioners increasingly rely on AI tools, responsible inclusion of ACM content can broaden the reach and impact of authors’ ideas while maintaining pathways back to the version of record.
AI-assisted literature reviews, coding, writing, and discovery are already changing research workflows. Publishers absent from AI ecosystems risk becoming repositories of record rather than sources of influence. Beyond model training, conversational interfaces, semantic discovery systems, agentic workflows, and retrieval-augmented generation (RAG) are becoming important components of scholarly communication.
Research influence may increasingly depend not only on citations and downloads, but also on whether ideas are surfaced, synthesized, and operationalized within AI systems. Responsible participation can strengthen the long-term visibility and impact of ACM authors, who publish their work with ACM to ensure the maximum level of distribution and impact possible. In addition, while licensing revenue is not, nor should it ever be, the primary motivation for including ACM Publications in the growing AI ecosystem, doing so could provide additional financial support for editorial systems, preservation, conferences, and community programs that could play a role in ACM’s long-term financial sustainability.
Significant concerns remain. Current AI systems do not consistently provide adequate attribution to authors and publishers. Hallucinations and misrepresentations can distort research findings and create unwarranted confidence in AI-generated outputs. Licensing also requires legal, technical, and operational frameworks capable of ensuring transparency, compliance, and appropriate rights management.
There are also broader concerns regarding concentration of economic value among commercial AI providers. These issues deserve careful consideration and must be addressed through governance, safeguards, and continued community oversight. Not all LLMs and AI-related initiatives are equal, and ACM will evaluate opportunities and risks carefully before entering into agreements with AI companies and initiatives.
The strategic question facing ACM is no longer simply whether to permit LLMs to train on ACM Publications, but how to ensure that ACM authors, their scholarship, and the version of record remain visible, attributable, and influential within an increasingly AI-mediated knowledge ecosystem. Training rights represent only one aspect of a broader transformation in how scientific information is discovered, synthesized, and used.
Responsible participation offers opportunities to improve AI systems, expand global access to computing knowledge, and preserve the visibility and long-term impact of ACM authors. At the same time, concerns surrounding attribution, hallucinations, governance, commercial concentration, and intellectual property require careful attention.
Beyond licensing itself, ensuring ACM Publications remain visible, discoverable, attributable, and connected to the version of record within AI-mediated discovery systems is an equally important strategic objective for ACM. Increasingly, researchers, students, practitioners, and policymakers are interacting with scientific knowledge through AI systems rather than traditional search interfaces. The long-term value of ACM Publications in these environments will depend not only on whether they are available to AI systems, but also on whether attribution is preserved, users can easily return to the original source, and meaningful measures of AI-mediated scholarly impact can evolve alongside traditional citations, downloads, and usage metrics. Framed in this way, the central question facing ACM is not simply whether to permit training, but how to ensure the work of ACM authors remains visible, discoverable, and influential within the next generation of knowledge-discovery systems.
Ultimately, AI systems will shape the future of scholarly communication. The question is whether ACM and the computing community will actively help shape that future in a responsible and mission-aligned manner.
Before entering into formal licensing agreements, ACM seeks broad community feedback and will continue evaluating frameworks for responsible AI training and retrieval use.
Click here to share your feedback with ACM, so that ACM leadership can take your feedback into account in developing and implementing ACM’s AI content strategy over the coming months. Thank you!
Submit an Article to CACM
CACM welcomes unsolicited submissions on topics of relevance and value to the computing community.
You Just Read
View in the ACM Digital Library

This work is licensed under a Creative Commons Attribution International 4.0 License.
© 2026 Copyright is held by the owner/author(s).
ACM encourages its members to take a direct hand in shaping the future of the association. There are more ways than ever to get involved.
By opening CACM to the world, we hope to increase engagement among the broader computer science community and encourage non-members to discover the rich resources ACM has to offer.
The main question in the "AI image generator generates a Pikachu image" is whether the AI company serving that image generator to you is violating the copyright or not. Because they make money when doing so (API / subscription cost), and so it's like selling images of Pikachu. The user is likely in the clear as long as they don't go on sell that Pikachu further. But the AI company sold the Pikachu image to the user.
Not only are you not winning this one but I'm gonna laugh at you the entire time.