ArXiv declares independence from Cornell

The recent announcement to reject review articles and position papers already smelled like a shift towards a more "opinionated" stance, and this move smells worse.

The vacuum that arXiv originally filled was one of a glorified PDF hosting service with just enough of a reputation to allow some preprints to be cited in a formally published paper, and with just enough moderation to not devolve into spam and chaos. It has also been instrumental in pushing publishers towards open access (i.e., to finally give up).

Unfortunately, over the years, arXiv has become something like a "venue" in its own right, particularly in ML, with some decently cited papers never formally published and "preprints" being cited left and right. Consider the impression you get when seeing a reference to an arXiv preprint vs. a link to an author's institutional website.

In my view, arXiv fulfills its function better the less power it has as an institution, and I thus have exactly zero trust that the split from Cornell is driven by that function. We've seen the kind of appeasement prose from their statement and FAQ [1] countless times before, and it's now time for the usual routine of snapshotting the site to watch the inevitable amendments to the mission statement.

"What positive changes should users expect to see?" - I guess the negative ones we'll have to see for ourselves.

[1] https://tech.cornell.edu/arxiv/

> raised concerns about the proposed $300,000 salary for arXiv’s new CEO, saying it seemed high

Is a mid-to-high engineering salary outlandish for a CEO of what is likely to be a fairly major non-profit? Even non-profits have to be somewhat competitive when it comes to salary, and the ideal candidate is likely someone who would be balancing this against a tenured position at a major university

It's not that hard to make a mirror or arXiv. Basically, anybody who can pay for hosting (which, I suppose, isn't very cheap now when the whole world uses it). It's a problem to make users switch, because academia seems to have this weird tradition of resisting all practices that, god forbid, might improve global research capabilities and move forward the scientific progress. But then, if arXiv actually becomes unusable, I suppose they won't really have much choice than to switch?

And, FWIW, I do think that arXiv truly has a vast potential to be improved. It is currently in the position to change the whole process of how the research results are shared, yet it is still, as others have said, only a PDF hosting. And since the universities couldn't break out of the whole Elsevier & co. scam despite the internet existing for the 30 years, to me, breaking free from the university affiliation sounds like a good thing.

But, of course, I am talking only about the possibilities being out there. I know nothing about the people in charge of the whole endeavor, and ultimately in depends on them only, if it sails or sinks.

I'm not sure why we're so focused on filtering what gets into arxiv (which is an uphill battle and DOA at this point) vs fixing the indexing, i.e. the page rank of academia.

Google "sorted out" a messy web with pagerank. Academic papers link to each others. What prevents us from building a ranking from there?

I'm conscious I might be over-simplifying things, but curious to see what I am missing.

I might be missing something, but I still don't get the why. I don't see any "problem" that needs to be solved.

Given that Cornell charges what, $50k a year as an Ivy League, $300k feels like almost nothing.

From my limited experience, arXiv appears to include many low-quality, unreproducible papers, and some are straight-up self-marketing rather than serious scientific work.

What is worrisome about this development, and corollary actions like the hiring of a CEO with a $300,000/year salary, is that the essentially independent and community based platform will disappear. The ArXiv exists because mathematicians and physicists, and later computer scientists and engineers, posted there, freely, their work, with minimal attention to licensing and other commercial aspects. It has thrived because it required no peer review and made interesting things accessible quickly to whomever cared to read them.

A setup as a US-based "non-profit" is worrisome, if only because 300K is an obscene salary even in a for-profit setting. That the US-based posters can't see this is evidence of the basic problem which is that the US, both left and right, has been taken over by a neoliberal feudal antidemocratic nativist mindset that is anathema to the sort of free interchange of ideas that underlay the ArXiv's development in the hands of mathematicians and physicists now swept aside and ignored by machine learning grifters and technicians who program computers.

I fear their Mozilla-ification and Wikipedia-ification. Scope creep, various outreach feel-good programs, ballooning costs, lost focus etc. And other types of enshittification.

Any change to the basic premise will be a negative step.

They should just be boring quiet unopininionated neutral background infrastructure.

I wonder if there are plans to licence the content for AI training

This sounds terrible. Of course there's a huge risk of it becoming made for-profit. It almost makes you wonder if the academic publishers are behind this push somehow.

Could they not have made it into some legal structure that puts universities at the top? Say, with a bunch of universities owning shares that comprise the entirety of the ownership of arXiv, but that would allow arXiv to independently raise funds?

Do research papers published on Elsevier's sort of media remain more prestigious?

I read a dozen papers a month, typically on arxiv, never from paywalled journals. I find the quality on par. But maybe I'm missing something.

Now the question is, will arxiv wage a decade long bloody war with Cornell, using heavy infantry (PhD students), archers (reviewers) and field artillery (AI slop papers), or will the independence be mostly peaceful? Only time can tell.

This is exactly what happened last time when scientific publishing got cornered. Journals run by departments and research groups were spun out or sold off to publishers and independent orgs. And they continued to slowly boil the frog over 50 years with fees and gate keeping.

Its especially problematic because while ArXiv love to claim to be working for open science, they don't default to open licensing. Much of the publications they host are not Open Access, and are only read access. So there is definitely the potential to close things off at some point in the future, when some CEO need to increase value.

And they hired a LinkedIn business idiot to run the new organization - so the aim is for an infinite growth tech startup in terms of governance, despite the technical legal status of non-profit. It shows in the language they use in the announcement, too ("improved financial viability in the long run")

OpenAI shows exactly how well that works and what that kind of governance does to a company and to its support of science and the commons.

TL;DR, it's fucked.

>Cornell, for example, had a limited capacity to pay software developers to maintain and upgrade the site, which still has a very no-frills look and feel.

arXiv is doomed. It was nice while it lasted.

arXiv is great. It's just a problem that there's so much slop. What if arXiv offered a subscription service that people in different fields could use to just see a curated selection of the top papers in their field each month. Established researchers in each field could then review some of the preprints for putting into the curated monthly list.

Oh, wait.

we got this before gta 6

Maybe they should implement a graph based trust system:

You need your favourite academic gatekeeper (= thesis advisor) to vouch for you in order to be allowed to upload.

Then AI slop gets flagged and the shame spreads through the graph. And flaggings need to have evidence attached that can again be flagged.

With 300K for the CEO, its enshittification will commence imminently. It will now serve to maximize revenue. Just wait and watch while they issue a premium membership, payment requirements for authors, and other revenue generators to please their investors.

.. and soon to be dependent on US military funding? Controlled by someone who has run-ins with universities? This'll end in tears.

ArXiv is dead. Expect a paywall within three years, or other enshittification and slop added.

"Recently arXiv’s growth has accelerated. Since 2022, it has expanded its staff to 27, in large part to deal with a 50% increase in submitted manuscripts."

I am wary of that. IMO the business model is damaged therein. You can say in 2022 we had 27; bankrupt in 2030.

Good call, ArXiv seems like one of the most important institutions out there right now.

Frankly, the only beef I have with arXiv as is: its insistence on blocking AI access.

I had to tell my AI to set up an MCP for "fetch while bypassing arXiv's rate limit" so that it doesn't burn 40k tokens looking for workarounds every time it wants to look at a paper and gets hit with a "sorry, meatbags only" wall.

Very annoying, given how relevant arXiv papers are for ML specifically, and how many of papers there are. Can't "human flesh search" through all of them to pick the relevant ones for your work, and they just had to insist on making it harder for AIs to do it too.

Very unrelated to the article, but I think 'arXiv' as a brand is bad, and really detrimental to what the institution aims to accomplish.

That is, it's not readily parseable, it really gives an insider term vibe - like this isn't for you if you don't already know what it means or how you should read or say it. It sort of reminds me of the overuse of latin and latinate terms generally in the old professions and, well, the academy.

Just always struck me as being somewhat at odds with the goal.

OpenAI shows exactly how well that works and what that kind of governance does to a company and to its support of science and the commons.

TL;DR, it's fucked.

we got this before gta 6

.. and soon to be dependent on US military funding? Controlled by someone who has run-ins with universities? This'll end in tears.

"Recently arXiv’s growth has accelerated. Since 2022, it has expanded its staff to 27, in large part to deal with a 50% increase in submitted manuscripts."

I am wary of that. IMO the business model is damaged therein. You can say in 2022 we had 27; bankrupt in 2030.

The recent announcement to reject review articles and position papers already smelled like a shift towards a more "opinionated" stance, and this move smells worse.

"What positive changes should users expect to see?" - I guess the negative ones we'll have to see for ourselves.

[1] https://tech.cornell.edu/arxiv/

> and with just enough moderation to not devolve into spam and chaos

arXiv has become a target for grifters in other domains like health and supplements. I’ve seen several small scale health influencers who ChatGPT some “papers” and then upload them to arXiv, then cite arXiv as proof of their “published research”. It’s not fooling anyone who knows how research work but it’s very convincing to an average person who thinks that that they’re doing the right thing when they follow sources that have done academic research.

I’ve been surprised as how bad and obviously grifty some of the documents I’ve seen on arXiv have become lately. Is there any moderation, or is it a free for all as long as you can get an invite?

> Unfortunately, over the years, arXiv has become something like a "venue" in its own right, ...

In my experience as a publishing scientist, this is partly because publishing with "reputable" journals is an increasingly onerous process, with exorbitant fees, enshittified UIs, and useless reviews. The alternative is to upload to arXiv and move on with your life.

I came here to say something similar. As someone who works in a field that applies machine learning but is not purely focused on it, I interact with people who think that arXiv is the only relevant platform and that they don't need to submit their work to any journal, as well as people who still think that preprints don't count at all and that data isn't published until it's printed in an academic journal. It can feel like a clash of worlds.

I think both sides could learn from the other. In the case of ML, I understand the desire to move fast and that average time to publication of 250-300 days in some of the top-tier journals can feel like an unnecessary burden. But having been on both sides of peer review, there is value to the system and it has made for better work.

Not doing any of it follows the same spirit as not benchmarking your approach against more than maybe one alternative and that already as an after-thought. Or benchmaxxing but not exploring the actual real-world consequences, time and cost trade offs, etc.

Now, is academic publishing perfect? Of course not, very very far from it. It desperately needs to be reformed to keep it economically accessible, time efficient for both authors, editors and peer reviewers and to prevent the "hot topic of the day" from dominating journals and making sure that peer review aligns with the needs of the community and actually improves the quality of the work, rather than having "malicious peer review" to get some citations or pet peeves in.

Given the power that the ML field holds and the interesting experiments with open review, I would wish for the field to engage more with the scientific system at large and perhaps try to drive reforms and improve it, rather than completely abandoning it and treating a PDF hosting service as a journal (ofc, preprints would still be desirable and are important, but they can not carry the entire field alone).

> arXiv fulfills its function better the less power it has as an institution

It is an interesting instance of the rule of least power, https://en.wikipedia.org/wiki/Rule_of_least_power.

> Unfortunately, over the years, arXiv has become something like a "venue" in its own right, particularly in ML, with some decently cited papers never formally published and "preprints" being cited left and right. Consider the impression you get when seeing a reference to an arXiv preprint vs. a link to an author's institutional website.

This just isn't true. arXiv is not a venue. There's no place that gives you credit for arXiv papers. No one cares if you cite an arXiv paper or some random website. The vast vast majority of papers that have any kind of attention or citations are published in another venue.

My observation is that research, especially in AI has left universities, which are now focusing their research to a lesser degree on STEM. It appears research is now done by companies like Meta, OpenAI, Anthropic, Tencent, Alibaba, among many others.

> raised concerns about the proposed $300,000 salary for arXiv’s new CEO, saying it seemed high

Salaries in the US are so bonkers. Everywhere else outside of the US, $300,000 is an outlandish high salary. To call it "mid to high" is insane.

For anybody outside the SV, and especially outside the US, this seems high, yes.

arXiv does not need to and should not optimize for “shareholder value”, which is at least nominally the justification for outlandish CEO pay packages.

arXiv's CEO doesn't need to be a tenured professor equivalent it is a preprint repository ffs.

Should be the main link. The original article is based on the CEO job posting.

I might be missing something, but I still don't get the why. I don't see any "problem" that needs to be solved.

The article lists the reasons quite clearly.

I think the problem described in 6th paragraph needs to be solved.

As a US based academic, I have to say when I saw the salary I immediately gawked. I think it's not americans but silicon valley-ites and tech bros on here who have lived with inflated salary/net worth that think it's just a middle of the road salary. As I regularly interact with friends in engineering who make like $200k + benefits ($), and I wonder why I don't jump ship to that weird land.

I fear their Mozilla-ification and Wikipedia-ification. Scope creep, various outreach feel-good programs, ballooning costs, lost focus etc. And other types of enshittification.

Any change to the basic premise will be a negative step.

They should just be boring quiet unopininionated neutral background infrastructure.

> Mozilla-ification

All the Mozilla executives have done for the last 15+ years is

* lay off developers

* spend lots of money on stupid side projects nobody asked for or wants

* increase their own salaries

and all that with the backdrop of falling quality, market share, and relevance.

I would happily donate to Firefox, but this fucked up organization will never see a single cent from me. They will spend it on anything but Firefox, which is the only thing anybody wants them to spend it on.

It might already be too late, and we will be left with a browser monopoly.

> They should just be quiet unopininionated neutral background infrastructure.

Exactly. It should be a utility. Not quite dumb pipe, but not too far either.

I wonder if there are plans to licence the content for AI training

It's been available all along: https://info.arxiv.org/help/bulk_data.html

Id guess OAI & co have already copied without asking?

This sounds terrible. Of course there's a huge risk of it becoming made for-profit. It almost makes you wonder if the academic publishers are behind this push somehow.

> Of course there's a huge risk of it becoming made for-profit.

The article says that "it will become an independent nonprofit corporation", and as OpenAI's failed attempt showed, converting a non-profit to a for-profit organization is either really hard or impossible.

> Could they not have made it into some legal structure that puts universities at the top?

As a corporation (even a non-profit one), it will have a board of directors. I have no idea what their charter will look like, but I would be surprised if at least one seat wasn't reserved for a university representative, and more than that seems quite likely as well.

PhD students are levy infantry at best with Postdocs being the armoured levies.

Maybe they should implement a graph based trust system:

You need your favourite academic gatekeeper (= thesis advisor) to vouch for you in order to be allowed to upload.

Then AI slop gets flagged and the shame spreads through the graph. And flaggings need to have evidence attached that can again be flagged.

They already had a basic form of this for a while [1]

> arXiv requires that users be endorsed before submitting their first paper to arXiv or a new category.

[1] https://info.arxiv.org/help/endorsement.html

The endorsement system already works along that line: https://info.arxiv.org/help/endorsement.html

It's probably not perfect but in practice, it seems to have been enough to get rid of the worst crackpotty spam.

I've often thought that similar trust systems would work well in social media, web search, etc., but I've never seen it implemented in a meaningful way. I wonder what I'm missing.

Science reduced to people with a phd?

they'll just turn into a shitty journal at this point, they just need to introduce peer review and they can start competing with the real journals on price point.

another will need to rise to take its place.

Good call, ArXiv seems like one of the most important institutions out there right now.

It’s so important, in fact, that there should be more than one such institution.

People keep falling into the same trap. They love monopolies, then are shocked when those monopolies jerk them around.

Frankly, the only beef I have with arXiv as is: its insistence on blocking AI access.

> and with just enough moderation to not devolve into spam and chaos

I’ve been surprised as how bad and obviously grifty some of the documents I’ve seen on arXiv have become lately. Is there any moderation, or is it a free for all as long as you can get an invite?

Should be the main link. The original article is based on the CEO job posting.

> Mozilla-ification

All the Mozilla executives have done for the last 15+ years is

* lay off developers

* spend lots of money on stupid side projects nobody asked for or wants

* increase their own salaries

and all that with the backdrop of falling quality, market share, and relevance.

It might already be too late, and we will be left with a browser monopoly.

> They should just be quiet unopininionated neutral background infrastructure.

Exactly. It should be a utility. Not quite dumb pipe, but not too far either.

It's been available all along: https://info.arxiv.org/help/bulk_data.html

Id guess OAI & co have already copied without asking?

> Of course there's a huge risk of it becoming made for-profit.

> Could they not have made it into some legal structure that puts universities at the top?

PhD students are levy infantry at best with Postdocs being the armoured levies.

They already had a basic form of this for a while [1]

> arXiv requires that users be endorsed before submitting their first paper to arXiv or a new category.

[1] https://info.arxiv.org/help/endorsement.html

The endorsement system already works along that line: https://info.arxiv.org/help/endorsement.html

It's probably not perfect but in practice, it seems to have been enough to get rid of the worst crackpotty spam.

I've often thought that similar trust systems would work well in social media, web search, etc., but I've never seen it implemented in a meaningful way. I wonder what I'm missing.

Science reduced to people with a phd?

they'll just turn into a shitty journal at this point, they just need to introduce peer review and they can start competing with the real journals on price point.

another will need to rise to take its place.

The French government put a bit of money on the table to help researchers fulfil their open science requirements for government and EU grants, and funded the HAL repository ( https://hal.science/ ). It’s much smaller than arXiv, but it exists. In other countries like the UK there are clusters of smaller repositories as well, but it’s not as well centralised.

It’s so important, in fact, that there should be more than one such institution.

People keep falling into the same trap. They love monopolies, then are shocked when those monopolies jerk them around.

it just hosts pdfs, no?

Very unrelated to the article, but I think 'arXiv' as a brand is bad, and really detrimental to what the institution aims to accomplish.

Just always struck me as being somewhat at odds with the goal.

I wonder what makes you feel that. I've been publishing preprints close to a decade on arxiv now and never had any particular feelings about it.

To me it's just a way to get out your work fast, so that there is already a trace of it on the Internets - nothing more and nothing less.

> That is, it's not readily parseable, it really gives an insider term vibe...

Isn't that normal with highly specialized research fields? I agree many papers could benefit from clearer wording, but working in a niche means you sometimes don't reach a broader audience

It's a classic story of someone having to pick a name quickly, which then gets established long before anyone who cares about branding is aware of its existence.

The original service didn't even have a name, only a description, and it was amusingly hosted at xxx.lanl.gov. But LANL wasn't really interested in it, and the founder eventually left for Cornell. At that point, the service needed a domain name, but archive.org was already taken.

And besides, the name has Ancient Greek influences. A similar Latinate term might be something like "archive".

> like this isn't for you if you don't already know what it means

Isn't that actually kindof a good brand signal for a repo of very specialized papers? "Fun with learning" in comic sans wouldn't help credibility.

This the type of guy that will suggest paper.ly as a better name with a straight face and then we wonder why the internet is turning to shit

Simply anticipating basic push backs from reviewers makes sure that you do a somewhat thorough job. Not 100% thorough and the reviews are sometimes frivolous and lazy and stupid. But just knowing that what you put out there has to pass the admittedly noisily gatekept gate of peer review overall improves papers in my estimation. There is also a negative side because people try to hide limitations and honest assessments and cherry pick and curate their tables more in anticipation of knee jerk reviewers but overall I think without any peer review, author culture would become much more lax and bombastic and generally trend toward engagement bait and social media attention optimized stuff.

The current balance where people wrote a paper with reviers in mind, upload it to Arxiv before the review concludes and keep it on Arxiv even if rejected is a nice balance. People get to form their own opinion on it but there is also enough self-imposed quality control on it just due to wanting it to pass peer review, that even if it doesn't pass peer review, it is still better than if people write it in a way that doesn't care or anticipate peer review. And this works because people are somewhat incentivized to get peer reviewed official publications too. But being rejected is not the end of the world either because people can already read it and build on it based on Arxiv.

> Unfortunately, over the years, arXiv has become something like a "venue" in its own right, ...

That’s true. But that’s separate than the use in ML in Blockchain circles as a form of a marketing - using academic appearances.

A Fields medal was awarded based mainly on this paper never published elsewhere: https://arxiv.org/abs/math/0211159

> arXiv fulfills its function better the less power it has as an institution

It is an interesting instance of the rule of least power, https://en.wikipedia.org/wiki/Rule_of_least_power.

The irony of the TBL quotes there being the entire problem with the semantic web is the ontological tarpit that results due to the excessive expressive power of a general triple store.

Salaries in the US are so bonkers. Everywhere else outside of the US, $300,000 is an outlandish high salary. To call it "mid to high" is insane.

Everyone outside the US doesn't deal with USD. Your comment is bonkers. Read up on purchasing power. All locations are not equal.

Even in the states, it’s more a distortion caused by the big tech centres. A software engineer in Ohio doesn’t command that kind of salary, but in San Francisco or Seattle that’ll buy you a moderately-senior engineer.

And while academic salaries are generally not great, tenured professors at big universities tend to make a fair bit (plus a lot more vacation time and perks than is normal in the US)

>Salaries in the US are so bonkers.

Sure, but the cost of living there is significantly higher as well. Anyway, I can hardly even comprehend these kinds of sums, though I am a bit of an outlier, as I earn around $27,700 as an SWE in Europe, which is low even by the standards of companies in my own country.

Note that you are seeing an explicit tradeoff of different economic systems.

Not everywhere. Switzerland exists. Also cost of living is a thing so if anything US/CH just ramp up to match that. The rest of Europe has high CoL but terrible salaries. Asia has bad salaries but low CoL (on average).

So is the living cost. Insurance, housing, etc. A better comparison is PPP.

Yes the obvious play is to move human labor to cheaper countries like France (including CEO of course).

That's a specific field at a very specific time. In general there is a difference between research and development, you're going to expect the early work to be done in academia but the work to turn that into a product is done by commercial organizations.

You get ahead as an academic computer scientist, for instance, by writing papers not by writing software. Now there really are brilliant software developers in academic CS but most researchers wrote something that kinda works and give a conference talk about it -- and that's OK because the work to make something you can give a talk about is probably 20% of the work it would take to make something you can put in front of customers.

Because of that there are certain things academic researchers really can't do.

As I see it my experience in getting a PhD and my experience in startups is essentially the same: "how do you do make doing things nobody has ever done before routine?" Talk to people in either culture and you see the PhD students are thinking about either working in academia or a very short list of big prestigious companies and people at startups are sure the PhDs are too pedantic about everything.

It took me a long time of looking at other people's side projects that are usually "I want to learn programming language X", "I want to rewrite something from Software Tools in Rust" to realize just how foreign that kind of creative thinking is to people -- I've seen it for a long time that a side project is not worth doing unless: (1) I really need the product or (2) I can show people something they've never seen before or better yet both. These sound different, but if something doesn't satisfy (2) you can can usually satisfy (1) off the shelf. It just amazes me how many type (2) things stay novel even after 20 years of waiting.

Universities (outside a few) just have much weaker PR machines so you never hear what they do. Also their work is not user facing products so regular people, even tech power users won't see them.

arXiv's CEO doesn't need to be a tenured professor equivalent it is a preprint repository ffs.

It's a bit more complex than an S3 bucket though because the value comes from the reputation network, which can't really be replicated easily.

Though, saying that, I suppose all the reputation data is kind of public. Apart from emails/accounts.

For anybody outside the SV, and especially outside the US, this seems high, yes.

arXiv does not need to and should not optimize for “shareholder value”, which is at least nominally the justification for outlandish CEO pay packages.

$300k for a top executive position isn't especially high for anywhere in the US. That's around what the administrative director of a hospital would be making, which seems like a much smaller scope than leading ArXiv. For comparison, my roommate works for a non-profit that serves Philadelphia whose CEO's salary is $1.1 million. General average for US CEOs including for profits is around $800k and for large organizations tens of millions is not atypical.

Non-profits aren't maximizing stock value, but they do need to optimize for stakeholder value - you want to maximize the amount of money being donated in and you want to make the most of the donations you receive, both to advance the primary mission of the non-profit and to instill confidence in donors. This demands competent leadership. The idea that just because something is not being done for profit means the value of the person's contributions is worth less is absurd.

arXiv doesn't need much. All they do is host static pdfs uploaded by someone else with free CDN services from Fastly [0]. I'm sure they could get academics to volunteer moderation services as well.

In reality you could host the entire thing for well under $50k/year in hardware and storage if someone else is providing a free CDN. Their costs could be incredibly low.

But just like Wikipedia I see them very likely very quickly becoming a money hole that pretends to barely be kept afloat from donations. All when in reality whats actually happening is that its a ridiculous number of rent seekers managed to ride the coattails of being the defacto preprint server for AI papers to land themselves cushy Jobs at a place that spends 90+% of their money on flights and hotels and wages for their staff.

I'm already expecting their financial reports to look ridiculously headcount heavy with Personnel Expenses, Meetings and Travel blowing up. As well as the classic Wikipedia style we spend a ton of money in unclear costs [1].

Whats already sad is they stopped having a real broken down report that used to actually showed things. Like look at this beautiful screenshot of a excel sheet. Imagine if Wikipedia produced anything this clear. [2]

[0] https://blog.arxiv.org/2023/12/18/faster-arxiv-with-fastly/

[1] https://info.arxiv.org/about/reports/FY26_Budget_Public.pdf

[2] https://info.arxiv.org/about/reports/2020_arXiv_Budget.pdf

From my limited experience, arXiv appears to include many low-quality, unreproducible papers, and some are straight-up self-marketing rather than serious scientific work.

Do research papers published on Elsevier's sort of media remain more prestigious?

I read a dozen papers a month, typically on arxiv, never from paywalled journals. I find the quality on par. But maybe I'm missing something.

Oh, wait.

Everyone outside the US doesn't deal with USD. Your comment is bonkers. Read up on purchasing power. All locations are not equal.

Note that you are seeing an explicit tradeoff of different economic systems.

>Salaries in the US are so bonkers.

Because of that there are certain things academic researchers really can't do.

We don't do 'utility' in America. Everything has S.V. brain rot - it's mixed with wall street brain rot, and now if you aren't extracting wealth out of what you have access to - you are failing.

> It might already be too late, and we will be left with a browser monopoly.

Ladybird continues to have the appearance of making progress, fwiw:

https://ladybird.org/newsletter/2026-02-28/

>They will spend it on anything but Firefox, which is the only thing anybody wants them to spend it on.

Mozilla certainly won’t spend it on Firefox, because the structure of the organization legally prohibits them from spending any of their donation money on Firefox. The ‘side projects’ are, at least officially, the real purpose of Mozilla.

And it is a risk for Arxiv too that once they start to drink the koolaid and start going to the same cocktail parties that these kinds of nonprofit board members and execs go to and will feel the need to prance around with some fancy stuff.

"oh no, you see we are not a preprint server host anymore, our mission is a values driven blablabla to make a meaningful change in the blablabla, we have spent X dollars to promote the blablabla, take me seriously please I'm also fancy like you! "

No need to ask - the whole point is open access. https://info.arxiv.org/help/bulk_data.html

Is this Gondor or Mordor?

OpenAI didn't get everything that they wanted, but I very much disagree with calling it a "failed attempt". The non-profit went from owning the entirety of OpenAI to having ~25% stake.

Is your argument really that "OpenAI was an independent nonprofit corporation and it worked out great, Arxiv will remain just as non-profit as OpenAI"?

Lobsters has this I think. But it also means I've never posted there.

not a bad first order filter.

can you think of a better one?

> they'll just turn into a shitty journal at this point

To this end, they added an endorsement requirement this year: https://blog.arxiv.org/2026/01/21/attention-authors-updated-...

I am using Zenodo for a while now instead. It is more user friendly, as well.

It is just a preprint repository. It is pretty open (the stories where a preprint was rejected or delayed unreasonably are extremely rare). It offers the basic services for a math/compsci/physics themed preprint repository.

I don't see much of a monopoly, nor any "moat" apart from it being recognised. You can already post preprints on a personal website or on github, and there are "alternatives" such as researchgate that can also host preprints, or zenodo. There are also some lesser known alternatives even. I do not see anything special in hosting preprints online apart from the convenience of being able to have a centralised place to place them and search for them (which you call "monopoly"). If anything, the recognisability and centrality of arxiv helped a lot the old, darker days to establish open access to papers. There was a time when many journals would not let you publish a preprint, or have all kinds of weird rules when you can and when you can't. Probably still to some degree.

I'm not sure why we're so focused on filtering what gets into arxiv (which is an uphill battle and DOA at this point) vs fixing the indexing, i.e. the page rank of academia.

Google "sorted out" a messy web with pagerank. Academic papers link to each others. What prevents us from building a ranking from there?

I'm conscious I might be over-simplifying things, but curious to see what I am missing.

I am of the same opinion, and ultimately ArXiv becoming a journal that can prevent one from publishing a paper — no matter how junk it is — would pretty much kill its purpose. But I suppose that now when flooding the interned with LLM-generated garbage is almost endorsed by some satanic people, it is pretty much a security issue to have some sort of filter on uploads.

Now, honestly, I have no idea why would one spend resources on uploading terabytes of LLM garbage to arXiv, but they sure can. Even if some crazy person is publishing like 2 nonsense papers daily, it is no harm and, if anything, valid data for psychology research. But if somebody actually floods it with non-human-generated content, well, I suppose it isn't even that expensive to make ArXiv totally unusable (and perhaps even unfeasible to host). So there has to be some filtering. But only to prevent the abuse.

Otherwise, I indeed think that proper ranking, linking and user-driven moderation (again, not to prevent anybody from posting anything, but to label papers as more interesting for the specific community) is the only right way to go.

tangentially related: https://readabstracted.com/

Page rank was inspired by bibliometrics and evaluation of science publications. It's messed up now because of the rankings. Further fiddling with ranking will not fix the problem.

Given that Cornell charges what, $50k a year as an Ivy League, $300k feels like almost nothing.

This is going to be in NYC where $300k does not go as far as it does in Ithaca.

Heh, you might want to look up what they’re charging young people now.

>Cornell, for example, had a limited capacity to pay software developers to maintain and upgrade the site, which still has a very no-frills look and feel.

arXiv is doomed. It was nice while it lasted.

I am not a software engineer, although I do write programs. What is it about digital infrastructure that requires maintenance? In the natural world, there is corrosion, thermal fluctuation, radiation, seismic activity, vandalism, whathaveyou. What are the issues facing the arxiv demanding the attention of multiple people 'round the clock?

ArXiv is dead. Expect a paywall within three years, or other enshittification and slop added.

Maybe they'll do something like what Anna’s Archive did

I really am not sure about that: https://biologue.plos.org/wp-content/uploads/sites/7/2020/05...

The problem is that "optimizing for peer-review" is not the same thing as optimizing for quality. E.g., I like to add a few tongue-in-cheeks to entertain the reader. But then I have to worry endlessly about anal-retentive reviewers who refuse to see the big picture.

A Fields medal was awarded based mainly on this paper never published elsewhere: https://arxiv.org/abs/math/0211159

I think there is a misunderstanding here. Does arXiv count as a publication? Yes, pretty much anything that gives you a DOI does, for example Zenodo. Does it function as a reputable anything? No.

The paper you link to counts as a publication, but its reputation stands on its own, it has nothing to do with arXiv as a venue. Ideally, that's how it is for all papers, but it isn't, just by publishing in certain venues your paper automatically gets a certain amount of reputation depending on the venue.

That’s true. But that’s separate than the use in ML in Blockchain circles as a form of a marketing - using academic appearances.

That sounds more like an issue of certain fields having crappy standards because the people in those fields benefit from crappy standards than an issue with the site they happen to host papers on.

The irony of the TBL quotes there being the entire problem with the semantic web is the ontological tarpit that results due to the excessive expressive power of a general triple store.

Well, I’d argue that many things in the semweb are not expressive enough and lead to the misunderstandings we have.

People think, for instance, that RDFS and OWL are meant to SHACL people into bad an over engineered ontologies. The problem is these standards add facts and don’t subtract facts. At risk of sounding like ChatGPT: it’s a data transformation system not a validation system.

That is, you’re supposed to use RDFS to say something like

  ?s :myTermForLength ?o -> ?s :yourTermForLength ?o .

The point of the namespace system is not to harass you, it is to be able to suck in data from unlimited sources and transform it. Trouble is it can’t do the simple math required to do that for real, like

  ?s :lengthInFeet ?o -> ?s :lengthInInches 12*?o .

Because if you were trying OWL-style reasoning over arithmetic you would run into Kurt Gödel kinds of problems. Meanwhile you can’t subtract facts that fail validation, you can’t subtract facts that you just don’t need in the next round of processing. It would have made sense to promote SHACL first instead of OWL because garbage-in-garbage out, you are not going to reason successfully unless you have clean data… but what the hell do I know, I’m just an applications programmer who models business processes enough to automate them.

Similarly the problem of ordered collections has never been dealt with properly in that world. PostgreSQL, N1QL and other post-relational and document DB languages can write queries involving ordered collections easily. I can write rather unobvious queries by hand to handle a lot of cases (wrote a paper about it) but I can’t cover all the cases and I know back in the day I could write SPAQL queries much better than the average RDF postdoc or professor.

As for underengineering, Dublin Core came out when I worked at a research library and it just doesn’t come close in capability to MARC from 1970. Larry Masinter over at Adobe had to hack the standard to handle ordered collections because… the authors of a paper sure as hell care what order you write their names in. And it is all like that: RDF standards neglect basic requirements that they need to be useful and then all the complex/complicated stuff really stands out. If you could get the basics done maybe people would use them but they don’t.

And while academic salaries are generally not great, tenured professors at big universities tend to make a fair bit (plus a lot more vacation time and perks than is normal in the US)

It's also caused by progressive tax rates. People take harder jobs based on net wage, not gross wage, so gross wage has to compensate.

So is the living cost. Insurance, housing, etc. A better comparison is PPP.

Living costs are similarly high in many places that have nowhere near the salaries of the US.

It's still the land of opportunities. It's easier to find ways to reduce your living costs than ways to increase your salary.

According to swissdevjobs.ch[1], the top 10% salary for a senior software developer in Switzerland is 135,000 swiss franc; that's roughly $170,000 per year.

So if this is correct, then even in Switzerland, it seems like $300,000 per year would be an obscenely high salary for a senior developer.

[1]: https://swissdevjobs.ch/salaries/all/all/Senior

Yes the obvious play is to move human labor to cheaper countries like France (including CEO of course).

The net salary in France might be low but the overall cost of hiring is quite high. Besides, why go to the middle when you can just find even cheaper places, if that's your prime metric?

The reason the French can’t build these things is the same reason they shouldn’t be allowed to be in charge. It’s a preprint PDF host. Just make your own if you can run this one.

It's a bit more complex than an S3 bucket though because the value comes from the reputation network, which can't really be replicated easily.

Though, saying that, I suppose all the reputation data is kind of public. Apart from emails/accounts.

> It's a bit more complex than an S3 bucket

It’s even less. I would bet if it’s not now, for the vast majority of its life it was a machine at someone’s desk at Cornell.

Universities (outside a few) just have much weaker PR machines so you never hear what they do. Also their work is not user facing products so regular people, even tech power users won't see them.

Not sure about that. How would a university test scaling hypotheses in AI, for example? The level of funding required is just not there, as far as I know.

I think there is a misunderstanding here. Does arXiv count as a publication? Yes, pretty much anything that gives you a DOI does, for example Zenodo. Does it function as a reputable anything? No.

I really am not sure about that: https://biologue.plos.org/wp-content/uploads/sites/7/2020/05...

Well, I’d argue that many things in the semweb are not expressive enough and lead to the misunderstandings we have.

That is, you’re supposed to use RDFS to say something like

  ?s :myTermForLength ?o -> ?s :yourTermForLength ?o .

  ?s :lengthInFeet ?o -> ?s :lengthInInches 12*?o .

Living costs are similarly high in many places that have nowhere near the salaries of the US.

It's still the land of opportunities. It's easier to find ways to reduce your living costs than ways to increase your salary.

It's also caused by progressive tax rates. People take harder jobs based on net wage, not gross wage, so gross wage has to compensate.

I think the problem described in 6th paragraph needs to be solved.

arXiv doesn't need much. All they do is host static pdfs uploaded by someone else with free CDN services from Fastly [0]. I'm sure they could get academics to volunteer moderation services as well.

In reality you could host the entire thing for well under $50k/year in hardware and storage if someone else is providing a free CDN. Their costs could be incredibly low.

[0] https://blog.arxiv.org/2023/12/18/faster-arxiv-with-fastly/

[1] https://info.arxiv.org/about/reports/FY26_Budget_Public.pdf

[2] https://info.arxiv.org/about/reports/2020_arXiv_Budget.pdf

That sounds more like an issue of certain fields having crappy standards because the people in those fields benefit from crappy standards than an issue with the site they happen to host papers on.

I don’t buy “some fields are just more honorable”. Everyone uses publishing for personal gain.

But yes it’s a people problem, not an arxiv problem.

The article lists the reasons quite clearly.

For everyone else,

The reason is because arxiv is growing significantly leading to 297,000 deficit in operating costs for 2025 alone. Corenell has helped with donation a long with other organizations that pay membership fees.

As a result, donors + leaders of arxiv think it's best to spin off to increase funding.

We don't do 'utility' in America. Everything has S.V. brain rot - it's mixed with wall street brain rot, and now if you aren't extracting wealth out of what you have access to - you are failing.

> It might already be too late, and we will be left with a browser monopoly.

Ladybird continues to have the appearance of making progress, fwiw:

https://ladybird.org/newsletter/2026-02-28/

No need to ask - the whole point is open access. https://info.arxiv.org/help/bulk_data.html

Lobsters has this I think. But it also means I've never posted there.

not a bad first order filter.

can you think of a better one?

> they'll just turn into a shitty journal at this point

To this end, they added an endorsement requirement this year: https://blog.arxiv.org/2026/01/21/attention-authors-updated-...

This is going to be in NYC where $300k does not go as far as it does in Ithaca.

Heh, you might want to look up what they’re charging young people now.

tangentially related: https://readabstracted.com/

Maybe they'll do something like what Anna’s Archive did

I don’t buy “some fields are just more honorable”. Everyone uses publishing for personal gain.

But yes it’s a people problem, not an arxiv problem.

The net salary in France might be low but the overall cost of hiring is quite high. Besides, why go to the middle when you can just find even cheaper places, if that's your prime metric?

> It's a bit more complex than an S3 bucket

It’s even less. I would bet if it’s not now, for the vast majority of its life it was a machine at someone’s desk at Cornell.

Is this Gondor or Mordor?

Hacker Times

Hacker Times

ArXiv declares independence from Cornell

Discussion

Discussion