arXiv Leaves Cornell After 35 Years to Go Independent
The nonprofit platform hosting cutting-edge AI and physics research is severing its 35-year tie to Cornell University.

Opening
Dear reader, have you ever come across the name arXiv while searching for research papers? It’s the place where you can encounter the latest research in AI, physics, and mathematics before anyone else. From the technology underlying ChatGPT to the frontiers of quantum computing, the leading edge of science is now surfacing first not in academic journals, but on this preprint1 server.
Launched in 1991 by physicist Paul Ginsparg at Los Alamos National Laboratory, the service now hosts more than 2.4 million papers free of charge, with monthly downloads reaching 10 million. And yet, this massive knowledge infrastructure is run by a staff of just 27 people, on an annual operating budget of only $6.7 million (~₩9 billion).
News broke that on July 1 this year, arXiv will leave Cornell University—its home for the past 25 years—and transition into an independent nonprofit. This is not a simple organizational relocation. Behind this decision lie two structural pressures created by the AI era.
Walking a $6.7 Million Tightrope — Why Is arXiv Leaving?
arXiv’s independence isn’t a sudden decision. Founder Paul Ginsparg himself has recommended this direction for years. The core reason is money. arXiv has posted operating deficits in both 2024 and 2025. The 2025 deficit alone was $297,000 (~₩400 million). Cornell covered the shortfall and threw in an additional $819,000 (~₩1.1 billion) of in-kind support—but from the university’s standpoint, arXiv was always a line item competing against other budget demands.
The root cause of the problem is explosive growth. Submissions have surged more than 50% since 2022, and 2026 is expected to surpass 300,000 papers a year. There’s even a record of 26,000 papers received in the single month of September 2025.
This growth has two drivers.
First, the explosion of AI research itself. The number of papers in AI and machine learning is growing exponentially. Thomas Dietterich (Oregon State University), arXiv’s editor-in-chief for computer science, named this as a direct driver of growth.
Second, papers made by AI—so-called “AI slop.” Low-quality papers generated by LLMs are pouring into arXiv. According to Ralf Beers, chair of the editorial board (University of Amsterdam), submissions of AI slop have grown exponentially since early 2025, and the rejection rate has spiked from 4% to 10-12%. A study published in Nature Human Behaviour reported that, as of September 2024, roughly a quarter of computer science abstracts showed traces of LLM editing.
Within the walls of Cornell, it was hard to withstand these two pressures. The university couldn’t hire enough software developers, nor could it freely raise outside donations. Ginsparg himself acknowledged this—that universities, as institutions, have no track record of sustaining this kind of global research infrastructure over the long haul.

AI Slop — Science’s Own Tool Now Threatening Science’s Speed
The AI slop problem deserves a closer look, because it isn’t unique to arXiv—it’s a structural crisis for scholarly communication as a whole. A study published in Science in 2025 by a Cornell University team led by Yian Yin illustrates this well. Analyzing more than 2 million papers across major preprint platforms, the team found that publication volume among researchers estimated to be using LLMs increased by 33-50% or more. Among non-English-speaking researchers in particular, the increase reached 43-89%. Productivity did rise—but there’s a catch: over the same period, the cost of telling meaningful research apart from mass-produced content also shot up.
arXiv has responded to this problem in stages. In October 2025, it effectively banned submissions of review papers and position papers2 in the computer science category. In January 2026, it expanded its endorsement system for first-time submitters to all fields—previously, anyone with an institutional email address could submit, but now submitters need endorsement from an existing arXiv author in the relevant field.
But here’s the dilemma. arXiv’s core value is “openness that lets anyone share research quickly.” Raising the gate will cut down on AI slop, but it also raises the barrier to entry for emerging researchers. AI safety researcher Stephen Casper pointed out that this policy could disproportionately affect early-career researchers, researchers with limited computing resources, and researchers in ethics and governance fields.
The tension between openness and quality—this is a problem every platform faces in the AI era. From Wikipedia to X (Twitter), every platform open to anyone stands before the same question.
The Great Preprint Migration — Not Just arXiv’s Story
What’s interesting is that arXiv’s independence is not an isolated event. In March 2025, bioRxiv (life sciences) and medRxiv (medicine) left their parent institution, Cold Spring Harbor Laboratory, to establish an independent nonprofit called openRxiv. The reasons are nearly identical to arXiv’s—expanding donations and modernizing technology is difficult within the walls of a university or research institute. openRxiv is sailing along smoothly, having taken in more than 64,000 preprints in its first year.
What’s unfolding now is a kind of “Great Preprint Migration.” The core infrastructure of scholarly communication is leaving the embrace of universities and research institutes to reorganize as independent nonprofits. Behind this trend lies a shared recognition:
Global knowledge infrastructure is not a side project of a single institution—it’s an independent mission that requires a dedicated organization.
Oz’s Lens
As I read this story, one comparison kept coming back to me. $6.7 million versus $3.2 billion. arXiv’s annual operating budget is $6.7 million (~₩9 billion), while RELX—the parent company of Elsevier, the largest academic publisher—posted an operating profit of £3.2 billion (~₩5.8 trillion) in 2024. Elsevier’s operating margin comes in at 38.4%—on par with Microsoft or Google.
From a GTM strategy perspective, this is a bizarre market. Researchers do the research with public funding, other researchers peer-review it unpaid, and universities pay subscription fees to access it all over again—every core piece of labor in the value chain is done for free, while the middleman platform pockets a margin approaching 40%.
arXiv is the oldest alternative to this structure. It made papers freely available and bypassed the slow clock of academic publishing to accelerate the pace of research. But it’s ironic that this very alternative has always struggled with funding shortages. I think that for arXiv’s independence to succeed, designing a sustainable revenue model matters more than the legal form of “nonprofit” status. The current setup—relying on membership dues from roughly 270 institutions (up to $10,000 a year each) and foundation grants—is too thin to cover operating costs in an era of 300,000 papers a year. Sure, you could object, “how dare you reduce sacred research to a matter of capital!”—but that sounds a bit too much like a fairy tale to me… lol.
Let me be honest about one concern, too. The worry circulating among scientists about independent nonprofits—“won’t this just end up getting commercialized?”—isn’t entirely baseless. Looking at the history of academic publishing, the pattern of nonprofit-founded platforms being acquired by for-profit companies under funding pressure has repeated itself. The controversy over the new CEO’s $300,000 salary comes out of this very context. In the end, the core question is this: while keeping the distribution of knowledge a public good, who pays for it? arXiv’s independence is a 35-years-in-the-making experiment testing this question. Until the answer arrives, it’s worth all of us watching closely.
Closing
- arXiv becomes independent from Cornell on July 1, but this isn’t an organizational relocation—it’s the surfacing of the question, “who maintains global knowledge infrastructure?”
- The dual pressure created by AI—the research explosion and AI slop—has exposed the limits of the existing operating model.
- The chain of preprint servers going independent signals a structural shift in the governance of scholarly communication, moving from universities to dedicated nonprofits.
Here’s one thought I’ll leave you with. The AI research we read every day is first published on a platform that runs on ₩9 billion a year—about the size of one or two tech-startup seed rounds. Are we thinking enough about the cost of keeping knowledge moving at speed?
References & Further Reading
- Jeffrey Brainard, “ArXiv, the pioneering preprint server, declares independence from Cornell”, Science, 2026. : This is the central article behind today’s newsletter. The interviews with Ginsparg and Morisset are the highlight.
- Keigo Kusumegi, Xinyu Yang, Paul Ginsparg et al., “Scientific production in the era of large language models”, Science, Vol.390, 2025. : A study analyzing the impact of LLMs on academic productivity across roughly 2 million papers.
- Jeffrey Brainard, “ArXiv preprint server clamps down on AI slop”, Science, 2026. : An article covering the background behind arXiv’s introduction of the endorsement system. Useful for grasping the scale of AI slop.
- Ranjit Singh, “On arXiv, an Influx of AI Slop Pits Surface Against Substance”, Data & Society, 2025. : A piece analyzing the structural impact of AI slop on academic quality-control systems.
- Cornell Tech, “arXiv Transition to Independent Nonprofit”, 2026. : The official announcement page for arXiv’s transition to independence. It lays out the CEO job posting and governance structure.
- openRxiv, “2025 Year in Review”, 2025. : A good resource for understanding the independence case of bioRxiv/medRxiv.
- Karl Huang, “Academic publishing is a multibillion-dollar industry. It’s not always good for science”, The Conversation, 2026. : Background for understanding the structure of the academic publishing market and Elsevier’s revenue model.

The author, Kwangseob Ahn, is a professor of business administration at Sejong University and lead consultant at OBF (Oswarld Boutique Consulting Firm). He teaches statistics and data analysis — business data management and business analytics — while leading GTM and AI strategy consulting in the field, designing the seam between technology and business. He has published academic research on a memory architecture for AI dialogue systems (HEMA) and runs Daily Arxiv, a daily curation of global AI papers. He holds a master’s from Korea University’s Graduate School of Technology Management and a KMBA. He is the author of Homo Brainless: The People Who Outsource Their Thinking.
Footnotes
-
Preprint: A paper made available online before it goes through a journal’s official peer review. Think of a restaurant analogy—it’s similar to a chef putting out a tasting dish before it ever makes it onto the official menu. Fast knowledge-sharing is the core value here. ↩
-
Position Paper: An academic document laying out an author’s stance or views on a specific topic. Unlike papers that report new experimental results, position papers build arguments on top of existing research—which makes them a type that’s relatively easy to generate with an LLM. ↩
Your take shapes the next issue
What resonated most in this issue, or where has your experience been different?