What Is Actually In physics.gen-ph
Everyone in physics knows what physics.gen-ph is for, in the way everyone knows things nobody has checked. It’s the general-physics category, and it functions as a soft landing: arXiv’s moderators don’t reject your paper, they file it somewhere nobody looks. Perpetual motion, one-page proofs that Einstein was wrong, the usual.
AI-generated illustration.
That’s the folklore. I wanted the number.
Disclosure, so you can weight what follows. I have a paper sitting in gen-ph. I have opinions about how it got there, and I’ll write about that separately. This post is about the corpus, not about me — and I’ve published the scripts and the raw data at the bottom so you can run it without taking my word for anything.
The method
arXiv’s metadata carries a journal-ref field. If a preprint later appears in a journal, the author can record it there — and many never bother, so it’s a floor rather than a count. But the floor sits at the same height for everyone, which makes comparisons between categories meaningful even when the absolute number isn’t.
I took every gen-ph and every gr-qc paper submitted between 2020 and 2022. Not a sample — the complete census: 869 papers and 17,897 papers respectively. Then I asked three questions.
Number one: do they get published?
Yes, and this surprised me.
44.6% of gen-ph papers acquire a journal reference. For gr-qc it’s 51.8%.
If gen-ph were the junk drawer of folklore, that figure should be in the single digits. Nearly half of what lands there is subsequently published somewhere by somebody. Whatever the routing is detecting, it isn’t unpublishability.
Number two: where do they publish?
Here the folklore recovers.
Of papers that publish, the share reaching the field’s flagship venues — Physical Review D and Letters, Classical and Quantum Gravity, JHEP, JCAP, Physics Letters B, MNRAS, ApJ — is 56.2% for gr-qc and 4.9% for gen-ph. Elevenfold. MDPI-family journals run the other direction, 11.6% against 5.2%.
So the honest one-line summary is that gen-ph routing barely predicts whether your work survives peer review, and strongly predicts where it lands.
Which is a good result for the moderators, and I want to put their case as strongly as they would: gen-ph papers publish, but not where it counts, and that is precisely what you’d expect if the routing tracks something real about quality.
It might. But notice that from outside you cannot tell which way the arrow points. A paper in gen-ph is unread, and unread work doesn’t reach flagship journals. So the routing could be diagnosing the outcome or producing it. Prophecy and diagnosis are indistinguishable without the counterfactual, and the only party holding the data that would separate them is the party under evaluation.
I can’t resolve that. But there’s a third measurement that constrains it.
Number three: which came first?
Normal scholarly order is preprint, then journal. You post to arXiv, then you submit, then months later it appears. The reverse ordering — arXiv deposit after journal publication — is rare, and it’s a fingerprint of a specific route: the author didn’t get onto arXiv the usual way, published first, and came back afterwards.
So: how enriched is gen-ph in that fingerprint?
Of published gen-ph papers, 10.1% were deposited one to three years after their journal publication. In gr-qc, 1.9%.
That’s 5.2-fold. And it isn’t people archiving ancient work — the decades-late tail is small in both and doesn’t account for it, 2.1% against 0.5%. One to three years is the publication cycle. These are papers that went to a journal first and arrived at arXiv holding a receipt.
Ten percent is a floor twice over, incidentally. It counts only papers carrying a journal-ref at all, and year granularity makes same-year readmissions invisible.
What that implies
gen-ph is not simply a low-prestige category. A measurable slice of it is a holding pen — work that was turned away at the door, went and got published, and came back.
That matters for number two, because it partly dissolves it. If a real fraction of gen-ph is readmitted work, then some of that 4.9% flagship rate just reflects where a paper can publish when it has no preprint to point at. Which is a treatment effect wearing the costume of selection.
There’s a tidy loop in that: the routing produces the evidence that justifies the routing.
And one more figure, the sharpest in the set. Of the 39 readmitted gen-ph papers in this window, the number that reached a flagship venue is zero. Not low. Zero. They went to EPJ D, IJTP, Modern Physics Letters A, a Mexican society bulletin, and a few venues whose names I’d describe as aspirational.
Caveats, and one I went and tested
journal-ref is author-reported, and reporting habits may differ between categories. That’s the one I can’t fix from outside. “Flagship” is my keyword list, not a ranking body’s. The window is three years. And there’s a reading I can’t exclude: gen-ph may simply attract authors further from the mainstream who find arXiv late, for reasons having nothing to do with rejection.
One objection I can answer, because I went back and checked. My first pass sampled gr-qc at 200 papers a month rather than taking all of them, and a capped, non-randomly-ordered sample is exactly the sort of thing that manufactures a result. So I ran the full census — all 17,897 — and every figure moved by less than a percentage point. The sample was fine. I mention it because I wouldn’t have believed me either.
What I’d like someone to do next
Three things I can’t do from here, in rising order of how much I want them:
Resolve the same-year cases. Journal months are recoverable through Crossref, which would turn the 10.1% floor into an actual estimate, and I suspect the actual estimate is considerably higher.
Extend backwards. Three years is enough to see the effect and not enough to see whether it’s growing.
And the one that would settle it: citation trajectories for matched pairs. Take gen-ph papers that published in a given journal, match them to gr-qc papers in the same journal and year, and compare citations. If the gen-ph twin is systematically less cited despite the identical venue, that’s the treatment effect isolated, and the prophecy-versus-diagnosis question stops being unanswerable from outside.
Scripts, raw data, and the full census are here: github.com/beastraban/WhatsInGenPhys. If you think this is wrong, everything needed to demonstrate it is in that directory. I’d rather be corrected in public than right in private.
Dr. Ira Wolfson is a physicist and Senior Lecturer at Braude College of Engineering, Israel. He works on thermodynamics, Bayesian epistemology, and the philosophy of science.