服务调研关于联系
← Back to Research
2026-05-22内容架构·

The Paper FactoryA Hundred Years of the Assembly Line for Human Knowledge

Prologue: An Anonymous PubPeer Post That Predated Mr. Geng by Ten Months

The true starting point of this story is not a short-video blogger.

In June 2025, on an English-language academic website called PubPeer, two anonymous researchers posted about a paper that had made it into Nature—the work of a team led by Wang Ping, dean of the School of Life Sciences and Technology at Tongji University: Human HDAC6 senses valine abundancy to regulate DNA damage.

The post was short—a few close-up comparisons of gel-electrophoresis images, with one calm line attached: "There are inexplicable local repetitions in the images."

PubPeer is a peculiar institution in academia: anyone can anonymously flag suspicious points in a published paper, and both the authors and the journal are notified. It is one of the most important pieces of "post-publication peer review" infrastructure of the past decade.

But criticism on PubPeer is in English, tucked away in a specialist corner, and almost no one outside the field knows it exists.

For ten full months, no one picked up the baton.

Until April 2026.


A science-popularization blogger who goes by "Mr. Geng Tells Stories" translated that PubPeer post into Chinese, added his own frame-by-frame dissection video, and posted it simultaneously across Bilibili, Douyin, WeChat Channels, and other platforms (as of May 2026, roughly 1.88 million followers on Bilibili and roughly 1.54 million on Douyin). He enlarged the images, compared pixels, tallied the shapes of the bands, and reproduced exactly the problems those two anonymous researchers had pointed out long before:

He delivered a line that made the whole internet smile knowingly:

"I strongly recommend that Professor Wang use a random number generator the next time he fabricates data."

Right on its heels, Rao Yi, president of Capital Medical University—a molecular biologist known for speaking his mind—published three long essays in quick succession on his WeChat public account "Rao's Science Comments," openly questioning whether two of the Wang Ping team's top-journal Nature papers involved fabrication of raw data. Rao Yi's intervention was the decisive blow that broke the story into the mainstream: the dissection videos on the short-video platforms were grassroots public opinion, but Rao Yi's essays were the academy's own admission.

On May 6, Tongji University issued its notice: Wang Ping was removed from his post as dean and demoted two ranks; the paper's first author, Jin Jiali, had her employment terminated.

Then the dominoes fell: Chen Quan of Nankai University, Kang Tiebang and Kuang Dongming of Sun Yat-sen University, Su Jiacan of Shanghai University—the number of publicly named-and-reported holders of the "Distinguished Young Scholar" / Changjiang Scholar titles reached five.

On May 17, Mr. Geng released another video, announcing that he still held a new batch of fabrication evidence against 5 "Distinguished Young Scholars" across 4 universities—Tongji, East China Normal, Hunan University, and Sun Yat-sen—and issued a public ultimatum: "I'm giving you one chance to inspect and correct yourselves. Refuse, and the next hammer comes down."

What is a "Distinguished Young Scholar"? A recipient of the National Natural Science Foundation's "National Science Fund for Distinguished Young Scholars," it is one of the highest titles a young scholar can hold in Chinese academia—roughly a "reserve corps for academicians." The Changjiang Scholar is the Ministry of Education's counterpart title. Every scholar named here wears at least one such hat.

Nature's official response was a rather cautious line: "We are aware of the criticisms of this paper and are carefully evaluating them."


More telling still is Mr. Geng the person himself.

His real name is Geng Hongwei—bachelor's and master's degrees in biology from Jilin University, and a doctoral candidate at the School of Biomedical Engineering, Beihang University (Beijing University of Aeronautics and Astronautics). In his fifth year, in 2025, he withdrew from Beihang. The version that circulates is that he was "harassed by his advisor, dropped out, and turned around to start exposing fraud."

This is an atypical life trajectory: he is not a "grassroots" outsider to the system—he is a defector from within it. A man the system had ground down for five years, who then voluntarily gave up his doctorate, now turns every ounce of the discernment he learned inside the system back upon the system itself.

At this point, the version the story is most easily told as is "grassroots figure topples a big shot." But the truth was always more complicated:

The truth is a combined force—
   anonymous peers (the English-language challenge on PubPeer)
 + a within-system dropout (Geng Hongwei's Chinese translation and amplification)
 + the academy's clear-eyed few (Rao Yi's public intervention)
 + public opinion (the simultaneous ferment across Bilibili, Douyin, Channels, Weibo, and Zhihu)
 + disciplinary bodies under university PR pressure
Only when these four forces filled the gaps at once did the system's self-cleaning mechanism grind, barely, into motion for a single turn.

Figure 1: The moment the system failed—the combined intervention of four unconventional forces

Figure 1 · PubPeer's anonymous peers (ten months early) + the within-system dropout Geng Hongwei + the academy's clear-eyed Rao Yi + the simultaneous ferment of public opinion—only when these four forces converged did a paper that had hung there as a visibly obvious forgery for ten months finally reach Tongji's disciplinary podium.

This is the fact worth taking seriously. And what it brings out is not the cheap conclusion "the system is entirely broken," but a sharper judgment—

The system's routine self-cleaning mechanisms (peer review, academic committees, administrative correction) have failed so completely that it takes PubPeer's anonymous peers, a dropout doctoral candidate, an outspoken university president, and the force of trending public opinion—these "unconventional forces"—all pushing together just to catch, barely, a visibly obvious forgery.

So the questions truly worth pressing are these three:

First, how could this paper pass Nature's peer review, the national review for a million yuan in funding, and layer upon layer of institutional academic-committee vetting, sailing straight through without obstruction?

Second, why did the anonymous English-language challenge on PubPeer hang there for ten months with no institutional reaction whatsoever from domestic academia—why did it take a short-video blogger shouting it out in Chinese and a veteran scholar lending his name before anything was done?

Third, if fraud of this intensity can hide in the most elite journals for so many years, then of all the "studies show" we read today, how much is actually true?

To answer these three questions, we have to pull the camera back and see clearly how the entire scholarly-paper system was built, where it came from, and where its sickness lies.

This essay is exactly such a manual.

The first half describes its skeleton and bloodline—from the undergraduate thesis to the academician's paper, from Aristotle to the RCT, and how the paper became the central currency of human knowledge.

The second half describes its pathology and its future—the tyrant p<0.05 left behind by the statistical revolution, the tsunami of dilution spawned by Publish or Perish, how papers are now poisoning the very AI models being trained on them, and a global reform that is far from finished.

Mr. Geng only lifted one corner. What we want to see is the two-hundred-year factory beneath the iceberg.


First Half: The Scale of the Paper

——The Ladder of Cognition from Undergraduate to Academician


I. Turning In an Assignment vs. Stepping Into the Arena: An Undergraduate Thesis and a Lecturer's Bid for Promotion Are Not the Same Thing at All

Many people wrongly assume that "the paper" is one unified thing. In fact it is two entirely different activities—different in purpose, audience, judge, and stakes of life and death.

The undergraduate thesis is, in essence, "turning in homework."

Its purpose is not to produce new knowledge but to prove that you have completed one full round of academic training: that you can search the literature, apply a method, and make yourself clear. The judges are your advisor and your school's defense committee; whether it is published is irrelevant, so long as it survives the defense. Its capacity for padding is inherently vast—a great many undergraduate theses are in fact glorified reading notes plus a compilation of source material.

The core-journal paper a lecturer publishes to earn a title is, in essence, "stepping into the arena."

It must undergo anonymous review under the gaze of same-field peers worldwide; it must survive in the strictly ranked journals of SCI / SSCI / CSSCI. Its judges are people you do not know, and they do not know who you are. It is not written for a teacher to read—it is written for the entire discipline. In theory, it must bring an increment of originality to the field, and it must be methodologically rigorous and its data reproducible.

One is a training ground; the other is a public battlefield. The logic of the former is "Teacher, I did my best"; the logic of the latter is "Whole industry, please inspect the goods."

Confusing these two things is the greatest cognitive error of the Chinese undergraduate—imagining that the thing they wrote, which cleared its defense at a 30% plagiarism-check score, is the same species as the stack of SCI papers piled on a professor's desk.


II. From Wielding the Tools to Framing the Question: The Real Thresholds of the Five Grades on the Academic Ladder

Slice the academic career into sections, and the demand each section makes of a paper is, at bottom, an escalating demand on "your relationship to knowledge":

GradeCore TaskPaper RequirementIn One Sentence
UndergraduateComplete academic initiationPass the defenseCan copy a method by rote
Master's studentMaster a field's system of methodsMost schools require 1 core-journal paperCan design research
Doctoral studentMake an original contribution to knowledge2–5 SCI / SSCI papers, at least 1 in a top journalCan find a real problem
Lecturer bidding for associate professorEstablish an independent research directionSeveral core papers within five years + leading a national-level projectCan form a research line
Associate professor bidding for professorBuild disciplinary influenceHighly cited papers + representative works + major projectsCan define the problem
Academician / top-tier scholarDefine the direction of the disciplineTop-journal covers / foundational papersCan found a school of thought

The essence of climbing the ladder: from "measuring things with rulers other people built" to "building a new ruler that everyone else adopts."

Figure 2: The six grades of the academic ladder—from wielding the tools to framing the question

Figure 2 · The six-level progression from undergraduate to academician. The real threshold of each stage is the shrinking distance between you and the "scale of the discipline"—from copying a method by rote to founding a school of thought.


III. The Roots of Two Kinds of Knowledge: The Humanities and the Sciences Differ Not Merely in Subject Matter, but in Epistemology

Many assume the divide between the humanities and the sciences is only a difference in what they study. In fact the difference runs far deeper—the two hold, from the outset, different positions on what even counts as knowledge.

The paradigm of the sciences / natural science: hypothesis-and-verification.

observe a phenomenon → propose a hypothesis → design an experiment → collect data
→ statistical analysis → support or falsify → publish a reproducible conclusion

Its core law is two words: reproducible and falsifiable.

The format is highly uniform (IMRAD: Introduction-Methods-Results-Discussion), mathematized, quantified, objectified.

Knowledge is cumulative; later generations stand on the shoulders of earlier ones, and old theories can be overturned.

The paradigm of the humanities: interpretation-and-construction.

pose a research question → choose a theoretical framework → text / fieldwork / case analysis
→ interpret meaning → build an argument → reach a bounded conclusion

Its core law is rigor of argument and explanatory power.

Quantification is not required; qualitative methods (ethnography, in-depth interviews, close reading of texts) dominate.

Knowledge is not accumulated in one direction—the classics can be reinterpreted again and again, and the Analects can still yield a new paper today.

The social sciences are caught in between. Economics, psychology, and sociology are the arena where the two paradigms wrestle. Over the past twenty years the causal-inference revolution has swept in, and the mainstream has tilted toward quantification, but qualitative research never died—and "Mixed Methods" has instead become a prominent school.

An unfinished truth is this: the essence of the humanities-versus-sciences dispute is whether "the world can be fully measured by a single ruler." The sciences assume it can; the humanities assume it cannot. This argument will never reach an end, because what it asks is, from the start, about the boundary of humanity itself.


IV. A Two-Thousand-Year Relay: Where the Methodology Came From

This methodology of the scholarly paper looks like a cold manual of formats, but behind it lies a relay of thought stretching across two thousand years.

First leg—the logical bedrock of ancient Greece (4th century BC).

Aristotle wrote the Organon, laying down deductive logic: the syllogism. Major premise, minor premise, conclusion. This is the underlying grammar of all scholarly argument. The "theoretical framework + case argument" structure most common in humanities papers today is, at its bones, Aristotle.

Second leg—the experimental paradigm of the Scientific Revolution (17th century).

Bacon stressed induction: observation before theory.

Galileo mathematized nature and made experiments reproducible.

Newton, in his Mathematical Principles of Natural Philosophy, simultaneously demonstrated the ultimate fusion of induction and deduction—this book is, in fact, the mother-template of every "scientific paper" that came after.

From this leg on, "hypothesis—experiment—verification" became the heartbeat of science.

Third leg—the statistical revolution (19th–20th centuries).

This is the pivotal leg in which the modern paper system truly took shape, and the second half of this essay is devoted to it. Galton, Pearson, Fisher, and Neyman-Pearson—four men in relay—built an entire industrial assembly line for quantitative judgment.

Fourth leg—the establishment of social-science methodology (19th–20th centuries).

Comte drove positivism, so that sociology too wanted to imitate physics.

Weber objected, proposing an interpretive sociology that stressed "understanding (Verstehen)."

Durkheim quantified social facts.

Geertz practiced "thick description" in anthropology.

The war of methodologies within this leg—positivist versus interpretive—remains unresolved to this day.

Fifth leg—the paradigm theory of the philosophy of science (20th century).

Popper proposed "falsifiability," drawing a boundary around science: a proposition that cannot be falsified is not science.

Kuhn, in The Structure of Scientific Revolutions, argued that science does not accumulate linearly but undergoes periodic paradigm shifts—the old paradigm collapses, a new one is built.

Lakatos proposed the concept of the "research programme."

When a doctoral student today writes the chapter on their "research paradigm," this is the very vocabulary they use.

Sixth leg—the causal-inference revolution (21st century to the present).

Angrist and others, wielding econometric instruments such as instrumental variables, regression discontinuity, and difference-in-differences, allowed economics and policy science, for the first time, to seriously discuss "causation" rather than "correlation."

Judea Pearl, using the graph theory of causation, proved at the mathematical level that correlation and causation are two different things and can be rigorously distinguished.

The RCT (randomized controlled trial) became the gold standard for medicine, economics, and policy evaluation.

The fusion of machine learning and causal inference is the frontier of the moment.

Figure 3: The two-thousand-year relay of methodology—six milestone legs

Figure 3 · From Aristotle's syllogism to the 21st-century causal-inference revolution, methodology has passed through six legs of a relay—each leg inheriting from the one before, and each completely rewriting the standard for "what counts as knowledge."

Over two thousand years of this relay, the paper has become more than a genre—it is a precision machine for domesticating human curiosity.


Second Half: The Crisis of the Paper

——Dilution, AI, and an Unfinished Revolution


V. Four Men Fire Up the Furnace: The Hidden Founders of the Statistical Revolution

If the paper is the central currency of modern knowledge, then the statistical revolution of the 19th and 20th centuries was the mint that struck that currency. Four men—Galton, Pearson, Fisher, and Neyman-Pearson—teacher succeeding teacher, hating one another, yet together completing the foundation.

Galton (Francis Galton, 1822–1911): The First Brick

Darwin's cousin. Spurred by evolutionary theory, he set out to measure "heredity."

He recorded the height data of 928 adult children and 205 pairs of parents (for narrative simplicity, the discussion below uses "father-son height" as its example; in his raw data he multiplied daughters' heights by 1.08 to "convert" them, and merged the parents into a single "mid-parent" metric—an operation that has itself remained controversial throughout the history of statistics).

He drew a scatter plot and asked: if a father is extremely tall, must the son be extremely tall too?

He found the answer was no. When the father was extremely tall, the son tended to be a little shorter; when the father was extremely short, the son tended to be a little taller. Both were drifting toward the mean.

He named this phenomenon "regression to the mean."

—and this is the origin of the word "regression." The "linear regression" in the first chapter of every machine-learning textbook today traces its name straight back to here.

At the same time he intuitively proposed the concept of the correlation coefficient: the degree to which two variables "move together." But his mathematics was not strong enough to write it down rigorously.

He was the observer who opened the game, waiting for a student to come and polish the tools.

(A footnote: Galton and his student Pearson were both fervent promoters of eugenics—a dark chapter in the history of statistics that has been examined again and again; UCL issued a formal public apology in the 2020s and is still reckoning with it. The birth of statistics is deeply entangled with racist ideology, and that alone is worth a separate essay.)

Pearson (Karl Pearson, 1857–1936): Building the Precision Instrument

Galton's student, trained as a mathematician, hard in temperament.

He did three things:

① He wrote down a rigorous mathematical definition of the correlation coefficient—today called the Pearson correlation coefficient r. It ranges from -1 to +1. The formula has been used from 1900 to today with almost no change.

② He invented the chi-squared test—which answers one core question: is the difference between the observed frequency distribution and the distribution that "ought" to hold in theory too large?

χ² = Σ (observed - expected)² / expected

Flip a coin 100 times and get 60 heads—is there something wrong with this coin? The chi-squared test tells you.

③ In 1911 he founded the world's first Department of Statistics at University College London (UCL), initially named the "Department of Applied Statistics," raising statistics from the "concubine of mathematics" to a discipline in its own right.

But Pearson had a fatal blind spot: he firmly believed science could discuss only "correlation," never "causation." In his view, causation was metaphysical nonsense.

This blind spot shackled statistics for a full 100 years, until Judea Pearl finally corrected it with the graph theory of causation in the 2000s.

Fisher (Ronald Fisher, 1890–1962): The Sovereign Takes the Stage

Pearson's sworn enemy. The two hated each other all their lives, yet together they laid the foundation of modern statistics.

Fisher's contribution runs deepest, in three layers:

① Analysis of Variance (ANOVA)

Three fertilization schemes—which is best? The traditional approach is a pairwise t-test, run 3 times. Fisher said no—that way you accumulate error.

His method was to test all groups at once—decomposing the total variation into "between-group variation" and "within-group variation," their ratio being the F value; the larger the F, the more significant the treatment effect.

This method is today the standard equipment of psychology, medicine, and agricultural experiments.

② The three principles of the randomized controlled trial (RCT)

While working at the Rothamsted agricultural experiment station, he found experimental design in chaos—some plots of land were good, some poor, and there was no way to tell "fertilizer effect" apart from "difference in land."

He gave three principles: randomization, replication, blocking.

The idea: randomly assign treatments, so that all "unknown confounders" are automatically distributed evenly in a statistical sense and thereby eliminated.

This body of thought later evolved directly into the RCT of medical clinical trials, becoming the gold standard of today's evidence-based medicine.

③ The p-value—the number that ruled the world for 90 years

In 1925, Fisher wrote this passage in Statistical Methods for Research Workers

"The value for which P = 0.05, or 1 in 20, is 1.96 or nearly 2; it is convenient to take this point as a limit in judging whether a deviation is to be considered significant or not."

Note the word he himself used: "convenient."

The line p<0.05, under his pen, was from the start a threshold-tool adopted "for convenience," not a final verdict. He himself used 0.01, used 0.02, and never once said 0.05 was sacred and inviolable.

But later scholars learned only the sentence "P=0.05, 1 in 20."

p < 0.05 → "significant" → published → written into the textbook → truth
p ≥ 0.05 → "not significant" → tossed into the drawer → gone forever

A "convenient" convenience-threshold was crowned by posterity as science's line of verdict. Fisher realized late in life that this was a disaster, but by the time he died in 1962, the dictatorship of p<0.05 was fully established.

The Neyman-Pearson Framework: Sharpening the Knife into a Framework

Neyman (Jerzy Neyman) + the younger Pearson (Egon Pearson, Karl Pearson's son), in the 1930s.

They plugged a hole in Fisher's system: the p-value only says "the probability of seeing this set of data if the null hypothesis is true," but says nothing about "what decision I should make."

They introduced a pair of terms:

H₀ (null hypothesis): assume there is no effect
H₁ (alternative hypothesis): assume there is an effect

Two kinds of error:
  α (Type I error) = judging there is an effect when there is none (false positive)
  β (Type II error) = judging there is no effect when there is one (false negative)

Statistical power = 1 - β

They turned "whether to run the experiment" into decision theory: you first set an acceptable error rate, then act by the rules.

Fisher and Neyman traded public insults for thirty years. Fisher said Neyman had turned science into "factory decision-making," betraying the spirit of the pursuit of truth; Neyman said Fisher's logic was not rigorous at all.

The irony is that today's statistics textbooks forcibly stitch the two systems together, and what students memorize is a hybrid whose internal logic contradicts itself. That contradiction is one of the hidden sources of the statistical crisis to come.

Figure 4: The four men who fired up the furnace of statistics—the web of mentorship and grudges

Figure 4 · Galton → Pearson (master and student) → the younger Pearson (father and son); Pearson ⟷ Fisher (30 years of mutual invective); Fisher ⟷ Neyman (public invective). The seemingly seamless "statistics" of today's textbooks is, in essence, the product of posterity forcibly stitching together this web of grudges.


VI. p<0.05: A Deified "For Convenience"

We pull this matter out to discuss on its own because it touches the very lifeline of all modern science.

Figure 5: p<0.05—two worlds on either side of a dividing line

Figure 5 · The same experiment, the same people, the same method—cross that 0.05 line and you enter the assembly line of "truth"; fail to cross it and you enter the "file drawer" and vanish. This line is the "convenient" that Fisher jotted down offhand, deified by posterity into science's verdict.

What are the real consequences of p<0.05?

First, publication bias. Journals love only "significant" results and never publish the non-significant. The result: the drawers are stuffed with "truth," and what is published is only the portion that "looks significant."

Second, p-hacking (manipulation of the p-value). Researchers repeatedly re-slice samples, add variables, and try models until the p-value crosses 0.05. This is not deception—often it is unconscious, because the human brain rationalizes to itself by nature.

Third, HARKing (Hypothesizing After the Results are Known). What was plainly a pattern stumbled upon by data mining is written up in the paper as "we hypothesized that…" A thoroughly exploratory study is packaged as a confirmatory one.

Fourth, statistical significance ≠ practical significance. With a large enough sample, almost any tiny difference can reach p<0.05, yet such a difference may be utterly meaningless in reality.

In 2016, the American Statistical Association (ASA) took the rare step of issuing an official statement formally acknowledging that the p-value has been seriously abused, and recommending against using p<0.05 as the standard for judging a scientific conclusion.

In 2019, more than 800 scientists co-signed a piece in Nature calling for the abolition of the very concept of "statistical significance."

But ninety years of inertia will not vanish overnight.


VII. Publish or Perish: The Structural Pathology of Global Dilution

Academia has a piece of jargon: Publish or Perish.

This is the assessment philosophy that began in the United States in the 1960s and was imported on a massive scale into China in the 2000s: it binds the sheer number of papers directly to titles, salaries, projects, household registration, and one's children's school admission. The result is to turn "publishing" from a means into an end in itself.

Dilution thereby becomes inevitable—and it is structural:

An explosion in the number of journals. In 1987 there were roughly 70,000 academic journals worldwide (Meadows's estimate); by 2020 there were roughly 40,000–50,000 active peer-reviewed journals; count predatory journals and semi-scholarly publications together, and the number far exceeds this. Among them, predatory journals are estimated to already account for 10–30% of academic-publishing volume—they publish anything so long as you pay the page fee, and do no peer review. The total number of journals has grown exponentially since World War II, and this is the most direct supply-side soil of dilution.

The Least Publishable Unit culture (LPU). A study that could have been written as one good paper is sliced into five "just-publishable little papers," each innocuous and painless, but the count looks good.

Citation rings and mutual-citation networks. Journal editors require authors to cite the journal's own articles; research teams sign "mutual-citation agreements"; the citation counts of certain scholars are artificially inflated to astonishing levels.

The replication crisis. Since 2011, psychology has begun replicating classic experiments on a large scale. The landmark study by the Open Science Collaboration (OSC), published in Science in 2015, found that of 100 classic psychology experiments, the fraction that could be replicated to a statistically significant result in the original direction was only about 36% (up to 47% under a looser standard). Economics, medicine, and neuroscience saw similar crises erupt in turn.

China's particular curve:

2000-2015: quantity surges, quality uneven
2015-2020: highly cited papers grow rapidly, some fields rank among the global top
                (AI, materials, chemistry, quantum information)
2020-now:  the Ministry of Education's "break the five onlys" policy begins to correct course
           (break away from paper-only, hat-only, title-only, degree-only, award-only)
           but actual enforcement is uneven

The whole system has by now formed a kind of self-reinforcing incentive to pad—the more you publish, the more resources you get; the more resources you get, the more you publish. A clear-eyed few quit the game; those who remain keep accelerating.

Return to the Wang Ping case at Tongji from the prologue—what it toppled was not a few individuals but a cross-section of this system: fraud that could be seen through with the naked eye, that had hung on PubPeer for ten months, could still be published in Nature, walk off with a million yuan in funding, and be crowned with a dean's hat. This proves that this screening mechanism has grown so sick it can no longer identify even the most basic forgery, and it takes the combined intervention of forces outside the system (PubPeer's anonymous peers), on the system's edge (a short-video-platform blogger), and the clear-eyed few within the system (Rao Yi) just to catch, barely, one exception.


VIII. Backflow Contamination: How Papers Poison AI

This is the most cutting-edge, and least publicly discussed, topic of the moment.

The composition of the pretraining data for large models runs roughly like this (estimated from the few models—GPT-3, Llama, and the like—that have publicly disclosed their training recipes; each company's actual recipe is a trade secret, so this is for reference only):

Common Crawl (web crawl)      ~60-70%
GitHub code                   ~10-15%
academic papers (arXiv, etc.) ~5-15%
books                         ~10%
Wikipedia                      ~3%
other high-quality sources     ~5%

The numbers make papers look like a small share, but their weight far exceeds their proportion—papers are often upsampled during training, because they contain humanity's most densely packed chains of explicit reasoning.

Conclusion: paper quality directly affects the model's quality of reasoning, not just the breadth of its knowledge.

So how, exactly, do junk papers poison AI?

First, contamination by false facts. False conclusions derived from p-hacking, "psychological common sense" over-generalized from small samples, "classic experiments" that cannot be replicated—once AI learns them, it recites them back in the tone of academic authority. Early GPT, for instance, described the Stanford Prison Experiment with high confidence, but after Le Texier and others published the declassified recordings and interview materials in 2018, that experiment has been seriously challenged by the field—Zimbardo's instructions to the guards were far more explicit than the official narrative, and the participants were partly performing; its academic standing today has retreated from "classic conclusion" to "ethically compromised research whose conclusions are seriously overstated."

Second, contamination of writing style. That kind of hollow, formulaic phrasing—"This study aims to explore the effect of X on Y; the results show… with important theoretical and practical significance"—has been ingested in bulk, and the model itself begins to write this way too, using complex vocabulary to cover hollow content.

Third, the deep source of hallucination. AI's hallucinations are often blamed on "problems with the model itself," but a considerable portion of them originate in the chains of error propagation formed by papers citing one another in the training data. One paper gets a fact wrong; it is cited by fifty papers; and in the end it becomes "fifty-one sources all say so," and what the AI sees is "a high-confidence consensus."

This is a hidden but profound feedback loop—the dilution produced by academia is turning into an invisible deficit in AI's reasoning ability.


IX. The Problem of Knowing What One Does Not Know: Can AI Identify Junk Papers?

Partly, yes—but far from solved, and with a fundamental difficulty.

What AI can currently do:

What AI currently cannot do:

"How large, exactly, is this paper's incremental contribution?"
→ requires a domain expert's judgment; AI has no real understanding.

"Was the experiment actually carried out?"
→ AI cannot see the laboratory and cannot verify the raw data.

"Is this conclusion a major breakthrough in the field?"
→ requires knowing "where the field's current boundary lies"—and AI's knowledge has a cutoff date.

The most fundamental difficulty is this: the core capability for judging a paper's quality is "knowing what one does not know."

And that is precisely AI's weakest link right now.

The irony is that the most effective "fraud detector" at present is still PubPeer's anonymous peers + a person like Geng Hongwei who is patient, has domain knowledge, and stands outside the system's chain of interests + a clear-eyed member of the academy like Rao Yi who dares to intervene publicly. Technology can assist, but it cannot replace them.


X. The Unfinished Revolution: A Map of Reform

Academia has not been idle. Over the past fifteen years, five clear paths of reform have formed in response to the dilution crisis:

① Pre-registration

Before reform: run the experiment → look at the results → write the paper → select the "significant" ones to publish
After reform:  first publicly declare "what experiment I will run, what hypothesis"
       → run the experiment → publish regardless of the outcome

This cuts the roots of p-hacking and HARKing directly. The Open Science Framework (OSF) has by now hosted over 100,000 pre-registered studies, with psychology and medicine advancing fastest.

② Registered Reports

More radical still—the journal decides whether to accept before the experiment is even completed:

The researcher submits: research question + method design (no results)
The journal reviews: is the method good? Is it worth doing?
The journal promises: publish whether the result is positive or negative
The researcher runs the experiment
Publication (results written in)

Nature Human Behaviour and PLOS ONE have already introduced it. But its spread is slow, because journals gain no "exclusive blockbuster" benefit from it.

③ Open Data and Reproducibility Standards

When a paper is submitted, it must come with the raw dataset and the analysis code. Anyone can run it on their own computer to verify the results. arXiv + GitHub have become standard equipment in physics, mathematics, and computer science. In biomedicine, the FAIR principles (Findable, Accessible, Interoperable, Reusable) are being advanced.

④ Post-publication Peer Review

Traditional review happens before publication—inefficient and opaque. The new model moves review to after publication:

The Wang Ping case is the most dramatic real-world demonstration of this reform path: when traditional review has lost the ability to detect fraud, the combination of PubPeer + Chinese public opinion + the academy's clear-eyed few in fact constitutes a systemic "post-publication peer review" mechanism. It is not elegant, not formal, and it depends on individual heroism to fill the gaps—but it does work.

⑤ Bibliometric Reform—Beyond the Impact Factor

The Impact Factor, that outdated ruler, is being thrown out:

The DORA Declaration (2013) has by now been signed by 22,000+ institutions and individuals, publicly pledging no longer to use the impact factor to evaluate individual research.

Figure 6: The map of reform—five paths

Figure 6 · All five paths are trying to answer the same question—how to realign "publishing" with "the pursuit of truth." Among them, "post-publication peer review" (PubPeer / eLife) is exactly the real-world demonstration ground of the prologue's Wang Ping case.

This reform has a direction but no end. Its greatest resistance is not technical but incentive-based—no one is willing to be the first to give up the game rule of "number of papers published."


XI. Hard Bones and Soft Bones: A Ranking of Gold Content

Having covered the crisis and the reform, let us return to a plainer question: which papers today still command genuine respect?

The plain answer is: engineering and the hard sciences.

This is not prejudice but a matter of structural reasons.

First, the standards are harder.

Second, the cost of failure is extremely high.

Behind an engineering paper are products, patents, factories. Fabrication leads directly to engineering failure, the price of which is concrete: a product flops, an investment rots, a bridge collapses.

A map of gold content (based on reproducibility and community consensus):

Extremely high gold content (engineering + hard sciences):
  · Computer systems / top ML conferences (NeurIPS, ICML, CVPR)
     open-source code, reproducible results, rapid community verification
  · Materials science (Nature Materials, etc.)
     new-material performance is hardcore; fabrication is suicide
  · Physics (PRL, Nature Physics)
     particle physics routinely has thousands of co-signing authors; the whole body vets the data
  · Chemical synthesis (JACS, Angewandte Chemie)
     synthesis routes and spectra cannot be forged
  · Astronomy / cosmology
     data come from telescopes, shared globally, cross-verified globally

High gold content (applied + clinical):
  · Medical RCTs (NEJM, Lancet)
     large-sample, multi-center, high credibility
  · Structural biology
     protein structures and the like can be independently verified by third parties

Medium gold content (behavioral social science):
  · Economics (with RCTs or natural experiments)
  · Neuroscience (the replication crisis exists, and is improving)

Gold content in question:
  · Small-sample psychology (the WEIRD-sample problem)
  · Some humanities core journals
  · Fringe fields overrun by predatory journals

One unpleasant but necessary sentence must be added: the life sciences are an exception. They wear the badge of "hard science," but in reality they belong to a high-incidence zone of dilution—cell images, gels, and microscopy images are all extremely easy to Photoshop, replicating an experiment routinely costs hundreds of thousands of yuan, and their reproducibility is far below that of the hard sciences such as physics and chemistry. A 2025 Nature tally noted that of the world's top ten institutions by retraction volume, seven were hospitals or medical schools in China. The field on which Mr. Geng concentrated his fire is precisely the life sciences—this is no coincidence, but the outward opening of a structural loophole in the system.

Computer-science papers deserve a sentence of their own.

Over the past twenty years, a large share of the world's most influential papers has come from computer engineering:

2017 "Attention Is All You Need" (Transformer)
     nearly 200,000 Google Scholar citations; directly ignited the AI revolution

2012 "ImageNet Classification with Deep CNNs" (AlexNet)
     year zero of deep learning

1998 "Gradient-Based Learning Applied to Document Recognition" (LeNet)
     the starting point of convolutional neural networks

These papers share one common feature: the code can be reproduced, the benchmarks are public, industry can verify them within 6–12 months, and the academic community reproduces and improves them on a large scale within 1–2 years.

They are trustworthy not because a reviewer stamped them, but because the entire community reproduced and improved them in a very short time.

This is what a paper ought to be—an open, verifiable baton of cognition that pushes civilization forward.

Figure 7: The map of paper gold content—a four-tier pyramid

Figure 7 · Three plain dimensions for judging gold content: Is the data/code public? Is the cost of error concrete? Can the community reproduce it quickly? Yes to all three = extremely high gold content. The life sciences are a high-incidence zone of dilution wearing the badge of hard science, and require separate vigilance.


Conclusion: A Few Principles for the Clear-Eyed Use of Papers

Having written this far, let me leave the reader a few plain reminders:

One: when you see the words "studies show," first ask—which study? How large a sample? Which journal? Default vigilance is the starting point of clear-sightedness.

Two: when you see p<0.05, do not take it as gospel at once. Behind that number is Fisher's footnote of "convenient"—it is not as sacred as it wants you to believe.

Three: when you see the word "significant," distinguish "statistically significant" from "practically significant." The former is mathematics; the latter is life.

Four: the "classic conclusions" of psychology and sociology should be discounted by at least twenty percent. The replication crisis tells us that half of the "common sense" of the past few decades may not hold up.

Five: keep an extra measure of suspicion toward image-type data in the life sciences. This is the most concrete lesson the Tongji Wang Ping case leaves us.

Six: papers in engineering and the hard sciences are relatively trustworthy, because the cost of fabrication is concrete and brutal. This is not to say they contain no errors, but that the errors get exposed quickly.

Seven: the answer AI gives you is, in essence, the average product of the past few decades' pool of papers—and papers carry far more weight than their volume-share, because they are usually upsampled in training. However clean the pool is, that is how clean the answer will be.

And finally, the most important one:

Do not mistake "published" for "proven true." The paper system is one of the best truth-seeking tools humanity has so far, but it is not truth itself. It is an industrial assembly line—it produces defective goods, it gets abused, and it also incubates the genuinely great.

What is admirable in the Wang Ping case was never any single hero.

It was the two anonymous English-language challenges that hung on PubPeer for ten months, ignored;

it was Geng Hongwei, who gave up his doctorate in his fifth year and turned around to shout it out anew in Chinese;

it was Rao Yi, president of Capital Medical University, who dared to publish three signed essays in a row on his public account;

it was the retweets, comments, and questions that surged overnight on the trending lists.

Only when four unconventional forces filled the gaps at once did the system's routine self-cleaning mechanism grind, barely, into motion for a single turn. And this itself is the greatest reminder of our age: when a system's built-in immunity has failed to this degree, every clear-eyed outsider is, in fact, that system's last line of defense.

Use it clear-eyed; do not be defined by it. This is the most basic requirement this age makes of a modern person with self-knowledge.


Reference sources:

Share
← Back to Research

Related · TryWay Labs

長為試之印(盖印版·自然崩口)