AI & Decision Making Research Group
Paper Radar2026-07-28

🔮 How Will Research, Papers, and Peer Review Change? — Short-, Medium-, and Long-Term Predictions by Field

AI / Academia / Peer Review / Predictions

Written: July 27, 2026 Prerequisite reading: How Far Has AI Penetrated Research? A Field-by-Field Survey

A note on the nature of this document Chapters 1–2 organize the structure based on observed data; Chapter 3 onward is prediction. To make the confidence level explicit, each prediction is tagged with a confidence level (High / Medium / Low). The conditions under which these predictions would be wrong are summarized in Chapter 7.


0. Summary

The academic system was built on the premise that “writing a plausible-sounding claim is costly.” Generative AI has destroyed exactly that premise, while the cost of verification has a floor specific to each field and barely falls at all. So the differences between fields over the next decade will be determined not by differences in AI capability, but almost entirely by “where that field’s verification-cost floor sits.”

In three lines, the conclusion is:

  1. Short term (–2027): The quantitative inflation of papers and the collapse of peer review become visible, and the system applies emergency treatment via “detection and turning submissions away at the door.” The humanities and social sciences haven’t yet felt the full brunt of the damage, but that’s not because they’re immune — they’re just later in the queue.
  2. Medium term (2028–2030): “Number of papers” effectively dies as an evaluation metric, and the unit of value shifts from the paper to “verified deliverables” and “access to bottleneck resources.” Peer review splits into a machine-audit layer and a human-judgment layer.
  3. Long term (2031–2035): Fields branch into three regimes by mode of verification. In fields where formal verification works, the paper system survives; in fields rate-limited by physical verification, the paper stops being the main battlefield; and fields where interpretation is fundamentally intersubjective (humanities, qualitative social science, legal interpretation) face the most severe institutional crisis.

1. The Driving Mechanism: The Collapse of the G/V Ratio

1.1 The academic system’s implicit premise

Let’s set up notation.

The modern academic publishing system was implicitly designed on the premise that G ≳ V. Because writing itself was costly, the mere fact that something “had been written” functioned as a weak quality assurance, and peer review only needed to sit thinly on top as a final check. The reason the system worked even though reviewers didn’t fully reproduce every manuscript is that G was high enough.

Peer review was a mechanism for socially distributing the burden of V.

1.2 What happened between 2023 and 2026

What the observed data shows is an asymmetric change: G alone dropped by more than two orders of magnitude, while V barely moved at all.

20222026Change
Generating a paper draftWeeks of human labor$4–6, a few hours (a 10-stage pipeline, 15,000 words)G at 10⁻²–10⁻³×
Independently verifying one claimHours to days of a reviewer’s timeUnchanged (if anything, increased by the extra detection work)V essentially unchanged

Everything that followed from this can be explained by the collapse of the G/V ratio.

1.3 The theoretical limit: V has a floor

This is the crux of the prediction. G will keep falling further, but V has a physical/principled floor specific to each field.

Mode of verificationWhat sets V’s floorLevel of the floor
Formal verification (e.g., Lean)Compile time, compute resourcesSeconds to hours. Falls together with AI
Numerical reproduction (code + data)Compute resources, whether data is publicMinutes to days. Falls partially
Physical experimentsReaction time, sample procurement, equipment occupancyDays to years. Doesn’t fall in principle
Clinical trialsSubjects’ lived time, ethics review, follow-up periodOn the order of years. Doesn’t fall
Cross-checking primary sourcesPhysical access to archives, reading comprehensionFalls, but interpretation remains
Validity of an interpretationFormation of intersubjective agreementAutomation is impossible in principle

In fields where V can be automated, the G/V ratio recovers; in fields where V is rate-limited by physical reality or by interpretation, it doesn’t. This single point produces every field difference that follows.


2. Branching into Three Regimes

From the above, fields split into three groups, determined not by the strength of AI but by mode of verification.

Regime A: Formal verification works (V falls together with G)

Mathematics, theoretical computer science, formal methods, some of theoretical physics

The reason AlphaProof Nexus could solve 9 Erdős problems (2 of them unsolved for 56 years) and prove 44 of 492 OEIS conjectures is that the Lean compiler, a falsification device, automates V. Here, even as G falls, V falls simultaneously, so the G/V ratio doesn’t collapse. Hallucinations get automatically weeded out in the form of “it doesn’t compile.”

The paper system survives in this group. Though where the value sits does shift (more below).

Regime B: Physical verification is the rate-limiter (V doesn’t fall)

Biology, chemistry, materials science, clinical medicine, experimental physics, earth science

Kosmos can generate a cited report in 12 hours, but human verification finds the descriptions only 80% accurate, and dataset selection and validity checking remain the premise that humans handle. The fact that GSK is paying NOETIK a $50 million upfront fee, and Lilly is paying Chai Discovery tens of millions of dollars a year, means industry has judged that scarcity sits not in generating hypotheses but on the verification side.

The paper stops being the main battlefield in this group. Value shifts to the capacity to run experiments (autonomous labs, samples, access to subjects).

Regime C: Verification is fundamentally intersubjective (V can’t be automated)

Humanities (history, philosophy, literature), qualitative social science, legal interpretation, part of theoretical economics, normative research in general

There’s no falsification device here. Only a human community can judge “is this interpretation valid,” and that judgment cost doesn’t fall with AI. Meanwhile, G has fallen just as much as in other fields. The fact that economics/finance measures at 47% in the arXiv data shows that the gateway to this group has already started to crumble.

This is the most severe group. The paper as a format stops functioning as a vehicle of trust at all.


3. Short-Term Predictions: Late 2026 – End of 2027

3.1 Common across all fields

PredictionConfidence
“Detection + mechanical rejection at the door” becomes the standard operating procedure at major conferences and publishers. The NeurIPS 2026 approach (scoring with Pangram etc., desk-rejecting anything over threshold with no chance to appeal) spreads horizontallyHigh
Caps on submission volume are introduced. Annual per-author submission caps, co-author count limits, refundable submission deposits, and the like start experimentallyHigh
Mandatory submission of edit history spreads. NeurIPS 2026 has already made “presenting an online version with edit history” a condition for lifting a conditional rejection, and this becomes standardMedium–High
The obligation to disclose AI use stays a dead letter. The picture PNAS revealed — 70% policy adoption, 0.1% disclosure rate — doesn’t improve. Norm-based control doesn’t workHigh
Several wrongful-accusation incidents from detector false positives become public disputes, concentrated especially among non-native-English speakers and early-career researchersMedium–High
Reviewer shortage hits a critical point, and journals/conferences that introduce monetary compensation for reviewing appearMedium

3.2 Short term, field by field

CS / Machine Learning — Being furthest ahead, this is the first field to enter the phase where “acceptance at a major conference stops functioning as a trust signal.” Submission counts keep rising further from NeurIPS’s 21,575 (2025), AAAI’s roughly 29,000 (2026), and ICLR’s 19,490 (2026). “Denominator gaming” (inflating submission counts rather than quality to drive down the acceptance rate and manufacture apparent prestige) becomes an openly discussed issue. The acceptance rate dies as a prestige metric within 2027. (Confidence: Medium–High)

Mathematics — Almost unaffected. But the question of authorship — “a human writes up, as a paper, a theorem AI solved” — becomes the first point of controversy. As cases of solved Erdős problems increase, consensus starts forming on who counts as the author. (Confidence: High)

Life Science / Drug Discovery — The confluence of paper mills (equivalent to 1.5–2% of papers published in 2022) and AI-generated papers deepens, and a systematic-review credibility crisis surfaces. Already, 0.15% of 200,000 reviews cite a retracted paper-mill paper. Because contaminated evidence synthesis spills over into clinical guidelines, this becomes the first field where regulators step in. (Confidence: Medium–High)

Physics / Astrophysics — Norms still hold up (33% for writing assistance, 48% for programming). But the capability gap — under 20% reproduction on ReplicationBench, 38.8% on PaperArena versus 83.5% for PhD holders — carries the risk of creating a mistaken sense of safety, a “we can’t hand this to AI” that’s actually false comfort. (Confidence: Medium)

Economics / Quantitative Social Science — The gap between 47% on the text side and 20% regular use of coding agents stabilizes temporarily as an intermediate norm of “AI writes, I analyze.” But the adoption gap (postdocs at 2× professors, male-sounding names at over 2× female-sounding names, top-25 universities at 40% higher) becomes visible as a productivity gap and turns into a fairness controversy. (Confidence: Medium–High)

Humanities / Qualitative Social ScienceStill quiet on the surface, but this isn’t immunity — it’s just a matter of sequence. What happens in the short term: (a) LLM use for qualitative coding becomes the de facto standard; (b) the IRB/ethics-review problem of feeding sensitive data into cloud LLMs surfaces; (c) reviewers start writing their reviews with LLMs (the humanities have a small reviewer pool with heavy per-review load, so the incentive is, if anything, stronger). (Confidence: Medium)

Law — Pressure from practice runs ahead of academia. Court cases involving AI-fabricated citations reached 1,598 by June 2026 (up from about 200 a year earlier), with sanction amounts up 11× in 18 months. “Did you actually read and verify the citation yourself” has become an explicit requirement of professional ethics, and that comes straight down into the norms of legal education and legal-paper writing. (Confidence: High)


4. Medium-Term Predictions: 2028–2030

4.1 The unit of value shifts away from “the paper”

This is the central change of the medium term. A paper was a “package of a claim,” but once the cost of generating a claim approaches zero, no value remains in the package itself. What remains is one of the following.

New unit of valueApplicable fieldsWhat it looks like
Verified deliverableMathematics, theoretical CS, formal methodsA repository machine-verified in Lean etc. The paper is demoted to its explanatory document
Capacity to run experimentsBiology, chemistry, materialsOccupancy time of autonomous labs, samples, equipment. SDL 2.0 becomes infrastructure
Access to subjects / clinical accessMedicineThe right to run a prospective trial. This is, in principle, irreplaceable by AI
Proprietary data and identification strategyEconomics, quantitative social scienceAccess rights to administrative/corporate data, the validity of causal identification
Access to primary sources and responsibility for interpretationHumanities, qualitative research, lawArchives, fieldwork, relationships with those concerned. Who bears responsibility

4.2 “Number of papers” dies as an evaluation metric

The 2026 OECD report and the research-evaluation reform discourse (over 20,000 DORA signatories) already say that “publication counts and citation counts are inappropriate proxies for meaningful impact.” What’s medium-term about the change is that this shifts from a matter of principle to a practical necessity. Once researchers with 200 co-authored papers a year stop being rare, the count carries no information.

What rises as a substitute (Confidence: Medium):

This last point matters, and it’s one instance of a general rule: “having come first in time” becomes a primary source of trust. The mandatory submission of edit history (short term) rests on the same logic.

4.3 The two-tiering of peer review

Peer review stops being a single step and splits into the following two layers (Confidence: Medium–High).

Layer 1: Machine Audit Statistical consistency, whether citations actually exist, compliance with reporting guidelines, whether data/code actually runs, overlap with prior publications, image manipulation. This becomes fully automated, and humans stop being involved. Passing is a necessary, not sufficient, condition for publication.

Layer 2: Human Judgment The significance of the problem set up, its positioning within the field, the validity of the interpretation, normative implications. Only this remains as human work, and it starts being compensated.

Publishers are already, at this point, pushing a design where “AI handles reviewer matching, checking against reporting standards, and summarizing key points, while humans concentrate on scientific merit” — an early form of this two-tiering. The emerging consensus between COPE and STM (discussed at WCRI 2026) is along the same lines.

4.4 The right to submit becomes a scarce resource

The mechanism arXiv has already introduced — “first-time submitters need an endorsement” — is a shift toward making the right to submit itself a scarce resource. This generalizes in the medium term, and the following appear (Confidence: Medium).

The side effect is obvious: the exclusion of unaffiliated researchers, the Global South, and early-career researchers. This is rational as a countermeasure against AI spam, yet it collides head-on with the value of academic openness. This becomes the biggest political flashpoint of the medium term. (Confidence: High)

4.5 Medium term, field by field

Mathematics — AI proof becomes routine, and the value of a paper shifts from “possession of a proof” to “invention of a problem setup.” Only the ability to decide what should be asked becomes scarce. Peer review is almost entirely replaced by formal verification, and human reviewers only ask, “is this theorem interesting?” (Confidence: Medium–High)

Life Science / Chemistry / Materials — Autonomous labs become infrastructure, and the paper comes to resemble “an audit record proving the experiment really happened.” Machine-readable submission of raw data, equipment logs, and experimental protocols becomes a submission requirement. Papers that are hypothesis-only lose value. (Confidence: Medium)

Medicine — The evidence hierarchy gets rebuilt. A rating that explicitly places “AI-synthesized findings” below in the hierarchy gets introduced. The capacity to run clinical trials becomes a decisive scarce resource, and AI gets pushed into preclinical work and evidence synthesis. (Confidence: Medium)

Physics / Astronomy — Foundation models like AION-1 become the shared infrastructure of the field, and contribution to a model/dataset is valued over an individual paper. Observation time at large facilities remains the center of value. (Confidence: Medium)

CS / Machine LearningThe conference-centric model hits its limit. Unable to recover from a state where 21% of reviews are AI-generated, some or all of the following happen: (a) a retreat to invitation-only/smaller scale; (b) authority shifting to in-house corporate evaluation and benchmarks; (c) a public track record of code and models becoming a stronger signal than the paper itself. The value of “getting into a top conference” clearly declines over the medium term. (Confidence: Medium–High)

Economics / Quantitative Social Science — Once automatic generation of analysis code becomes normal, the validity of the identification strategy and proprietary data become the only differentiators. The value of “being able to run a regression” goes to zero. Registered Reports and pre-registration standardize earliest in this field. (Confidence: Medium–High)


5. Humanities and Qualitative Social Science: A Detailed Prediction for the Most Severe Field

The reason to treat this group separately is that it’s the only one with no path to automating verification. Other fields can retreat into a machine-audit layer; the humanities have no such refuge.

5.1 Why it’s severe

The validity of a humanistic claim rests on faithfulness to a chain of details — edition, page, paragraph, footnote, marginal note, context, counterexample, position in historiography. LLMs are extremely good at “plausibly” reconstructing this chain, while whether the primary source was actually consulted cannot, in principle, be distinguished from the output.

The numbers observed in law are the proof. Even legal-specialized RAG tools hallucinate at a rate of 17–34%, and court cases involving AI-fabricated citations rose from 200 to 1,598 in a single year. Law is a rare humanities-adjacent field with an external verification mechanism — “mistakes get exposed in court” — and even there, these are the numbers. There’s no reason to think the same thing isn’t happening in history or literary studies, fields with no comparable exposure mechanism.

5.2 Short term (–2027): Quiet infiltration

5.3 Medium term (2028–2030): The institutionalization of provenance

The core prediction: since the humanities can’t verify “the correctness of a claim,” they shift instead toward verifying “the provenance of the work.” (Confidence: Medium)

A proposal in this direction has already appeared. A July 2026 arXiv paper, “Traceable Scholarship: Page Anchors and Ariadne’s Thread for Humanistic Inquiry in the Age of Generative AI,” proposes a framework that brings page-level anchors and a trace of the work into humanistic inquiry. There’s a good chance this becomes the institutional solution for the field.

The concretely expected shape:

5.4 Long term (2031–2035): A reversal of value

A paradoxical prediction: once AI can produce interpretations without limit, the value of interpretation itself falls, and the value of “who bears responsibility” rises. (Confidence: Medium)

The product of the humanities is “a reading,” but in a world where readings come infinitely cheap, the novelty of a reading stops being scarce. What remains is the following three things.

  1. Physical/legal access to primary sources (unpublished documents, fieldwork, trust relationships with those concerned)
  2. The person as the subject who bears responsibility for an interpretation (whose judgment this reading is being presented as)
  3. The judging function of the community (the human collective that decides what counts as a good reading, in itself)

In other words, the humanities are forced to change their self-definition from “the practice of producing text” to “the practice of maintaining a responsible subject of judgment.” This is a redefinition, not a shrinking, but because it’s completely inconsistent with the current evaluation system (publication counts), the transitional turmoil will be larger than in STEM.

5.5 The particular circumstances of qualitative social science

Qualitative research follows a somewhat different branch from the humanities (Confidence: Medium).


6. Long-Term Predictions: 2031–2035

6.1 Institutional reorganization: separation into three academic spheres

Over the long term, I predict that what’s currently lumped together as “academia” will, in effect, separate into three spheres with different modes of verification (Confidence: Medium).

Formal SphereEmpirical SphereInterpretive Sphere
FieldsMathematics, theoretical CS, formal methodsBiology, chemistry, materials, medicine, experimental physicsHumanities, qualitative social science, legal interpretation, normative research
Guarantee of trustMachine verificationAuditing the reality of the experimentProvenance and a responsible subject
Role of the paperCommentary on a verified deliverableSummary of experimental recordsSignature on a judgment
Peer reviewAlmost fully automatic + humans only judge significanceMachine audit + experimental auditHumans only. Becomes the most costly
Scarce resourceAbility to set up problemsCapacity to run experiments, subjectsAccess to sources, a responsible subject
AI’s positionCo-proverHypothesis generator/analyzerMaterial generator (not trusted)

Evaluation criteria across these three spheres become mutually incommensurable. Cross-field evaluation is already difficult today, but over the long term, “comparing researchers on the same footing” becomes impossible in principle, and university hiring and budget allocation are the first things to hit a breaking point. (Confidence: Medium)

6.2 The fate of “the paper” as a format

Prediction: the paper doesn’t disappear, but its function shifts from “transmitting knowledge” to “recording where responsibility lies.” (Confidence: Medium)

Historically, the paper served several functions at once — transmitting knowledge, asserting priority, proving achievement, and making clear where responsibility lies. Of these, the only one AI can’t substitute for is the last. Transmitting knowledge is replaced by AI summaries and dialogue; priority is served well enough by a timestamped registry; and proof of achievement has stopped functioning due to volume inflation.

The long-term paper therefore becomes something short, heavy on signature, revocable, and traceable in provenance. The current format of “a 30-page PDF” is likely not to survive.

6.3 The redefinition of “researcher” as a profession

Once autonomous AI systems can carry out everything from hypothesis generation through data analysis to reporting, the literature on research evaluation already argues the human role shifts to system designer, verifier, curator. Building on that, I predict (Confidence: Medium–Low):

Conversely, the current core training of a doctoral program — “read the literature and organize it, run a standard analysis, write a paper” — loses almost all professional value. Redesigning graduate education becomes the biggest unresolved problem over the long term. (Confidence: Medium–High)

6.4 The upshot for the publishing industry

Publishers’ revenue base already shows signs of a shift. Five publishers — Elsevier, Cengage, Hachette, Macmillan, and McGraw Hill — filed a joint copyright infringement lawsuit against Meta on May 5, 2026, and Cambridge University Press has rolled out an opt-in AI licensing scheme premised on author consent.

Long-term prediction (Confidence: Low–Medium): publishers’ revenue shifts from “selling papers” to “selling verification services and provenance guarantees.” Value remains in being the institution that operates the machine-audit layer, guarantees provenance, and manages retractions. Conversely, the value of the function of “typesetting and distributing manuscripts” goes to zero.


7. Conditions Under Which This Prediction Would Be Wrong

Let me make the falsifiability of this prediction explicit. If the following happen, the framework above needs revision.

Falsifying conditionImpact
AI-based verification catches up with generation — technology emerges that dramatically lowers V via automated falsification/automated reproductionThe split between Regime B and C disappears, and the paper system survives broadly. The single most important branch point
Detection technology becomes highly accurate and stable — the false-positive rate falls to a practical level, and identifying AI generation becomes reliable“Detection + rejection at the door” lets the system hold up, and the transition becomes gentler
AI capability plateaus — stalls at the current 20–40% (ReplicationBench, PaperArena)The medium-term changes get pushed back a few years, but the direction doesn’t change
Strong legal/regulatory intervention — legal responsibility for AI-generated research gets clarified, and penalties actually functionAlready happening in law (sanction amounts up 11× in 18 months). If it spills over to other fields quickly, the short-term turmoil gets suppressed
The humanities invent their own mode of verification — a solution other than provenance emergesChapter 5’s predictions change substantially. No strong candidate is visible at present

The line to watch most closely is the first one. The assumption that “V doesn’t fall” is the foundation of this entire prediction, and if it collapses, the conclusions change substantially. Put the other way: investment in automated verification technology is the most fundamental intervention available against this structural problem.


8. Practical Implications

Assuming the predictions are broadly correct, here are moves available right now, across fields.

As an individual researcher

As an educator

As someone running an institution


References